Commit Graph

5 Commits

Author SHA1 Message Date
c74fe7e074 Add an OpenAI-compatible gateway, so the model provider is a config value
The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.

`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.

`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.

THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.

Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.

Three configuration mistakes are refused at startup rather than at runtime:

  - MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
    endpoint and ignores the field, so this is a deployment that believes it
    switched providers and did not — every run still goes to Anthropic and is
    still billed there, with nothing in the logs to say so. Cost is the whole
    reason this change exists, and that is the one mistake that silently
    defeats it.
  - An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
  - A production deployment with no credential — except against localhost,
    which needs none, and demanding one would make the free local path
    impossible to configure.

reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.

`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.

Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.

NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-01 11:47:53 +05:30
57c2a52c1e Refuse an HTTP write timeout that would cut off a legal agent run
Some checks failed
CI / test (push) Failing after 5m32s
CI / fixture (push) Failing after 59s
Production answered 502 Bad Gateway on a non-streamed agent run. Nothing about
that was a gateway fault: krow-proxy already had proxy_read_timeout 3600s, and
the API pods were healthy with zero restarts throughout.

HTTP_WRITE_TIMEOUT was 30s. Every shipped agent runs at the `balanced` tier,
whose deadline is 60s, and the `deep` tier allows 120s. So the server aborted
the response on any run over half the time the runtime considered legal, the
proxy saw its upstream vanish mid-response, and it reported the only thing it
could. A gateway error for something no gateway did — which is why it looked
like infrastructure for as long as it did.

Delegation did not cause this; it made it routine. A parent that asks two
subagents takes longer than one answering alone, so a latent misconfiguration
became a reliable one. Verified: the exact request that returned 502 now
answers 200 in 18s.

Streaming is what hid it, and that is the part worth keeping in mind. The chat
panel uses SSE, so the product looked healthy while every non-streaming caller
got 502 on a slow question. A bug only reachable by the callers who do not yet
exist is one nobody reports.

So the value is now derived from the thing that constrains it — the default is
DeepestAgentDeadline plus headroom rather than a number typed once — and
validate() refuses anything below that deadline at startup. A slow,
intermittent, misattributed failure becomes a message on the first boot.

DeepestAgentDeadline is duplicated in internal/config rather than imported,
because internal/runtime already imports internal/config and a cycle to share
one number is a bad trade. TestConfigKnowsTheDeepestAgentDeadline asserts the
two agree, so drift is a build failure rather than a discovery. It also checks
that no tier exceeds it, or the name lies.

ORDERING, and it matters for the next deploy: the check refuses the old 30s, so
a pod carrying this image against an unpatched configmap will not boot.
Production's configmap is already 180s. The handover says so too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-31 12:06:26 +05:30
6849363a37 Implement delegation, so an agent's subagents are more than decoration
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
krow-workforce-agent has declared five subagents since it was written and
answered every question by itself. Everything for §6 existed except the
delegation: the parser read `subagents:`, runtime.Agent carried them, the
loader populated them, agent_runs had a parent_run_id column with a
self-reference and a no-self-parent constraint, and budget.go's comments
already described sharing a budget with subagents. Nothing called any of it.

A subagent is offered to the parent's model as a tool, because §6 says that is
what delegation is from the parent's side. Three rules are enforced rather than
assumed, each with a test that fails if it stops holding:

  I1  The subagent runs as the ORIGINAL caller. It cannot read anything the
      person could not read directly.
  §6  It SHARES the parent's budget. The test sets MaxSteps to 1, spends it in
      the parent, and asserts the child terminates BudgetExceeded — an
      assertion that only passes when the budget is shared, and that a fresh
      budget would quietly turn green.
  §3  Depth is capped at 2. At the cap no subagent is loaded or offered, so a
      cycle reaching run time is bounded rather than unbounded.

I4 survives too: a write a SUBAGENT wants approved still stops the whole run
and asks a person, rather than being performed because it happened one level
down.

Two bugs found by running it rather than by reading it:

  - delegate() read the error before the result. finish returns a non-nil
    error for every termination that is not Completed, INCLUDING
    ConfirmationPending — which is not a failure but a run that stopped to ask
    a question. Reading the error first discarded the result and with it the
    confirmation, so a subagent's write silently never happened and nobody was
    asked.

  - Delegated trajectories were never persisted at all. parent_run_id is a
    foreign key and a subagent finishes BEFORE the run that delegated to it,
    so every child insert named a parent row that did not exist yet. The
    database refused it; finish deliberately does not fail a run over a sink
    error; and the entry recording that the trajectory could not be saved was
    itself in the trajectory that was not saved. Children are now buffered and
    written by finish after the parent's own row, each arriving with its
    descendants already ordered behind it, so one pass writes a whole tree
    parent-first. The regression test asserts on save ORDER, because a
    MemorySink has no foreign key and will pass either way.

Verified end to end against a live model: an agent with no tools of its own and
one subagent produced

    delegation-probe   run=run_16622d7de6  parent=(root)
    talent-pool-agent  run=run_64160684b1  parent=run_16622d7de6

with the subagent's answer reaching the parent's model. Full suite green, only
TestLive* skipped.

Not addressed: §3's publish-time cycle detection, which needs the whole agent
set in hand. The depth cap is what holds without it, and is the half that
matters at run time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 11:59:20 +05:30
f7df96c973 agent build 2026-08-28 12:21:44 +05:30
7d12ebef3d first commit 2026-08-24 13:06:29 +05:30