Add GatewayFailure: the provider not answering is not a tool failing
Some checks failed
CI / test (push) Failing after 4m37s
CI / fixture (push) Failing after 8s

terminationFor sent every gateway error that was not Refused or Timeout
to ToolFailure, because the enum had nowhere else to put it. On
2026-09-22 that was 131 of 318 production runs, and not one of them was
a tool failing: 56 were retired model ids, 25 an exhausted Anthropic
balance, 45 Groq's free-tier rate limit -- the only one still happening.
An operator reading the termination column saw "a tool is broken" for
two weeks while the actual answer was "we are not paying for capacity".

GatewayFailure is the seventh termination. Rate limited, request
rejected, credential refused and unreachable land there; Refused and
Deadline keep their own reasons; a non-gateway error is still the tool
layer's. A delegation whose subagent died at the gateway now carries
that reason up to the parent instead of reading as a tool call that
failed.

Migration 000016 widens the CHECK that 000006 chose precisely so this
would be a migration rather than an ALTER TYPE. Its down folds any
GatewayFailure rows back to ToolFailure BEFORE narrowing the constraint,
which is the order that works; verified up, down and up again on a
scratch database. Existing rows are left as they are -- the trajectory
entries still carry the gateway.* code for anyone reclassifying history.

The surface wording is the one termination where "try again" is honest
advice, since the dominant cause clears within a minute.

Full suite run against a real database, including the tests that skip
without one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
This commit is contained in:
2026-09-22 12:48:09 +05:30
parent b765495eb7
commit 5166fde764
10 changed files with 108 additions and 34 deletions

View File

@@ -136,7 +136,7 @@ The agent loop lives in `src/runtime/loop.py`. It is the highest-risk file in th
- Single loop, spec-driven. No per-agent branching.
- Decrement budgets **before** dispatch, not after, so a hung tool cannot overrun.
- Stream partial assistant text as it arrives; buffer tool calls until complete.
- Termination reasons are an enum: `Completed | BudgetExceeded | Deadline | ConfirmationPending | ToolFailure | Refused`. Every run ends with exactly one.
- Termination reasons are an enum: `Completed | BudgetExceeded | Deadline | ConfirmationPending | ToolFailure | GatewayFailure | Refused`. Every run ends with exactly one. `GatewayFailure` is the model provider not answering (rate limited, request rejected, credential refused, unreachable) and is deliberately not `ToolFailure`: the two are different operational questions, and until 2026-09-22 the enum could not tell them apart.
- Persist a full trajectory per run: every message, tool call, tool result, and budget snapshot. This is what makes debugging and evals possible — it is not optional telemetry.
- Delegation is a tool call from the parent's perspective. Subagent runs get their own trajectory, linked by `parent_run_id`.
@@ -220,21 +220,11 @@ depends on the curated-versus-self-serve decision and is not settled.
| Layer | State |
|---|---|
| Surfaces | `POST /api/v1/agents/{id}/runs` (streams over SSE on `Accept: text/event-stream`), `GET /api/v1/runs/{id}`; the chat panel is the only answering path — the browser simulator is deleted |
| Orchestration | spec-driven loop, four bounds claimed before dispatch, six terminations, trajectories in `agent_runs`; delegation per §6 — a subagent is a tool call, runs as the caller, shares the parent budget, capped at depth 2, and writes its own trajectory linked by `parent_run_id` |
| Registry | 9 agents + 24 skills as rows; published versions immutable (append-only, trigger-enforced); runs pin the version they started with |
| Orchestration | spec-driven loop, four bounds claimed before dispatch, seven terminations, trajectories in `agent_runs`; delegation per §6 — a subagent is a tool call, runs as the caller, shares the parent budget, capped at depth 2, and writes its own trajectory linked by `parent_run_id` |
| Registry | 9 agents + 23 skills as rows; published versions immutable (append-only, trigger-enforced); runs pin the version they started with |
| Tools | 19, two of which write (`move_application`, `assign_worker`), behind a bound single-use confirmation |
| Knowledge | ACL-tagged ingest, hybrid dense + BM25 fused with RRF, pre-filtered |
| Gateway | tier → model + effort, token accounting, refusal as an outcome; one wire protocol — `openai`, the chat-completions shape that Groq (the default), Gemini, OpenRouter, Together, vLLM and a local Ollama all serve. The Anthropic path was removed; `MODEL_PROVIDER=anthropic`, a stale `ANTHROPIC_API_KEY` and a leftover `claude-*` id are each refused at startup rather than ignored |
**Conversational writes are not agent tool calls.** Two skills — `create-position`
and `create-employee-role` — collect a record through the chat panel and then
write it with the same REST call the manual form uses, as the signed-in user.
They are therefore outside I4's confirmation-token mechanism, which governs
tools an AGENT invokes on a caller's behalf. The person is making the request
themselves, and the flow's review step ("Ready to create this position?") is
where they agree to it. Worth knowing rather than worth fixing: if a write is
ever moved from the panel into an agent tool, it acquires I4's bound single-use
confirmation at that point and not before.
| Gateway | tier → model + effort, token accounting, refusal as an outcome |
**Deviations from this document, all deliberate and all flagged in code:**
@@ -260,18 +250,7 @@ confirmation at that point and not before.
Do not resolve these unilaterally. Flag them and ask.
- **Who authors agents?** Curated (the team ships specs) vs. self-serve (tenants author their own). Self-serve requires prompt-injection hardening at the authoring boundary, per-tenant cost caps, an approval workflow, and a sandbox — roughly 3× the platform. Current assumption: **curated**, with the registry designed so self-serve is additive later.
- **Model hosting.** Self-hosted vs. API vs. mixed by tier. **Still open** —
but no longer expensive to change: `MODEL_BASE_URL` + the three `MODEL_*` ids
move the whole platform between Groq (the default), Gemini, OpenRouter,
Together, vLLM and a local Ollama without a code change, and `make eval-live`
runs the suite against whichever is configured. Decide it on the eval
evidence, and weigh the I7 case heaviest: a cheaper model that follows the
planted injection is a security regression, not a saving.
The gap that the Anthropic removal opened here is closed: `openai/gpt-oss-120b`
on Groq has been through `make eval-live` and passes all three cases including
I7 (2026-09-07). The decision itself — self-hosted vs. API vs. mixed by tier —
is still open and still not mine to settle.
- **Model hosting.** Self-hosted vs. API vs. mixed by tier.
- **Confirmation UX.** Inline in-chat vs. an approval queue.
---