I7 had a hole. ContextInstruction states the rule for <context> blocks —
retrieved documents — and SystemPrompt has always carried it. Nothing stated it
for tool results, which arrive as their own message carrying whatever the
records hold: a candidate's note, a job description, a worker's name. Any of
those is text a person outside the company can write, and the model was given no
reason to read it as data.
gateway.ToolResultInstruction sits beside ToolResult for the same reason
ContextInstruction sits beside its renderer: a prompt promising a rule the
transport does not frame is a defence that has quietly stopped existing.
What it is worth is small, and the comment says so with the numbers. Against a
local qwen3:0.6b with a tool result carrying "ignore your previous
instructions": 3 runs in 20 held the line without the sentence, 5 in 20 with it.
An n=10 pass first suggested 1-in-10 against 6-in-10 and did not replicate. So
it is hygiene, not a control — what makes an injection survivable is I1 and I4,
which cost a hijacked turn an answer and never an action.
qwen_probe_test.go is how those numbers were taken: a DB-free probe of a
candidate model's tool-calling and injection resistance, skipped unless
MODEL_BASE_URL is set. The live eval suites need PostgreSQL and SKIP without it,
so they pass while testing nothing on a machine with none.
Also carries the language selector: a closed enum, because the value arrives
from a browser and the directive it selects goes into the system prompt. A
client picks a constant by name; nothing it sends is ever written into a prompt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A failed save must not fail the run -- the answer already exists --
but until now the only record of the loss was an error entry
appended to the trajectory that had just failed to save. §6 says
the trajectory is not optional telemetry; losing one silently is
the worst version of losing one.
The runtime has no logger by design, so ExecutionResult gains an
Unsaved list the surface reads and turns into a §10 log line with
run_id, tenant_id, agent_key and agent_version. Never serialised
to the client. Covered on all three run paths, streaming included.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
terminationFor sent every gateway error that was not Refused or Timeout
to ToolFailure, because the enum had nowhere else to put it. On
2026-09-22 that was 131 of 318 production runs, and not one of them was
a tool failing: 56 were retired model ids, 25 an exhausted Anthropic
balance, 45 Groq's free-tier rate limit -- the only one still happening.
An operator reading the termination column saw "a tool is broken" for
two weeks while the actual answer was "we are not paying for capacity".
GatewayFailure is the seventh termination. Rate limited, request
rejected, credential refused and unreachable land there; Refused and
Deadline keep their own reasons; a non-gateway error is still the tool
layer's. A delegation whose subagent died at the gateway now carries
that reason up to the parent instead of reading as a tool call that
failed.
Migration 000016 widens the CHECK that 000006 chose precisely so this
would be a migration rather than an ALTER TYPE. Its down folds any
GatewayFailure rows back to ToolFailure BEFORE narrowing the constraint,
which is the order that works; verified up, down and up again on a
scratch database. Existing rows are left as they are -- the trajectory
entries still carry the gateway.* code for anyone reclassifying history.
The surface wording is the one termination where "try again" is honest
advice, since the dominant cause clears within a minute.
Full suite run against a real database, including the tests that skip
without one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
krow-workforce-agent has declared five subagents since it was written and
answered every question by itself. Everything for §6 existed except the
delegation: the parser read `subagents:`, runtime.Agent carried them, the
loader populated them, agent_runs had a parent_run_id column with a
self-reference and a no-self-parent constraint, and budget.go's comments
already described sharing a budget with subagents. Nothing called any of it.
A subagent is offered to the parent's model as a tool, because §6 says that is
what delegation is from the parent's side. Three rules are enforced rather than
assumed, each with a test that fails if it stops holding:
I1 The subagent runs as the ORIGINAL caller. It cannot read anything the
person could not read directly.
§6 It SHARES the parent's budget. The test sets MaxSteps to 1, spends it in
the parent, and asserts the child terminates BudgetExceeded — an
assertion that only passes when the budget is shared, and that a fresh
budget would quietly turn green.
§3 Depth is capped at 2. At the cap no subagent is loaded or offered, so a
cycle reaching run time is bounded rather than unbounded.
I4 survives too: a write a SUBAGENT wants approved still stops the whole run
and asks a person, rather than being performed because it happened one level
down.
Two bugs found by running it rather than by reading it:
- delegate() read the error before the result. finish returns a non-nil
error for every termination that is not Completed, INCLUDING
ConfirmationPending — which is not a failure but a run that stopped to ask
a question. Reading the error first discarded the result and with it the
confirmation, so a subagent's write silently never happened and nobody was
asked.
- Delegated trajectories were never persisted at all. parent_run_id is a
foreign key and a subagent finishes BEFORE the run that delegated to it,
so every child insert named a parent row that did not exist yet. The
database refused it; finish deliberately does not fail a run over a sink
error; and the entry recording that the trajectory could not be saved was
itself in the trajectory that was not saved. Children are now buffered and
written by finish after the parent's own row, each arriving with its
descendants already ordered behind it, so one pass writes a whole tree
parent-first. The regression test asserts on save ORDER, because a
MemorySink has no foreign key and will pass either way.
Verified end to end against a live model: an agent with no tools of its own and
one subagent produced
delegation-probe run=run_16622d7de6 parent=(root)
talent-pool-agent run=run_64160684b1 parent=run_16622d7de6
with the subagent's answer reaching the parent's model. Full suite green, only
TestLive* skipped.
Not addressed: §3's publish-time cycle detection, which needs the whole agent
set in hand. The depth cap is what holds without it, and is the half that
matters at run time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g