The ids always travelled TO the model; nothing a reader could look at ever
travelled back. So the panel stripped the citations the model wrote — there was
nowhere to put them — and a grounded answer became indistinguishable from an
invented one, which is the opposite of what citing is for.
A run now carries its sources: the id the model was told to cite, the document
title, the heading, and the opening of the passage. Both the JSON and the
streamed paths return them, because both build the same response.
Source is NOT knowledge.Result. That type carries ranks, scores and the whole
chunk, which exist to debug a retrieval rather than to be shown: an RRF score
is a rank and would be read as a percentage, and the full text would make the
response larger than the answer. This is the subset a citation needs.
Snippets are capped at 240 runes. A reader checking "the policy says X" needs
to recognise the passage, not to receive the corpus one answer at a time.
A run that retrieved nothing carries nothing rather than an empty list, so the
panel has no heading to draw for absent evidence — which is most runs, since
seven of nine agents answer from tools.
Next, and only now possible: the panel can render these and stop stripping the
citations that point at them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
I7 had a hole. ContextInstruction states the rule for <context> blocks —
retrieved documents — and SystemPrompt has always carried it. Nothing stated it
for tool results, which arrive as their own message carrying whatever the
records hold: a candidate's note, a job description, a worker's name. Any of
those is text a person outside the company can write, and the model was given no
reason to read it as data.
gateway.ToolResultInstruction sits beside ToolResult for the same reason
ContextInstruction sits beside its renderer: a prompt promising a rule the
transport does not frame is a defence that has quietly stopped existing.
What it is worth is small, and the comment says so with the numbers. Against a
local qwen3:0.6b with a tool result carrying "ignore your previous
instructions": 3 runs in 20 held the line without the sentence, 5 in 20 with it.
An n=10 pass first suggested 1-in-10 against 6-in-10 and did not replicate. So
it is hygiene, not a control — what makes an injection survivable is I1 and I4,
which cost a hijacked turn an answer and never an action.
qwen_probe_test.go is how those numbers were taken: a DB-free probe of a
candidate model's tool-calling and injection resistance, skipped unless
MODEL_BASE_URL is set. The live eval suites need PostgreSQL and SKIP without it,
so they pass while testing nothing on a machine with none.
Also carries the language selector: a closed enum, because the value arrives
from a browser and the directive it selects goes into the system prompt. A
client picks a constant by name; nothing it sends is ever written into a prompt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A failed save must not fail the run -- the answer already exists --
but until now the only record of the loss was an error entry
appended to the trajectory that had just failed to save. §6 says
the trajectory is not optional telemetry; losing one silently is
the worst version of losing one.
The runtime has no logger by design, so ExecutionResult gains an
Unsaved list the surface reads and turns into a §10 log line with
run_id, tenant_id, agent_key and agent_version. Never serialised
to the client. Covered on all three run paths, streaming included.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
terminationFor sent every gateway error that was not Refused or Timeout
to ToolFailure, because the enum had nowhere else to put it. On
2026-09-22 that was 131 of 318 production runs, and not one of them was
a tool failing: 56 were retired model ids, 25 an exhausted Anthropic
balance, 45 Groq's free-tier rate limit -- the only one still happening.
An operator reading the termination column saw "a tool is broken" for
two weeks while the actual answer was "we are not paying for capacity".
GatewayFailure is the seventh termination. Rate limited, request
rejected, credential refused and unreachable land there; Refused and
Deadline keep their own reasons; a non-gateway error is still the tool
layer's. A delegation whose subagent died at the gateway now carries
that reason up to the parent instead of reading as a tool call that
failed.
Migration 000016 widens the CHECK that 000006 chose precisely so this
would be a migration rather than an ALTER TYPE. Its down folds any
GatewayFailure rows back to ToolFailure BEFORE narrowing the constraint,
which is the order that works; verified up, down and up again on a
scratch database. Existing rows are left as they are -- the trajectory
entries still carry the gateway.* code for anyone reclassifying history.
The surface wording is the one termination where "try again" is honest
advice, since the dominant cause clears within a minute.
Full suite run against a real database, including the tests that skip
without one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g