The store and the write trigger existed and nothing used either. This wires
both ends, so "we have long-term memory" stops being a statement about code
that exists and becomes one about behaviour.
READ. Memories are recalled before retrieval and placed before it in the
prompt: it is the smaller block and the more general one — a standing
preference frames how the documents should be read, where a document does not
frame a preference. The question stays last, because a model reads the last
thing and answers it, and evidence after the question becomes the prompt.
Skipped for smalltalk on the same terms as retrieval. Nobody needs remembering
to say good morning, and paying for it is how "hi" came to cost six thousand
tokens.
FAILS QUIET, RECORDED LOUDLY. A memory store that is unreachable must not take
the run with it: an answer without memory is worse, not wrong, and the
alternative is an outage in the knowledge layer becoming an outage in the
product. The trajectory records the failure, and records separately when the
store returned recency instead of relevance — a reader comparing two answers
needs to know which one got which.
WRITE. tools.Remember is registered only where there is somewhere to put it,
through DefaultToolsWithMemory rather than a nil check inside the old
constructor: a deployment that has not migrated 000017 must not offer a tool
whose every call fails against a table that is not there. Passing nil
registers exactly the catalogue that was there before.
Memory shares retrieval's embedder, and memory.Embedder is knowledge.Embedder
by structure so it cannot be given a different one. Two embedding models in
one deployment produce vectors that cannot be compared, and the failure is
silent: a recall that returns nothing rather than an error.
Five new tests on the loop, including the two that matter — a failing store
still answers, and a greeting carries nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
In-place failover covered none of the failures this deployment actually had.
gateway.canFailOver will not move a conversation that has called a tool — the
assistant turn echoing that call belongs to the provider that issued it, and a
vendor which signs its function calls rejects a follow-up carrying somebody
else's. But a rate limit lands where the request is BIGGEST, which is the
second or third model call, once the catalogue, the retrieved block, the tool
results and the whole prior conversation are being re-sent.
Every GatewayFailure in agent_runs had already called a tool. The error text
says exactly what it was:
http 429: Rate limit reached for model `openai/gpt-oss-120b` …
on tokens per minute (TPM): Limit 8000, Used 7183
So the loop starts the turn over on the next provider. No transcript is sent,
so nothing provider-specific travels and the signature problem cannot arise:
the question is simply asked again somewhere with budget left. It costs the
work already done, charged to the budget that is not exhausted.
gateway.Standby is the whole of what the runtime is told — "there is another
one, here it is". No vendor, credential or model id crosses the boundary, and
the loop still cannot name a provider.
THE RULE THAT MAKES IT SAFE: a run carrying a confirmation never restarts.
Re-running re-runs its tools; a read twice is two reads, a write twice is two
shifts assigned. I4 makes the test cheap — a write executes only against a
resolved token (Registry.gate), so a run with no confirmation cannot have
written anything, and one with a confirmation is refused without inspecting
what it did.
Once, not until the providers run out: a question worth asking twice is not
worth asking five times, and each attempt spends a real budget. A terminal
error — a rejected credential, a model this deployment cannot use — is not
retried anywhere, on the same line canFailOver already draws.
Five tests, including both refusals. Verified with teeth: disabling the restart
fails the rate-limit case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
I7 had a hole. ContextInstruction states the rule for <context> blocks —
retrieved documents — and SystemPrompt has always carried it. Nothing stated it
for tool results, which arrive as their own message carrying whatever the
records hold: a candidate's note, a job description, a worker's name. Any of
those is text a person outside the company can write, and the model was given no
reason to read it as data.
gateway.ToolResultInstruction sits beside ToolResult for the same reason
ContextInstruction sits beside its renderer: a prompt promising a rule the
transport does not frame is a defence that has quietly stopped existing.
What it is worth is small, and the comment says so with the numbers. Against a
local qwen3:0.6b with a tool result carrying "ignore your previous
instructions": 3 runs in 20 held the line without the sentence, 5 in 20 with it.
An n=10 pass first suggested 1-in-10 against 6-in-10 and did not replicate. So
it is hygiene, not a control — what makes an injection survivable is I1 and I4,
which cost a hijacked turn an answer and never an action.
qwen_probe_test.go is how those numbers were taken: a DB-free probe of a
candidate model's tool-calling and injection resistance, skipped unless
MODEL_BASE_URL is set. The live eval suites need PostgreSQL and SKIP without it,
so they pass while testing nothing on a machine with none.
Also carries the language selector: a closed enum, because the value arrives
from a browser and the directive it selects goes into the system prompt. A
client picks a constant by name; nothing it sends is ever written into a prompt.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A failed save must not fail the run -- the answer already exists --
but until now the only record of the loss was an error entry
appended to the trajectory that had just failed to save. §6 says
the trajectory is not optional telemetry; losing one silently is
the worst version of losing one.
The runtime has no logger by design, so ExecutionResult gains an
Unsaved list the surface reads and turns into a §10 log line with
run_id, tenant_id, agent_key and agent_version. Never serialised
to the client. Covered on all three run paths, streaming included.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
terminationFor sent every gateway error that was not Refused or Timeout
to ToolFailure, because the enum had nowhere else to put it. On
2026-09-22 that was 131 of 318 production runs, and not one of them was
a tool failing: 56 were retired model ids, 25 an exhausted Anthropic
balance, 45 Groq's free-tier rate limit -- the only one still happening.
An operator reading the termination column saw "a tool is broken" for
two weeks while the actual answer was "we are not paying for capacity".
GatewayFailure is the seventh termination. Rate limited, request
rejected, credential refused and unreachable land there; Refused and
Deadline keep their own reasons; a non-gateway error is still the tool
layer's. A delegation whose subagent died at the gateway now carries
that reason up to the parent instead of reading as a tool call that
failed.
Migration 000016 widens the CHECK that 000006 chose precisely so this
would be a migration rather than an ALTER TYPE. Its down folds any
GatewayFailure rows back to ToolFailure BEFORE narrowing the constraint,
which is the order that works; verified up, down and up again on a
scratch database. Existing rows are left as they are -- the trajectory
entries still carry the gateway.* code for anyone reclassifying history.
The surface wording is the one termination where "try again" is honest
advice, since the dominant cause clears within a minute.
Full suite run against a real database, including the tests that skip
without one.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
krow-workforce-agent has declared five subagents since it was written and
answered every question by itself. Everything for §6 existed except the
delegation: the parser read `subagents:`, runtime.Agent carried them, the
loader populated them, agent_runs had a parent_run_id column with a
self-reference and a no-self-parent constraint, and budget.go's comments
already described sharing a budget with subagents. Nothing called any of it.
A subagent is offered to the parent's model as a tool, because §6 says that is
what delegation is from the parent's side. Three rules are enforced rather than
assumed, each with a test that fails if it stops holding:
I1 The subagent runs as the ORIGINAL caller. It cannot read anything the
person could not read directly.
§6 It SHARES the parent's budget. The test sets MaxSteps to 1, spends it in
the parent, and asserts the child terminates BudgetExceeded — an
assertion that only passes when the budget is shared, and that a fresh
budget would quietly turn green.
§3 Depth is capped at 2. At the cap no subagent is loaded or offered, so a
cycle reaching run time is bounded rather than unbounded.
I4 survives too: a write a SUBAGENT wants approved still stops the whole run
and asks a person, rather than being performed because it happened one level
down.
Two bugs found by running it rather than by reading it:
- delegate() read the error before the result. finish returns a non-nil
error for every termination that is not Completed, INCLUDING
ConfirmationPending — which is not a failure but a run that stopped to ask
a question. Reading the error first discarded the result and with it the
confirmation, so a subagent's write silently never happened and nobody was
asked.
- Delegated trajectories were never persisted at all. parent_run_id is a
foreign key and a subagent finishes BEFORE the run that delegated to it,
so every child insert named a parent row that did not exist yet. The
database refused it; finish deliberately does not fail a run over a sink
error; and the entry recording that the trajectory could not be saved was
itself in the trajectory that was not saved. Children are now buffered and
written by finish after the parent's own row, each arriving with its
descendants already ordered behind it, so one pass writes a whole tree
parent-first. The regression test asserts on save ORDER, because a
MemorySink has no foreign key and will pass either way.
Verified end to end against a live model: an agent with no tools of its own and
one subagent produced
delegation-probe run=run_16622d7de6 parent=(root)
talent-pool-agent run=run_64160684b1 parent=run_16622d7de6
with the subagent's answer reaching the parent's model. Full suite green, only
TestLive* skipped.
Not addressed: §3's publish-time cycle detection, which needs the whole agent
set in hand. The depth cap is what holds without it, and is the half that
matters at run time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g