The write trigger, which was the open question. Three answers were available
and two are worse.
A second model call after each run, asked to extract durable facts, judges well
and costs an entire extra call against a ceiling of 8,000 tokens a minute — on
every run, most of which have nothing worth keeping. A heuristic in the loop is
cheap and remembers the wrong things: the shape of a run says nothing about
whether a fact outlives it, so the table fills with restatements of rows the
database already holds.
So: a tool. It costs nothing extra, because the model is already mid-run with a
catalogue in front of it and remembering is one more call it may make when it
has just learned something that will not be in the records next time. It is
automatic in the sense that matters — nobody types "remember this" — and it is
visible in the trajectory, which an extraction pass would not be.
IT IS A CONFIRMED WRITE, and that is the invariant working rather than an
obstacle. EffectWrite forces RequiresConfirmation, and this tool stores
personal data that will shape later hiring answers — the most consequential
thing a model can do here short of assigning somebody to a shift. A reader sees
the sentence before it is kept, who it is about, that an agent and not a person
decided it, and that it expires in ninety days. A memory about a person also
carries a warning that says so and says it can be erased.
If a deployment later wants workspace facts kept without asking, the honest
change is a SECOND tool scoped to workspace subjects. Loosening this one would
quietly make personal memories unconfirmed too, which is the whole thing this
gate is for.
Author is always "model" and is not a field the model can set: an inference
must never be readable later as though a person had written it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The deployment's provider ceiling is 8,000 tokens a minute and a three-call run
measured 12,123, so a single question could not fit inside a minute's budget.
That is the whole of the "the model did not answer" the chat panel has been
showing: the retry loop fires three times and the provider refuses all three.
Three cuts, measured against the real corpus and the real registry:
DefaultK 8 -> 4 ~705 -> ~352 tokens per call
periodSchema period help attached to THIRTEEN tools, re-sent every call
DefaultMaxResultBytes 262_144 -> 32_768
A three-call control-center run goes from ~12,000 to ~10,700 tokens, an 11%
cut. STATED PLAINLY BECAUSE IT IS NOT ENOUGH: that is still above 8,000, and an
earlier estimate of ~7,000 was wrong. The tool catalogue is 1,312 tokens for
seven tools — about 190 each, which is JSON Schema structure rather than
padding, so trimming prose cannot reach it. The remaining lever is giving an
agent fewer tools, and that is a decision about what the agent can answer, not
a cleanup.
The result cap is the one with no downside: 256KiB let a single tool result
outweigh everything else in the prompt put together. 32KiB is ~8,000 tokens,
still more evidence than one answer needs.
DefaultK is a real trade: half the evidence behind a grounded answer. The corpus
is 43 chunks, so four is still ~10% of it per query, and the eval suites pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>