The deployment's provider ceiling is 8,000 tokens a minute and a three-call run
measured 12,123, so a single question could not fit inside a minute's budget.
That is the whole of the "the model did not answer" the chat panel has been
showing: the retry loop fires three times and the provider refuses all three.
Three cuts, measured against the real corpus and the real registry:
DefaultK 8 -> 4 ~705 -> ~352 tokens per call
periodSchema period help attached to THIRTEEN tools, re-sent every call
DefaultMaxResultBytes 262_144 -> 32_768
A three-call control-center run goes from ~12,000 to ~10,700 tokens, an 11%
cut. STATED PLAINLY BECAUSE IT IS NOT ENOUGH: that is still above 8,000, and an
earlier estimate of ~7,000 was wrong. The tool catalogue is 1,312 tokens for
seven tools — about 190 each, which is JSON Schema structure rather than
padding, so trimming prose cannot reach it. The remaining lever is giving an
agent fewer tools, and that is a decision about what the agent can answer, not
a cleanup.
The result cap is the one with no downside: 256KiB let a single tool result
outweigh everything else in the prompt put together. 32KiB is ~8,000 tokens,
still more evidence than one answer needs.
DefaultK is a real trade: half the evidence behind a grounded answer. The corpus
is 43 chunks, so four is still ~10% of it per query, and the eval suites pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
§9 says no agent ships without evals. Eight of the nine had none: the two
other suites in evals/ are harness fixtures rather than agents in the
registry, so the rule was being met by one agent in nine.
Evals — 40 new cases, five per agent, every one carrying mustNotLeak:
- the agent is loaded from its real spec in agents/*.md rather than
written out again in Go. A hand-copied agent tests the copy: it keeps
passing after somebody edits the spec, which is the moment it most
needed to fail.
- callNamed calls the tool a case names. toolThenAnswer always called
tools[0], so seven of positions-agent's eight tools were unreachable,
and a boundary nothing calls is a boundary nothing tests.
- seedWorkspace fills BOTH tenants. A leak test against an empty second
tenant cannot fail.
Verified by breaking workersByScore's org predicate: six cases across four
agents fail with LEAKED "RIVAL".
Knowledge — six policy documents, taking the corpus from 2 to 8 (34
chunks). Three restricted to admin and employer, five tenant-wide. They
cover what the tools cannot: a tool reports how many shifts went unworked,
a policy says what cover costs inside 24 hours.
corpus_test.go treats those documents as product rather than fixtures. The
first version was tautological — it read audience: from a file and checked
that file's audience was enforced, so opening a restricted document passed.
mustNotBeTenantWide now holds that judgement apart from the files, with the
reason recorded for each.
CI — the checks this repository already had, made unskippable. testutil
calls t.Skipf on an unreachable database, so a dead service container would
produce a green build over a suite that ran almost nothing. Simulated: go
test exits 0 with 74 tests skipped, including every tenant-isolation test.
The guard exits 1 and names them, while still allowing TestLive* to skip
without a model key.
This CI tests; it does not deploy. The README's claim that migrations are
run by CI against the target database remains aspirational.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu