Spend fewer tokens per run: fewer chunks, a terser shared schema, a sane result cap

The deployment's provider ceiling is 8,000 tokens a minute and a three-call run
measured 12,123, so a single question could not fit inside a minute's budget.
That is the whole of the "the model did not answer" the chat panel has been
showing: the retry loop fires three times and the provider refuses all three.

Three cuts, measured against the real corpus and the real registry:

  DefaultK 8 -> 4            ~705 -> ~352 tokens per call
  periodSchema period help   attached to THIRTEEN tools, re-sent every call
  DefaultMaxResultBytes      262_144 -> 32_768

A three-call control-center run goes from ~12,000 to ~10,700 tokens, an 11%
cut. STATED PLAINLY BECAUSE IT IS NOT ENOUGH: that is still above 8,000, and an
earlier estimate of ~7,000 was wrong. The tool catalogue is 1,312 tokens for
seven tools — about 190 each, which is JSON Schema structure rather than
padding, so trimming prose cannot reach it. The remaining lever is giving an
agent fewer tools, and that is a decision about what the agent can answer, not
a cleanup.

The result cap is the one with no downside: 256KiB let a single tool result
outweigh everything else in the prompt put together. 32KiB is ~8,000 tokens,
still more evidence than one answer needs.

DefaultK is a real trade: half the evidence behind a grounded answer. The corpus
is 43 chunks, so four is still ~10% of it per query, and the eval suites pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-10-06 13:10:13 +05:30
parent 7120efa417
commit 8bc6c23770
3 changed files with 11 additions and 4 deletions

View File

@@ -49,7 +49,7 @@ const RRFConstant = 60.0
const CandidateMultiple = 3 const CandidateMultiple = 3
// DefaultK is how many chunks a retrieval returns when the caller does not say. // DefaultK is how many chunks a retrieval returns when the caller does not say.
const DefaultK = 8 const DefaultK = 4
// MaxK is the ceiling. Not a performance guard — a context guard. Retrieved // MaxK is the ceiling. Not a performance guard — a context guard. Retrieved
// text is prompt, prompt is money, and a caller asking for 500 chunks has made // text is prompt, prompt is money, and a caller asking for 500 chunks has made

View File

@@ -42,7 +42,11 @@ const (
// Not a performance guard. An unbounded result is an unbounded prompt on the // Not a performance guard. An unbounded result is an unbounded prompt on the
// next turn, which is an unbounded bill and eventually a context overflow that // next turn, which is an unbounded bill and eventually a context overflow that
// presents as the model ignoring the middle of its own evidence. // presents as the model ignoring the middle of its own evidence.
const DefaultMaxResultBytes = 262_144 // 32KiB is roughly 8,000 tokens — already more evidence than any one answer
// needs, and an order of magnitude below the 256KiB this used to be. That old
// ceiling let ONE result outweigh everything else in the prompt put together,
// on a deployment whose provider ceiling is 8,000 tokens a minute.
const DefaultMaxResultBytes = 32_768
// Context is what a handler is given about its caller. // Context is what a handler is given about its caller.
// //

View File

@@ -69,8 +69,11 @@ func periodSchema(limitHelp string) map[string]any {
"period": map[string]any{ "period": map[string]any{
"type": "string", "type": "string",
"enum": []string{"today", "last-7-days", "last-30-days", "this-month", "previous-month"}, "enum": []string{"today", "last-7-days", "last-30-days", "this-month", "previous-month"},
"description": "The window to read. Omit for all recorded history. " + // Terse on purpose: this schema is attached to thirteen tools and
"Windows are computed from the current date; do not pass a date.", // the whole catalogue is re-sent on EVERY model call, so a
// sentence here is paid for once per tool per call. The "do not
// pass a date" warning is enforced by the enum anyway.
"description": "The window to read. Omit for all history.",
}, },
"limit": map[string]any{ "limit": map[string]any{
"type": "integer", "minimum": 1, "maximum": 100, "type": "integer", "minimum": 1, "maximum": 100,