Spend fewer tokens per run: fewer chunks, a terser shared schema, a sane result cap
The deployment's provider ceiling is 8,000 tokens a minute and a three-call run measured 12,123, so a single question could not fit inside a minute's budget. That is the whole of the "the model did not answer" the chat panel has been showing: the retry loop fires three times and the provider refuses all three. Three cuts, measured against the real corpus and the real registry: DefaultK 8 -> 4 ~705 -> ~352 tokens per call periodSchema period help attached to THIRTEEN tools, re-sent every call DefaultMaxResultBytes 262_144 -> 32_768 A three-call control-center run goes from ~12,000 to ~10,700 tokens, an 11% cut. STATED PLAINLY BECAUSE IT IS NOT ENOUGH: that is still above 8,000, and an earlier estimate of ~7,000 was wrong. The tool catalogue is 1,312 tokens for seven tools — about 190 each, which is JSON Schema structure rather than padding, so trimming prose cannot reach it. The remaining lever is giving an agent fewer tools, and that is a decision about what the agent can answer, not a cleanup. The result cap is the one with no downside: 256KiB let a single tool result outweigh everything else in the prompt put together. 32KiB is ~8,000 tokens, still more evidence than one answer needs. DefaultK is a real trade: half the evidence behind a grounded answer. The corpus is 43 chunks, so four is still ~10% of it per query, and the eval suites pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -42,7 +42,11 @@ const (
|
||||
// Not a performance guard. An unbounded result is an unbounded prompt on the
|
||||
// next turn, which is an unbounded bill and eventually a context overflow that
|
||||
// presents as the model ignoring the middle of its own evidence.
|
||||
const DefaultMaxResultBytes = 262_144
|
||||
// 32KiB is roughly 8,000 tokens — already more evidence than any one answer
|
||||
// needs, and an order of magnitude below the 256KiB this used to be. That old
|
||||
// ceiling let ONE result outweigh everything else in the prompt put together,
|
||||
// on a deployment whose provider ceiling is 8,000 tokens a minute.
|
||||
const DefaultMaxResultBytes = 32_768
|
||||
|
||||
// Context is what a handler is given about its caller.
|
||||
//
|
||||
|
||||
Reference in New Issue
Block a user