Files
krow_backend/go-api/internal/tools/tools.go
Aravind 8bc6c23770 Spend fewer tokens per run: fewer chunks, a terser shared schema, a sane result cap
The deployment's provider ceiling is 8,000 tokens a minute and a three-call run
measured 12,123, so a single question could not fit inside a minute's budget.
That is the whole of the "the model did not answer" the chat panel has been
showing: the retry loop fires three times and the provider refuses all three.

Three cuts, measured against the real corpus and the real registry:

  DefaultK 8 -> 4            ~705 -> ~352 tokens per call
  periodSchema period help   attached to THIRTEEN tools, re-sent every call
  DefaultMaxResultBytes      262_144 -> 32_768

A three-call control-center run goes from ~12,000 to ~10,700 tokens, an 11%
cut. STATED PLAINLY BECAUSE IT IS NOT ENOUGH: that is still above 8,000, and an
earlier estimate of ~7,000 was wrong. The tool catalogue is 1,312 tokens for
seven tools — about 190 each, which is JSON Schema structure rather than
padding, so trimming prose cannot reach it. The remaining lever is giving an
agent fewer tools, and that is a decision about what the agent can answer, not
a cleanup.

The result cap is the one with no downside: 256KiB let a single tool result
outweigh everything else in the prompt put together. 32KiB is ~8,000 tokens,
still more evidence than one answer needs.

DefaultK is a real trade: half the evidence behind a grounded answer. The corpus
is 43 chunks, so four is still ~10% of it per query, and the eval suites pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-06 13:10:13 +05:30

192 lines
7.6 KiB
Go

// Package tools is the tool layer: everything an agent can do that is not
// talking.
//
// A tool is a named, schema'd function the model may call. The contract below
// is §4's, and three parts of it are load-bearing rather than stylistic:
//
// - **Every handler authorizes on ctx.Principal, first line.** A handler that
// reads ctx for anything except authorization is wrong. This is where I1
// lives: an agent may read exactly what its caller could read directly, and
// the only way to guarantee that is for the tool — not the model, not the
// prompt — to apply the caller's own permissions.
// - **Filters are pre-filters.** A handler narrows in SQL, never in Go over a
// fetched result set. I2: post-filtering leaks through counts and ranking
// positions even when no forbidden row is ever printed.
// - **Errors are returned, not raised.** A tool that panics or returns a bare
// error takes the whole run with it. The runtime decides whether the model
// sees a failure and retries, and it can only decide that if the failure
// arrives as data.
package tools
import (
"context"
"encoding/json"
"fmt"
"github.com/krow/krow-backend/go-api/internal/authctx"
)
// Effect is whether a tool changes anything.
type Effect string
const (
// EffectRead observes. Safe to call without asking anyone.
EffectRead Effect = "read"
// EffectWrite writes, sends, deletes, charges or notifies. Never runs
// without a resolved confirmation — see Tool.RequiresConfirmation.
EffectWrite Effect = "write"
)
// DefaultMaxResultBytes caps a tool result.
//
// Not a performance guard. An unbounded result is an unbounded prompt on the
// next turn, which is an unbounded bill and eventually a context overflow that
// presents as the model ignoring the middle of its own evidence.
// 32KiB is roughly 8,000 tokens — already more evidence than any one answer
// needs, and an order of magnitude below the 256KiB this used to be. That old
// ceiling let ONE result outweigh everything else in the prompt put together,
// on a deployment whose provider ceiling is 8,000 tokens a minute.
const DefaultMaxResultBytes = 32_768
// Context is what a handler is given about its caller.
//
// Carries the principal, the tenant, the run and what budget is left, per §4.
// Deliberately a struct and not a context.Context value: a handler must not be
// able to *forget* to read it, and a compile error is a better reminder than a
// convention.
type Context struct {
// Principal is the caller the agent is acting for. Never the agent.
Principal authctx.Identity
// RunID addresses the trajectory this call is recorded in.
RunID string
// RemainingTokens is what the run has left to spend. A handler may use it
// to decide how much to return; it must not use it to decide whether the
// caller is allowed something.
RemainingTokens int64
// Confirmation is the resolved token for a write. Empty on a read, and
// empty on a write that has not been confirmed yet — which the dispatcher
// refuses before a handler is ever reached.
Confirmation string
// AgentID is the running agent, for the record. Never for authorization —
// what a caller may do is decided by their principal, and an agent that
// could widen that by being named would be an agent that expands access.
AgentID string
// KnowledgeSources are the corpora the running agent's SPEC declares.
//
// Here rather than in a tool argument, and the difference is the whole
// security property: an argument is something a model can choose, and which
// documents an agent may read is not the model's to choose. The loop sets
// this from the agent record; nothing in the conversation can reach it.
//
// Empty means this agent has no knowledge. It does not mean "all of it" —
// retrieval refuses an empty source list for exactly that reason.
KnowledgeSources []string
}
// OrgID is the tenant this call runs inside. I5 — every handler's query
// narrows by it, and there is no path that produces a call without one.
func (c Context) OrgID() string { return c.Principal.OrgID }
// Handler runs one tool.
//
// The signature is `(inputs, ctx)` in §4's terms, with the Go context first by
// convention so cancellation and the deadline reach the query. A handler
// returns a Result and never an error: a failure is a value the runtime routes,
// not a panic that ends a run.
type Handler func(ctx context.Context, tc Context, inputs json.RawMessage) Result
// Tool is one callable capability.
type Tool struct {
Name string
Description string
// InputSchema is JSON Schema. Every field described, because the
// description is what the model reads instead of documentation — and a
// tool whose schema demands an id the model was never given is a design
// bug, not a prompt problem. Add a lookup tool instead.
InputSchema map[string]any
Effect Effect
// RequiresConfirmation is forced true for a write by Register. It is a
// field rather than a method so a read tool may opt in — some reads are
// expensive enough to be worth asking about — but a write can never opt
// out. The model does not get a say either way.
RequiresConfirmation bool
// Confirm renders, in plain language, what this tool will do if approved.
// Mandatory when RequiresConfirmation is set: Register refuses a write
// without one, because a confirmation a person cannot read is not a
// confirmation, it is a click. See confirm.go.
Confirm Confirmer
MaxResultBytes int
Handler Handler
}
// Result is what a tool returns.
//
// Structured data, never prose: formatting is the model's job, and a handler
// that returns a sentence has decided how the answer reads before the model has
// seen the question.
type Result struct {
Data any `json:"data,omitempty"`
Error *ToolError `json:"error,omitempty"`
// Truncated says the result was cut at MaxResultBytes. Set alongside the
// data that survived — never as a silent drop, because a model given a
// truncated list with no marker will reason about it as if it were whole.
Truncated bool `json:"truncated,omitempty"`
// Confirmation is set when a write was described but not performed. It is
// neither success nor error: nothing happened, and something must now be
// approved by a person before anything can. The loop reads this and ends
// the run at ConfirmationPending rather than handing it to the model —
// see I4, and the note on Registry.Dispatch.
Confirmation *Confirmation `json:"confirmation,omitempty"`
}
// ToolError is a failure a handler chose to report.
type ToolError struct {
Code string `json:"code"`
Message string `json:"message"`
}
// Failf builds an error result.
func Failf(code, format string, args ...any) Result {
return Result{Error: &ToolError{Code: code, Message: fmt.Sprintf(format, args...)}}
}
// OK builds a success result.
func OK(data any) Result { return Result{Data: data} }
// Standard tool error codes. A denial is deliberately one code with one
// wording — see Denied.
const (
CodeDenied = "tool.denied"
CodeInvalidInput = "tool.invalid_input"
CodeUnavailable = "tool.unavailable"
CodeFailed = "tool.failed"
)
// Denied is the single refusal every handler returns when a caller may not do
// something.
//
// One code, one message, no detail. §8: a denial must not reveal that the
// resource exists. Two different refusals — "no such venue" and "not your
// venue" — are an oracle: a caller who can tell them apart can enumerate what
// they cannot see, and the agent will happily run that enumeration for them one
// question at a time.
func Denied() Result {
return Result{Error: &ToolError{
Code: CodeDenied,
Message: "the caller does not have access to this",
}}
}