terminationFor sent every gateway error that was not Refused or Timeout to ToolFailure, because the enum had nowhere else to put it. On 2026-09-22 that was 131 of 318 production runs, and not one of them was a tool failing: 56 were retired model ids, 25 an exhausted Anthropic balance, 45 Groq's free-tier rate limit -- the only one still happening. An operator reading the termination column saw "a tool is broken" for two weeks while the actual answer was "we are not paying for capacity". GatewayFailure is the seventh termination. Rate limited, request rejected, credential refused and unreachable land there; Refused and Deadline keep their own reasons; a non-gateway error is still the tool layer's. A delegation whose subagent died at the gateway now carries that reason up to the parent instead of reading as a tool call that failed. Migration 000016 widens the CHECK that 000006 chose precisely so this would be a migration rather than an ALTER TYPE. Its down folds any GatewayFailure rows back to ToolFailure BEFORE narrowing the constraint, which is the order that works; verified up, down and up again on a scratch database. Existing rows are left as they are -- the trajectory entries still carry the gateway.* code for anyone reclassifying history. The surface wording is the one termination where "try again" is honest advice, since the dominant cause clears within a minute. Full suite run against a real database, including the tests that skip without one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
247 lines
8.3 KiB
Go
247 lines
8.3 KiB
Go
package runtime
|
|
|
|
import (
|
|
"context"
|
|
"fmt"
|
|
"sync"
|
|
"time"
|
|
)
|
|
|
|
// Termination is how a run ended. Exactly one per run, always.
|
|
//
|
|
// An enum rather than a boolean and a message, because "what happened" is the
|
|
// first question asked of every trajectory — in a debugger, in an eval report,
|
|
// and in a support conversation — and a free-text reason cannot be grouped,
|
|
// counted or asserted on.
|
|
type Termination string
|
|
|
|
const (
|
|
TerminationCompleted Termination = "Completed"
|
|
TerminationBudgetExceeded Termination = "BudgetExceeded"
|
|
TerminationDeadline Termination = "Deadline"
|
|
TerminationConfirmationPending Termination = "ConfirmationPending"
|
|
TerminationToolFailure Termination = "ToolFailure"
|
|
TerminationRefused Termination = "Refused"
|
|
|
|
// GatewayFailure is the model provider failing to answer at all: rate
|
|
// limited, rejected the request, refused the credential, or unreachable.
|
|
// Added 2026-09-22 because until then every one of those was recorded as
|
|
// ToolFailure, and 131 of 318 production runs read as "a tool is broken"
|
|
// when no tool had failed — 45 of them were Groq's free-tier rate limit,
|
|
// which is a capacity decision, not a bug. The two are different questions
|
|
// to an operator ("what did we break" versus "what are we not paying
|
|
// for"), and an enum that could not tell them apart hid the answer for
|
|
// two weeks. Refused and Deadline keep their own reasons; this is the
|
|
// rest of the gateway's vocabulary.
|
|
TerminationGatewayFailure Termination = "GatewayFailure"
|
|
)
|
|
|
|
// Valid reports whether t is one of the seven.
|
|
func (t Termination) Valid() bool {
|
|
switch t {
|
|
case TerminationCompleted, TerminationBudgetExceeded, TerminationDeadline,
|
|
TerminationConfirmationPending, TerminationToolFailure, TerminationRefused,
|
|
TerminationGatewayFailure:
|
|
return true
|
|
}
|
|
return false
|
|
}
|
|
|
|
// Limits are the four bounds every run carries.
|
|
//
|
|
// I3: there is no "run until done" path. A run that reaches any of these ends
|
|
// with a structured result, never an exception into user-facing text.
|
|
type Limits struct {
|
|
MaxSteps int
|
|
MaxToolCalls int
|
|
MaxTokens int64
|
|
Deadline time.Duration
|
|
}
|
|
|
|
// LimitsForTier is what a run gets when its spec declares no limits of its own.
|
|
//
|
|
// Derived from the reasoning tier because that is the only thing a Krow agent
|
|
// definition says today about how much work it is worth. A `limits:` block in
|
|
// the frontmatter would override this per agent; adding one changes the spec
|
|
// contract and the authoring UI together, so it is a deliberate schema
|
|
// decision rather than something to infer here.
|
|
//
|
|
// The numbers are chosen so that the cheapest tier cannot quietly become the
|
|
// expensive one: a fast run gets a third of a deep run's steps and a sixth of
|
|
// its deadline, so a misrouted spec shows up as a truncated answer rather than
|
|
// as a bill.
|
|
func LimitsForTier(tier string) Limits {
|
|
switch tier {
|
|
case "fast":
|
|
return Limits{MaxSteps: 3, MaxToolCalls: 4, MaxTokens: 40_000, Deadline: 20 * time.Second}
|
|
case "deep":
|
|
return Limits{MaxSteps: 12, MaxToolCalls: 20, MaxTokens: 300_000, Deadline: 120 * time.Second}
|
|
default: // balanced, and anything unrecognised — ParseTier has already normalised it
|
|
return Limits{MaxSteps: 8, MaxToolCalls: 12, MaxTokens: 120_000, Deadline: 60 * time.Second}
|
|
}
|
|
}
|
|
|
|
// Snapshot is what a budget had left at one moment. Recorded into the
|
|
// trajectory before every dispatch, so "where did the budget go" is answerable
|
|
// after the fact instead of being reconstructed from timings.
|
|
type Snapshot struct {
|
|
StepsUsed int `json:"stepsUsed"`
|
|
StepsLeft int `json:"stepsLeft"`
|
|
ToolCallsUsed int `json:"toolCallsUsed"`
|
|
ToolCallsLeft int `json:"toolCallsLeft"`
|
|
TokensUsed int64 `json:"tokensUsed"`
|
|
TokensLeft int64 `json:"tokensLeft"`
|
|
MillisLeft int64 `json:"millisLeft"`
|
|
}
|
|
|
|
// Budget tracks one run against its limits.
|
|
//
|
|
// **Everything is claimed before dispatch, never after.** A step is spent the
|
|
// moment the loop decides to take it, not when it returns — otherwise a call
|
|
// that hangs until the context dies has consumed nothing on the ledger, and a
|
|
// loop that retries it can go round forever while the budget reads full.
|
|
//
|
|
// Tokens are the exception that proves the rule: their true cost is only known
|
|
// once a response comes back, so the budget is checked before dispatch and
|
|
// charged after. That leaves one turn of overshoot, bounded by MaxOutputTokens
|
|
// on the request, which is why the model gateway takes a hard per-call ceiling
|
|
// as well as this soft per-run one.
|
|
//
|
|
// Safe for concurrent use: subagents share their parent's budget, and two
|
|
// delegated branches must not both see the last step as available.
|
|
type Budget struct {
|
|
limits Limits
|
|
start time.Time
|
|
|
|
mu sync.Mutex
|
|
steps int
|
|
toolCalls int
|
|
tokens int64
|
|
}
|
|
|
|
// NewBudget starts a budget. The wall clock starts now: a run's deadline is
|
|
// measured from when it began, not from when it first reached a model.
|
|
func NewBudget(limits Limits) *Budget {
|
|
return &Budget{limits: limits, start: time.Now()}
|
|
}
|
|
|
|
// Limits returns the bounds this budget enforces.
|
|
func (b *Budget) Limits() Limits { return b.limits }
|
|
|
|
// ClaimStep takes one step up front, reporting the termination to end with if
|
|
// there was nothing left to take.
|
|
//
|
|
// The returned Termination is empty when the claim succeeded. Callers branch on
|
|
// that rather than on a boolean, so the reason a run stopped travels with the
|
|
// refusal instead of being re-derived at the call site.
|
|
func (b *Budget) ClaimStep() Termination {
|
|
if t := b.expired(); t != "" {
|
|
return t
|
|
}
|
|
b.mu.Lock()
|
|
defer b.mu.Unlock()
|
|
if b.steps >= b.limits.MaxSteps {
|
|
return TerminationBudgetExceeded
|
|
}
|
|
b.steps++
|
|
return ""
|
|
}
|
|
|
|
// ClaimToolCall takes one tool call up front.
|
|
func (b *Budget) ClaimToolCall() Termination {
|
|
if t := b.expired(); t != "" {
|
|
return t
|
|
}
|
|
b.mu.Lock()
|
|
defer b.mu.Unlock()
|
|
if b.toolCalls >= b.limits.MaxToolCalls {
|
|
return TerminationBudgetExceeded
|
|
}
|
|
b.toolCalls++
|
|
return ""
|
|
}
|
|
|
|
// CheckTokens reports whether there is token budget left to dispatch against.
|
|
//
|
|
// Checked before, charged after — see the type comment. A run that has already
|
|
// spent its allowance stops here rather than issuing one more call it cannot
|
|
// pay for.
|
|
func (b *Budget) CheckTokens() Termination {
|
|
if t := b.expired(); t != "" {
|
|
return t
|
|
}
|
|
b.mu.Lock()
|
|
defer b.mu.Unlock()
|
|
if b.tokens >= b.limits.MaxTokens {
|
|
return TerminationBudgetExceeded
|
|
}
|
|
return ""
|
|
}
|
|
|
|
// ChargeTokens records what a completed call actually cost.
|
|
//
|
|
// Called for refused and failed calls too. A turn that produced no text was
|
|
// still billed, and a ledger that forgives it is a ledger a loop will happily
|
|
// repeat against.
|
|
func (b *Budget) ChargeTokens(n int64) {
|
|
if n <= 0 {
|
|
return
|
|
}
|
|
b.mu.Lock()
|
|
defer b.mu.Unlock()
|
|
b.tokens += n
|
|
}
|
|
|
|
// expired reports the deadline having passed. Separate from the step and tool
|
|
// checks because it is a different termination reason: a run that ran out of
|
|
// time did not run out of budget, and conflating them hides which bound is
|
|
// actually being hit in production.
|
|
func (b *Budget) expired() Termination {
|
|
if time.Since(b.start) >= b.limits.Deadline {
|
|
return TerminationDeadline
|
|
}
|
|
return ""
|
|
}
|
|
|
|
// Context returns a context that is cancelled at the run's deadline.
|
|
//
|
|
// The same deadline the budget enforces, so an in-flight model call is torn
|
|
// down rather than being allowed to return into a run that has already ended.
|
|
func (b *Budget) Context(parent context.Context) (context.Context, context.CancelFunc) {
|
|
return context.WithDeadline(parent, b.start.Add(b.limits.Deadline))
|
|
}
|
|
|
|
// Snapshot reads the budget without changing it.
|
|
func (b *Budget) Snapshot() Snapshot {
|
|
b.mu.Lock()
|
|
defer b.mu.Unlock()
|
|
|
|
left := b.limits.Deadline - time.Since(b.start)
|
|
if left < 0 {
|
|
left = 0
|
|
}
|
|
return Snapshot{
|
|
StepsUsed: b.steps,
|
|
StepsLeft: max(0, b.limits.MaxSteps-b.steps),
|
|
ToolCallsUsed: b.toolCalls,
|
|
ToolCallsLeft: max(0, b.limits.MaxToolCalls-b.toolCalls),
|
|
TokensUsed: b.tokens,
|
|
TokensLeft: maxInt64(0, b.limits.MaxTokens-b.tokens),
|
|
MillisLeft: left.Milliseconds(),
|
|
}
|
|
}
|
|
|
|
func (s Snapshot) String() string {
|
|
return fmt.Sprintf("steps %d/%d, tools %d/%d, tokens %d, %dms left",
|
|
s.StepsUsed, s.StepsUsed+s.StepsLeft,
|
|
s.ToolCallsUsed, s.ToolCallsUsed+s.ToolCallsLeft,
|
|
s.TokensUsed, s.MillisLeft)
|
|
}
|
|
|
|
func maxInt64(a, b int64) int64 {
|
|
if a > b {
|
|
return a
|
|
}
|
|
return b
|
|
}
|