Files
Suriya 5166fde764
Some checks failed
CI / test (push) Failing after 4m37s
CI / fixture (push) Failing after 8s
Add GatewayFailure: the provider not answering is not a tool failing
terminationFor sent every gateway error that was not Refused or Timeout
to ToolFailure, because the enum had nowhere else to put it. On
2026-09-22 that was 131 of 318 production runs, and not one of them was
a tool failing: 56 were retired model ids, 25 an exhausted Anthropic
balance, 45 Groq's free-tier rate limit -- the only one still happening.
An operator reading the termination column saw "a tool is broken" for
two weeks while the actual answer was "we are not paying for capacity".

GatewayFailure is the seventh termination. Rate limited, request
rejected, credential refused and unreachable land there; Refused and
Deadline keep their own reasons; a non-gateway error is still the tool
layer's. A delegation whose subagent died at the gateway now carries
that reason up to the parent instead of reading as a tool call that
failed.

Migration 000016 widens the CHECK that 000006 chose precisely so this
would be a migration rather than an ALTER TYPE. Its down folds any
GatewayFailure rows back to ToolFailure BEFORE narrowing the constraint,
which is the order that works; verified up, down and up again on a
scratch database. Existing rows are left as they are -- the trajectory
entries still carry the gateway.* code for anyone reclassifying history.

The surface wording is the one termination where "try again" is honest
advice, since the dominant cause clears within a minute.

Full suite run against a real database, including the tests that skip
without one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-22 12:48:09 +05:30

247 lines
8.3 KiB
Go

package runtime
import (
"context"
"fmt"
"sync"
"time"
)
// Termination is how a run ended. Exactly one per run, always.
//
// An enum rather than a boolean and a message, because "what happened" is the
// first question asked of every trajectory — in a debugger, in an eval report,
// and in a support conversation — and a free-text reason cannot be grouped,
// counted or asserted on.
type Termination string
const (
TerminationCompleted Termination = "Completed"
TerminationBudgetExceeded Termination = "BudgetExceeded"
TerminationDeadline Termination = "Deadline"
TerminationConfirmationPending Termination = "ConfirmationPending"
TerminationToolFailure Termination = "ToolFailure"
TerminationRefused Termination = "Refused"
// GatewayFailure is the model provider failing to answer at all: rate
// limited, rejected the request, refused the credential, or unreachable.
// Added 2026-09-22 because until then every one of those was recorded as
// ToolFailure, and 131 of 318 production runs read as "a tool is broken"
// when no tool had failed — 45 of them were Groq's free-tier rate limit,
// which is a capacity decision, not a bug. The two are different questions
// to an operator ("what did we break" versus "what are we not paying
// for"), and an enum that could not tell them apart hid the answer for
// two weeks. Refused and Deadline keep their own reasons; this is the
// rest of the gateway's vocabulary.
TerminationGatewayFailure Termination = "GatewayFailure"
)
// Valid reports whether t is one of the seven.
func (t Termination) Valid() bool {
switch t {
case TerminationCompleted, TerminationBudgetExceeded, TerminationDeadline,
TerminationConfirmationPending, TerminationToolFailure, TerminationRefused,
TerminationGatewayFailure:
return true
}
return false
}
// Limits are the four bounds every run carries.
//
// I3: there is no "run until done" path. A run that reaches any of these ends
// with a structured result, never an exception into user-facing text.
type Limits struct {
MaxSteps int
MaxToolCalls int
MaxTokens int64
Deadline time.Duration
}
// LimitsForTier is what a run gets when its spec declares no limits of its own.
//
// Derived from the reasoning tier because that is the only thing a Krow agent
// definition says today about how much work it is worth. A `limits:` block in
// the frontmatter would override this per agent; adding one changes the spec
// contract and the authoring UI together, so it is a deliberate schema
// decision rather than something to infer here.
//
// The numbers are chosen so that the cheapest tier cannot quietly become the
// expensive one: a fast run gets a third of a deep run's steps and a sixth of
// its deadline, so a misrouted spec shows up as a truncated answer rather than
// as a bill.
func LimitsForTier(tier string) Limits {
switch tier {
case "fast":
return Limits{MaxSteps: 3, MaxToolCalls: 4, MaxTokens: 40_000, Deadline: 20 * time.Second}
case "deep":
return Limits{MaxSteps: 12, MaxToolCalls: 20, MaxTokens: 300_000, Deadline: 120 * time.Second}
default: // balanced, and anything unrecognised — ParseTier has already normalised it
return Limits{MaxSteps: 8, MaxToolCalls: 12, MaxTokens: 120_000, Deadline: 60 * time.Second}
}
}
// Snapshot is what a budget had left at one moment. Recorded into the
// trajectory before every dispatch, so "where did the budget go" is answerable
// after the fact instead of being reconstructed from timings.
type Snapshot struct {
StepsUsed int `json:"stepsUsed"`
StepsLeft int `json:"stepsLeft"`
ToolCallsUsed int `json:"toolCallsUsed"`
ToolCallsLeft int `json:"toolCallsLeft"`
TokensUsed int64 `json:"tokensUsed"`
TokensLeft int64 `json:"tokensLeft"`
MillisLeft int64 `json:"millisLeft"`
}
// Budget tracks one run against its limits.
//
// **Everything is claimed before dispatch, never after.** A step is spent the
// moment the loop decides to take it, not when it returns — otherwise a call
// that hangs until the context dies has consumed nothing on the ledger, and a
// loop that retries it can go round forever while the budget reads full.
//
// Tokens are the exception that proves the rule: their true cost is only known
// once a response comes back, so the budget is checked before dispatch and
// charged after. That leaves one turn of overshoot, bounded by MaxOutputTokens
// on the request, which is why the model gateway takes a hard per-call ceiling
// as well as this soft per-run one.
//
// Safe for concurrent use: subagents share their parent's budget, and two
// delegated branches must not both see the last step as available.
type Budget struct {
limits Limits
start time.Time
mu sync.Mutex
steps int
toolCalls int
tokens int64
}
// NewBudget starts a budget. The wall clock starts now: a run's deadline is
// measured from when it began, not from when it first reached a model.
func NewBudget(limits Limits) *Budget {
return &Budget{limits: limits, start: time.Now()}
}
// Limits returns the bounds this budget enforces.
func (b *Budget) Limits() Limits { return b.limits }
// ClaimStep takes one step up front, reporting the termination to end with if
// there was nothing left to take.
//
// The returned Termination is empty when the claim succeeded. Callers branch on
// that rather than on a boolean, so the reason a run stopped travels with the
// refusal instead of being re-derived at the call site.
func (b *Budget) ClaimStep() Termination {
if t := b.expired(); t != "" {
return t
}
b.mu.Lock()
defer b.mu.Unlock()
if b.steps >= b.limits.MaxSteps {
return TerminationBudgetExceeded
}
b.steps++
return ""
}
// ClaimToolCall takes one tool call up front.
func (b *Budget) ClaimToolCall() Termination {
if t := b.expired(); t != "" {
return t
}
b.mu.Lock()
defer b.mu.Unlock()
if b.toolCalls >= b.limits.MaxToolCalls {
return TerminationBudgetExceeded
}
b.toolCalls++
return ""
}
// CheckTokens reports whether there is token budget left to dispatch against.
//
// Checked before, charged after — see the type comment. A run that has already
// spent its allowance stops here rather than issuing one more call it cannot
// pay for.
func (b *Budget) CheckTokens() Termination {
if t := b.expired(); t != "" {
return t
}
b.mu.Lock()
defer b.mu.Unlock()
if b.tokens >= b.limits.MaxTokens {
return TerminationBudgetExceeded
}
return ""
}
// ChargeTokens records what a completed call actually cost.
//
// Called for refused and failed calls too. A turn that produced no text was
// still billed, and a ledger that forgives it is a ledger a loop will happily
// repeat against.
func (b *Budget) ChargeTokens(n int64) {
if n <= 0 {
return
}
b.mu.Lock()
defer b.mu.Unlock()
b.tokens += n
}
// expired reports the deadline having passed. Separate from the step and tool
// checks because it is a different termination reason: a run that ran out of
// time did not run out of budget, and conflating them hides which bound is
// actually being hit in production.
func (b *Budget) expired() Termination {
if time.Since(b.start) >= b.limits.Deadline {
return TerminationDeadline
}
return ""
}
// Context returns a context that is cancelled at the run's deadline.
//
// The same deadline the budget enforces, so an in-flight model call is torn
// down rather than being allowed to return into a run that has already ended.
func (b *Budget) Context(parent context.Context) (context.Context, context.CancelFunc) {
return context.WithDeadline(parent, b.start.Add(b.limits.Deadline))
}
// Snapshot reads the budget without changing it.
func (b *Budget) Snapshot() Snapshot {
b.mu.Lock()
defer b.mu.Unlock()
left := b.limits.Deadline - time.Since(b.start)
if left < 0 {
left = 0
}
return Snapshot{
StepsUsed: b.steps,
StepsLeft: max(0, b.limits.MaxSteps-b.steps),
ToolCallsUsed: b.toolCalls,
ToolCallsLeft: max(0, b.limits.MaxToolCalls-b.toolCalls),
TokensUsed: b.tokens,
TokensLeft: maxInt64(0, b.limits.MaxTokens-b.tokens),
MillisLeft: left.Milliseconds(),
}
}
func (s Snapshot) String() string {
return fmt.Sprintf("steps %d/%d, tools %d/%d, tokens %d, %dms left",
s.StepsUsed, s.StepsUsed+s.StepsLeft,
s.ToolCallsUsed, s.ToolCallsUsed+s.ToolCallsLeft,
s.TokensUsed, s.MillisLeft)
}
func maxInt64(a, b int64) int64 {
if a > b {
return a
}
return b
}