Run the turn again on another provider when a rate limit lands mid-run
In-place failover covered none of the failures this deployment actually had. gateway.canFailOver will not move a conversation that has called a tool — the assistant turn echoing that call belongs to the provider that issued it, and a vendor which signs its function calls rejects a follow-up carrying somebody else's. But a rate limit lands where the request is BIGGEST, which is the second or third model call, once the catalogue, the retrieved block, the tool results and the whole prior conversation are being re-sent. Every GatewayFailure in agent_runs had already called a tool. The error text says exactly what it was: http 429: Rate limit reached for model `openai/gpt-oss-120b` … on tokens per minute (TPM): Limit 8000, Used 7183 So the loop starts the turn over on the next provider. No transcript is sent, so nothing provider-specific travels and the signature problem cannot arise: the question is simply asked again somewhere with budget left. It costs the work already done, charged to the budget that is not exhausted. gateway.Standby is the whole of what the runtime is told — "there is another one, here it is". No vendor, credential or model id crosses the boundary, and the loop still cannot name a provider. THE RULE THAT MAKES IT SAFE: a run carrying a confirmation never restarts. Re-running re-runs its tools; a read twice is two reads, a write twice is two shifts assigned. I4 makes the test cheap — a write executes only against a resolved token (Registry.gate), so a run with no confirmation cannot have written anything, and one with a confirmation is refused without inspecting what it did. Once, not until the providers run out: a question worth asking twice is not worth asking five times, and each attempt spends a real budget. A terminal error — a rejected credential, a model this deployment cannot use — is not retried anywhere, on the same line canFailOver already draws. Five tests, including both refusals. Verified with teeth: disabling the restart fails the rate-limit case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -127,7 +127,67 @@ func (m *ModelExecutor) ExecuteAgent(ctx context.Context, agent *Agent, input Ex
|
||||
func (m *ModelExecutor) executeWithLimits(
|
||||
ctx context.Context, agent *Agent, input ExecutionInput, limits Limits,
|
||||
) (*ExecutionResult, error) {
|
||||
return m.executeRun(ctx, agent, input, limits, delegation{})
|
||||
res, err := m.executeRun(ctx, agent, input, limits, delegation{})
|
||||
if next, ok := m.standbyFor(res, input); ok {
|
||||
return next.executeRun(ctx, agent, input, limits, delegation{})
|
||||
}
|
||||
return res, err
|
||||
}
|
||||
|
||||
// standbyFor decides whether to run the whole turn again on another provider.
|
||||
//
|
||||
// WHY A RESTART AND NOT A HANDOVER. gateway.canFailOver will not move a
|
||||
// conversation that has called a tool: the assistant turn echoing that call is
|
||||
// the provider's own, and a vendor that signs its function calls rejects a
|
||||
// follow-up carrying somebody else's. But a rate limit lands where the request
|
||||
// is BIGGEST, which is the second or third call, once the catalogue, the
|
||||
// retrieved block, the tool results and the whole prior conversation are being
|
||||
// re-sent. Measured on this deployment: every GatewayFailure recorded had
|
||||
// already called a tool, so in-place failover covered none of them —
|
||||
//
|
||||
// http 429 … Rate limit reached … tokens per minute (TPM): Limit 8000, Used 7183
|
||||
//
|
||||
// Starting over sends no transcript, so nothing provider-specific travels and
|
||||
// the question is simply asked again somewhere with budget left. It costs the
|
||||
// work already done, charged to the budget that is NOT exhausted.
|
||||
//
|
||||
// THE RULE THAT MAKES IT SAFE: a run that carried a confirmation never
|
||||
// restarts. Re-running re-runs its tools, and a read twice is two reads while a
|
||||
// write twice is two shifts assigned. I4 is what makes the test this cheap —
|
||||
// a write executes ONLY against a resolved token (Registry.gate), so a run with
|
||||
// no confirmation cannot have written anything, and one with a confirmation is
|
||||
// refused here without inspecting what it did.
|
||||
//
|
||||
// Once. Not a loop over every provider: a question worth asking twice is not
|
||||
// worth asking five times, and each attempt spends a real budget. The second
|
||||
// result is returned as it stands, whatever it says.
|
||||
func (m *ModelExecutor) standbyFor(res *ExecutionResult, input ExecutionInput) (*ModelExecutor, bool) {
|
||||
if res == nil || res.Termination != TerminationGatewayFailure {
|
||||
return nil, false
|
||||
}
|
||||
// An approved write may already have happened. Nothing below is worth a
|
||||
// double assignment.
|
||||
if input.Confirmation != "" {
|
||||
return nil, false
|
||||
}
|
||||
// Only a transient fault moves, on the same line gateway.canFailOver draws:
|
||||
// a rejected credential or a model this deployment cannot use fails the
|
||||
// same way everywhere, and asking twice only doubles the bill.
|
||||
var gwErr *gateway.Error
|
||||
if !errors.As(res.Error, &gwErr) || !gwErr.Retryable() {
|
||||
return nil, false
|
||||
}
|
||||
sb, ok := m.gw.(gateway.Standby)
|
||||
if !ok {
|
||||
return nil, false
|
||||
}
|
||||
next, ok := sb.Standby()
|
||||
if !ok {
|
||||
return nil, false
|
||||
}
|
||||
clone := *m
|
||||
clone.gw = next
|
||||
return &clone, true
|
||||
}
|
||||
|
||||
// executeRun is the loop. `del` is what a SUBAGENT inherits from its parent —
|
||||
|
||||
Reference in New Issue
Block a user