In-place failover covered none of the failures this deployment actually had.
gateway.canFailOver will not move a conversation that has called a tool — the
assistant turn echoing that call belongs to the provider that issued it, and a
vendor which signs its function calls rejects a follow-up carrying somebody
else's. But a rate limit lands where the request is BIGGEST, which is the
second or third model call, once the catalogue, the retrieved block, the tool
results and the whole prior conversation are being re-sent.
Every GatewayFailure in agent_runs had already called a tool. The error text
says exactly what it was:
http 429: Rate limit reached for model `openai/gpt-oss-120b` …
on tokens per minute (TPM): Limit 8000, Used 7183
So the loop starts the turn over on the next provider. No transcript is sent,
so nothing provider-specific travels and the signature problem cannot arise:
the question is simply asked again somewhere with budget left. It costs the
work already done, charged to the budget that is not exhausted.
gateway.Standby is the whole of what the runtime is told — "there is another
one, here it is". No vendor, credential or model id crosses the boundary, and
the loop still cannot name a provider.
THE RULE THAT MAKES IT SAFE: a run carrying a confirmation never restarts.
Re-running re-runs its tools; a read twice is two reads, a write twice is two
shifts assigned. I4 makes the test cheap — a write executes only against a
resolved token (Registry.gate), so a run with no confirmation cannot have
written anything, and one with a confirmation is refused without inspecting
what it did.
Once, not until the providers run out: a question worth asking twice is not
worth asking five times, and each attempt spends a real budget. A terminal
error — a rejected credential, a model this deployment cannot use — is not
retried anywhere, on the same line canFailOver already draws.
Five tests, including both refusals. Verified with teeth: disabling the restart
fails the rate-limit case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
ace4db8 refused to move a conversation whose ToolCall carried Extra — provider
metadata echoed back verbatim, Gemini 3's thought signature being the case it
was written for. That test is wrong, and in the exact direction that breaks
production.
Extra is populated by the provider that ISSUED the call. A conversation begun
on Groq carries none at all, so it read as movable; moving it hands Gemini an
assistant turn holding a function call with no thought signature, which is the
400 that took the cluster down on 2026-09-22. An absent field meant "came from
somewhere that does not sign", and it was read as "safe to move".
So the test is the tool call, not the metadata: any ToolCalls or ToolResults in
the conversation pin it to whoever has been answering. Failover stays available
on the first model call of a run, which is where a rate limit lands anyway.
Found while configuring Groq primary with Gemini as the fallback — the exact
pairing that triggers it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
A free tier's ceiling is tokens per MINUTE, and one run can exceed a whole
minute's worth by itself: a three-call run measured 12,123 against a ceiling of
8,000. withRetry already fires three times and all three are refused, because
1.6 seconds of backoff does not buy back a minute's budget. The run ends
GatewayFailure and somebody reads "the model did not answer".
Retrying harder cannot fix a ceiling. Asking somebody else can — the ceilings
are per provider, so a second key is a second budget. Groq, Cerebras, Gemini,
Mistral and OpenRouter all serve the same chat-completions shape, which is why
this is a list of Configs and not a second implementation.
Configured as MODEL_FALLBACK_<n>_BASE_URL / _API_KEY / _FAST / _BALANCED /
_DEEP, numbered because five fields times three providers packed into one
delimited string is a parser nobody can read under pressure. Empty is the
ordinary case and returns the primary unwrapped, so a single-provider
deployment carries no wrapper and behaves exactly as before.
Failover is NOT unconditional, and the two guards are the design:
- Only a transient failure moves. Error.Retryable() already draws that line
for retries and it is the same line here. A 401 is this deployment's own
credential and a 400 is a malformed request; both fail identically at every
vendor, so trying three turns one visible fault into three invisible ones.
- Only an unpinned conversation moves. ToolCall.Extra carries provider
metadata echoed back verbatim — Gemini 3's thought signature — and a vendor
rejects a follow-up that drops its own. A conversation carrying any belongs
to whoever started it, so failover is available on the first model call,
which is where a rate limit usually lands anyway.
Streaming falls over only before the first fragment: once text is in the
reader's window, a second provider would continue that sentence in a different
voice.
What this does not do, since the gap is where the next bug lives: it does not
make a run cheaper, does not raise any one ceiling, and does not help when every
provider is exhausted at once. It turns one busy provider into a slower answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>