ace4db8 refused to move a conversation whose ToolCall carried Extra — provider
metadata echoed back verbatim, Gemini 3's thought signature being the case it
was written for. That test is wrong, and in the exact direction that breaks
production.
Extra is populated by the provider that ISSUED the call. A conversation begun
on Groq carries none at all, so it read as movable; moving it hands Gemini an
assistant turn holding a function call with no thought signature, which is the
400 that took the cluster down on 2026-09-22. An absent field meant "came from
somewhere that does not sign", and it was read as "safe to move".
So the test is the tool call, not the metadata: any ToolCalls or ToolResults in
the conversation pin it to whoever has been answering. Failover stays available
on the first model call of a run, which is where a rate limit lands anyway.
Found while configuring Groq primary with Gemini as the fallback — the exact
pairing that triggers it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
133 lines
4.9 KiB
Go
133 lines
4.9 KiB
Go
package gateway
|
|
|
|
// Failover: a second and third provider, for when the first one says no.
|
|
//
|
|
// THE PROBLEM THIS SOLVES IS A CEILING, NOT A BUG. A free tier is a token
|
|
// budget per minute, and one agent run can exceed a whole minute's worth by
|
|
// itself — a three-call run measured 12,123 tokens against a ceiling of 8,000.
|
|
// withRetry already fires three times, and on a rate limit all three are
|
|
// refused, because waiting 1.6 seconds does not buy back a minute's budget. The
|
|
// run then ends GatewayFailure and a person reads "the model did not answer".
|
|
//
|
|
// Retrying harder cannot fix that. Asking somebody else can: the ceilings are
|
|
// per provider, so a second key is a second budget. Groq, Cerebras, Gemini,
|
|
// Mistral and OpenRouter all serve the same chat-completions shape, which is
|
|
// the whole reason this is a list of Configs and not a second implementation.
|
|
//
|
|
// WHAT IT DOES NOT DO, stated because the gap is where the next bug lives:
|
|
// it does not make a run cheaper, it does not raise any one provider's ceiling,
|
|
// and it does not help when every configured provider is exhausted at once. It
|
|
// converts "one busy provider" from an outage into a slower answer.
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
)
|
|
|
|
// failover tries each provider in order until one answers.
|
|
type failover struct {
|
|
providers []Gateway
|
|
}
|
|
|
|
// NewFailover builds a gateway that falls back through `rest` when `primary`
|
|
// cannot answer. With no fallbacks it returns the primary unchanged, so a
|
|
// single-provider deployment carries no wrapper and behaves exactly as before.
|
|
func NewFailover(primary Gateway, rest ...Gateway) Gateway {
|
|
if len(rest) == 0 {
|
|
return primary
|
|
}
|
|
return &failover{providers: append([]Gateway{primary}, rest...)}
|
|
}
|
|
|
|
func (f *failover) Complete(ctx context.Context, req Request) (*Response, error) {
|
|
var last error
|
|
for i, p := range f.providers {
|
|
if i > 0 && !canFailOver(req, last) {
|
|
break
|
|
}
|
|
resp, err := p.Complete(ctx, req)
|
|
if err == nil {
|
|
return resp, nil
|
|
}
|
|
last = err
|
|
// The caller's deadline governs. A deployment with four providers must
|
|
// not spend four timeouts' worth of a person's patience discovering
|
|
// that none of them is available.
|
|
if ctx.Err() != nil {
|
|
break
|
|
}
|
|
}
|
|
return nil, last
|
|
}
|
|
|
|
// Stream falls over only before the first fragment has been delivered.
|
|
//
|
|
// After a delta reaches the client, the answer has begun in the reader's own
|
|
// window. Starting a second provider would continue that sentence in a
|
|
// different voice from a different model, or repeat its opening — so once text
|
|
// is out, the error is the answer.
|
|
func (f *failover) Stream(ctx context.Context, req Request, onDelta func(string)) (*Response, error) {
|
|
var last error
|
|
for i, p := range f.providers {
|
|
if i > 0 && !canFailOver(req, last) {
|
|
break
|
|
}
|
|
var delivered bool
|
|
wrapped := func(s string) {
|
|
delivered = true
|
|
onDelta(s)
|
|
}
|
|
resp, err := StreamComplete(ctx, p, req, wrapped)
|
|
if err == nil {
|
|
return resp, nil
|
|
}
|
|
last = err
|
|
if delivered || ctx.Err() != nil {
|
|
break
|
|
}
|
|
}
|
|
return nil, last
|
|
}
|
|
|
|
// canFailOver decides whether asking a DIFFERENT provider is sound.
|
|
//
|
|
// Two conditions, and both are necessary.
|
|
//
|
|
// 1. THE FAILURE MUST BE TRANSIENT. Error.Retryable() already draws that line
|
|
// for retries and it is the same line here: a rate limit or a 5xx is the
|
|
// provider being unable, and somebody else may be able. A 400 is a
|
|
// malformed request and will be malformed for everyone; a 401 is this
|
|
// deployment's own credential. Failing over on those turns one provider's
|
|
// configuration error into every provider's, and buries the fault.
|
|
//
|
|
// 2. THE CONVERSATION MUST CARRY NO TOOL CALL AT ALL. Not merely "no
|
|
// provider metadata" — ANY tool call pins the conversation, and the
|
|
// difference is a bug this got wrong first time round.
|
|
//
|
|
// The reasoning that failed: ToolCall.Extra carries provider metadata
|
|
// echoed back verbatim (Gemini 3's thought signature), so it looked
|
|
// sufficient to refuse only when Extra was present. But Extra is populated
|
|
// by the provider that ISSUED the call. A conversation begun on Groq
|
|
// carries no Extra at all, so it looked movable — and moving it hands
|
|
// Gemini an assistant turn containing a function call with no thought
|
|
// signature, which is exactly the 400 that took production down on
|
|
// 2026-09-22. The absent field was read as "safe to move" when it meant
|
|
// "came from somewhere that does not sign".
|
|
//
|
|
// So the test is the tool call, not the metadata. A conversation that has
|
|
// called a tool belongs to whoever has been answering it. Failover is
|
|
// available on the first model call of a run, which is where a rate limit
|
|
// lands anyway, and nowhere else.
|
|
func canFailOver(req Request, err error) bool {
|
|
var gwErr *Error
|
|
if !errors.As(err, &gwErr) || !gwErr.Retryable() {
|
|
return false
|
|
}
|
|
for _, m := range req.Messages {
|
|
if len(m.ToolCalls) > 0 || len(m.ToolResults) > 0 {
|
|
return false
|
|
}
|
|
}
|
|
return true
|
|
}
|