package gateway // Failover: a second and third provider, for when the first one says no. // // THE PROBLEM THIS SOLVES IS A CEILING, NOT A BUG. A free tier is a token // budget per minute, and one agent run can exceed a whole minute's worth by // itself — a three-call run measured 12,123 tokens against a ceiling of 8,000. // withRetry already fires three times, and on a rate limit all three are // refused, because waiting 1.6 seconds does not buy back a minute's budget. The // run then ends GatewayFailure and a person reads "the model did not answer". // // Retrying harder cannot fix that. Asking somebody else can: the ceilings are // per provider, so a second key is a second budget. Groq, Cerebras, Gemini, // Mistral and OpenRouter all serve the same chat-completions shape, which is // the whole reason this is a list of Configs and not a second implementation. // // WHAT IT DOES NOT DO, stated because the gap is where the next bug lives: // it does not make a run cheaper, it does not raise any one provider's ceiling, // and it does not help when every configured provider is exhausted at once. It // converts "one busy provider" from an outage into a slower answer. import ( "context" "errors" ) // failover tries each provider in order until one answers. type failover struct { providers []Gateway } // NewFailover builds a gateway that falls back through `rest` when `primary` // cannot answer. With no fallbacks it returns the primary unchanged, so a // single-provider deployment carries no wrapper and behaves exactly as before. func NewFailover(primary Gateway, rest ...Gateway) Gateway { if len(rest) == 0 { return primary } return &failover{providers: append([]Gateway{primary}, rest...)} } // Standby is a gateway that has somewhere else to go. // // The runtime needs this and must NOT learn what a provider is. A mid-run // failure cannot be moved by this package — the conversation is half built and // its tool calls belong to whoever issued them (see canFailOver) — so the only // thing that can rescue it is starting the run again somewhere else, and only // the loop can do that. This is the whole of what the loop is told: "there is // another one, here it is", with no vendor, credential or model id crossing the // boundary. type Standby interface { // Standby returns a gateway beginning at the NEXT provider, and whether // there was one. The receiver is unchanged. Standby() (Gateway, bool) } // Standby drops the provider that just failed and returns the rest. // // The remainder keeps its own fallbacks, so a second failure on a three // provider deployment still has somewhere to go. With one provider left there // is no wrapper at all, which is NewFailover's own rule. func (f *failover) Standby() (Gateway, bool) { if len(f.providers) < 2 { return nil, false } return NewFailover(f.providers[1], f.providers[2:]...), true } func (f *failover) Complete(ctx context.Context, req Request) (*Response, error) { var last error for i, p := range f.providers { if i > 0 && !canFailOver(req, last) { break } resp, err := p.Complete(ctx, req) if err == nil { return resp, nil } last = err // The caller's deadline governs. A deployment with four providers must // not spend four timeouts' worth of a person's patience discovering // that none of them is available. if ctx.Err() != nil { break } } return nil, last } // Stream falls over only before the first fragment has been delivered. // // After a delta reaches the client, the answer has begun in the reader's own // window. Starting a second provider would continue that sentence in a // different voice from a different model, or repeat its opening — so once text // is out, the error is the answer. func (f *failover) Stream(ctx context.Context, req Request, onDelta func(string)) (*Response, error) { var last error for i, p := range f.providers { if i > 0 && !canFailOver(req, last) { break } var delivered bool wrapped := func(s string) { delivered = true onDelta(s) } resp, err := StreamComplete(ctx, p, req, wrapped) if err == nil { return resp, nil } last = err if delivered || ctx.Err() != nil { break } } return nil, last } // canFailOver decides whether asking a DIFFERENT provider is sound. // // Two conditions, and both are necessary. // // 1. THE FAILURE MUST BE TRANSIENT. Error.Retryable() already draws that line // for retries and it is the same line here: a rate limit or a 5xx is the // provider being unable, and somebody else may be able. A 400 is a // malformed request and will be malformed for everyone; a 401 is this // deployment's own credential. Failing over on those turns one provider's // configuration error into every provider's, and buries the fault. // // 2. THE CONVERSATION MUST CARRY NO TOOL CALL AT ALL. Not merely "no // provider metadata" — ANY tool call pins the conversation, and the // difference is a bug this got wrong first time round. // // The reasoning that failed: ToolCall.Extra carries provider metadata // echoed back verbatim (Gemini 3's thought signature), so it looked // sufficient to refuse only when Extra was present. But Extra is populated // by the provider that ISSUED the call. A conversation begun on Groq // carries no Extra at all, so it looked movable — and moving it hands // Gemini an assistant turn containing a function call with no thought // signature, which is exactly the 400 that took production down on // 2026-09-22. The absent field was read as "safe to move" when it meant // "came from somewhere that does not sign". // // So the test is the tool call, not the metadata. A conversation that has // called a tool belongs to whoever has been answering it. Failover is // available on the first model call of a run, which is where a rate limit // lands anyway, and nowhere else. func canFailOver(req Request, err error) bool { var gwErr *Error if !errors.As(err, &gwErr) || !gwErr.Retryable() { return false } for _, m := range req.Messages { if len(m.ToolCalls) > 0 || len(m.ToolResults) > 0 { return false } } return true }