Ask a second provider when the first one is busy
A free tier's ceiling is tokens per MINUTE, and one run can exceed a whole
minute's worth by itself: a three-call run measured 12,123 against a ceiling of
8,000. withRetry already fires three times and all three are refused, because
1.6 seconds of backoff does not buy back a minute's budget. The run ends
GatewayFailure and somebody reads "the model did not answer".
Retrying harder cannot fix a ceiling. Asking somebody else can — the ceilings
are per provider, so a second key is a second budget. Groq, Cerebras, Gemini,
Mistral and OpenRouter all serve the same chat-completions shape, which is why
this is a list of Configs and not a second implementation.
Configured as MODEL_FALLBACK_<n>_BASE_URL / _API_KEY / _FAST / _BALANCED /
_DEEP, numbered because five fields times three providers packed into one
delimited string is a parser nobody can read under pressure. Empty is the
ordinary case and returns the primary unwrapped, so a single-provider
deployment carries no wrapper and behaves exactly as before.
Failover is NOT unconditional, and the two guards are the design:
- Only a transient failure moves. Error.Retryable() already draws that line
for retries and it is the same line here. A 401 is this deployment's own
credential and a 400 is a malformed request; both fail identically at every
vendor, so trying three turns one visible fault into three invisible ones.
- Only an unpinned conversation moves. ToolCall.Extra carries provider
metadata echoed back verbatim — Gemini 3's thought signature — and a vendor
rejects a follow-up that drops its own. A conversation carrying any belongs
to whoever started it, so failover is available on the first model call,
which is where a rate limit usually lands anyway.
Streaming falls over only before the first fragment: once text is in the
reader's window, a second provider would continue that sentence in a different
voice.
What this does not do, since the gap is where the next bug lives: it does not
make a run cheaper, does not raise any one ceiling, and does not help when every
provider is exhausted at once. It turns one busy provider into a slower answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -177,6 +177,17 @@ type ModelConfig struct {
|
||||
// OpenAI-compatible wire. Off by default: reasoning models accept the
|
||||
// field and most others reject the entire request rather than ignoring it.
|
||||
ReasoningEffort bool
|
||||
|
||||
// Fallbacks are further providers to ask when the one above cannot answer,
|
||||
// in order. Empty is the ordinary case and carries no wrapper at all.
|
||||
//
|
||||
// A FREE TIER'S CEILING IS PER PROVIDER, so a second key is a second
|
||||
// budget — which is the only thing that helps when a single run costs more
|
||||
// tokens than a provider allows in a minute. Each entry is a whole
|
||||
// ModelConfig because a fallback is a different service with its own
|
||||
// credential, its own base URL and its own model ids; sharing any of those
|
||||
// is what makes "the same request, somewhere else" impossible.
|
||||
Fallbacks []ModelConfig
|
||||
}
|
||||
|
||||
// SeedConfig locates the demo fixture. The file is generated from the frontend
|
||||
@@ -418,6 +429,7 @@ func Load() (*Config, error) {
|
||||
// unstreamed call, not the run's budget.
|
||||
MaxOutputTokens: intDefault("MODEL_MAX_OUTPUT_TOKENS", 16000),
|
||||
ReasoningEffort: boolDefault("MODEL_REASONING_EFFORT", false),
|
||||
Fallbacks: loadFallbacks(),
|
||||
},
|
||||
DB: DBConfig{
|
||||
Host: required("DATABASE_HOST"),
|
||||
@@ -943,3 +955,42 @@ func (c *Config) validateOAuth() error {
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
// loadFallbacks reads MODEL_FALLBACK_<n>_* for n = 1, 2, 3…
|
||||
//
|
||||
// Numbered rather than comma-separated because each provider needs five fields,
|
||||
// and a delimiter-packed string holding five fields times three providers is a
|
||||
// parser nobody can read and an operator cannot edit under pressure:
|
||||
//
|
||||
// MODEL_FALLBACK_1_BASE_URL=https://api.cerebras.ai/v1
|
||||
// MODEL_FALLBACK_1_API_KEY=…
|
||||
// MODEL_FALLBACK_1_BALANCED=<a model that endpoint serves>
|
||||
//
|
||||
// Stops at the first gap, so a deployment cannot half-configure a third
|
||||
// provider by deleting the second and have the third silently promoted.
|
||||
//
|
||||
// A fallback with no BASE_URL or no API_KEY is not a fallback, so both are
|
||||
// required and the entry is skipped without one. The model ids fall back to the
|
||||
// PRIMARY's — wrong for a different vendor, which is why each should be set,
|
||||
// but an unset id produces a visible invalid_request rather than silence.
|
||||
func loadFallbacks() []ModelConfig {
|
||||
var out []ModelConfig
|
||||
for n := 1; ; n++ {
|
||||
prefix := fmt.Sprintf("MODEL_FALLBACK_%d_", n)
|
||||
base := strings.TrimSpace(os.Getenv(prefix + "BASE_URL"))
|
||||
key := strings.TrimSpace(os.Getenv(prefix + "API_KEY"))
|
||||
if base == "" || key == "" {
|
||||
return out
|
||||
}
|
||||
out = append(out, ModelConfig{
|
||||
Provider: strings.ToLower(strings.TrimSpace(os.Getenv(prefix + "PROVIDER"))),
|
||||
APIKey: key,
|
||||
BaseURL: base,
|
||||
Fast: strings.TrimSpace(os.Getenv(prefix + "FAST")),
|
||||
Balanced: strings.TrimSpace(os.Getenv(prefix + "BALANCED")),
|
||||
Deep: strings.TrimSpace(os.Getenv(prefix + "DEEP")),
|
||||
MaxOutputTokens: intDefault(prefix+"MAX_OUTPUT_TOKENS", 16000),
|
||||
ReasoningEffort: boolDefault(prefix+"REASONING_EFFORT", false),
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user