A free tier's ceiling is tokens per MINUTE, and one run can exceed a whole
minute's worth by itself: a three-call run measured 12,123 against a ceiling of
8,000. withRetry already fires three times and all three are refused, because
1.6 seconds of backoff does not buy back a minute's budget. The run ends
GatewayFailure and somebody reads "the model did not answer".
Retrying harder cannot fix a ceiling. Asking somebody else can — the ceilings
are per provider, so a second key is a second budget. Groq, Cerebras, Gemini,
Mistral and OpenRouter all serve the same chat-completions shape, which is why
this is a list of Configs and not a second implementation.
Configured as MODEL_FALLBACK_<n>_BASE_URL / _API_KEY / _FAST / _BALANCED /
_DEEP, numbered because five fields times three providers packed into one
delimited string is a parser nobody can read under pressure. Empty is the
ordinary case and returns the primary unwrapped, so a single-provider
deployment carries no wrapper and behaves exactly as before.
Failover is NOT unconditional, and the two guards are the design:
- Only a transient failure moves. Error.Retryable() already draws that line
for retries and it is the same line here. A 401 is this deployment's own
credential and a 400 is a malformed request; both fail identically at every
vendor, so trying three turns one visible fault into three invisible ones.
- Only an unpinned conversation moves. ToolCall.Extra carries provider
metadata echoed back verbatim — Gemini 3's thought signature — and a vendor
rejects a follow-up that drops its own. A conversation carrying any belongs
to whoever started it, so failover is available on the first model call,
which is where a rate limit usually lands anyway.
Streaming falls over only before the first fragment: once text is in the
reader's window, a second provider would continue that sentence in a different
voice.
What this does not do, since the gap is where the next bug lives: it does not
make a run cheaper, does not raise any one ceiling, and does not help when every
provider is exhausted at once. It turns one busy provider into a slower answer.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
174 lines
6.6 KiB
Go
174 lines
6.6 KiB
Go
package gateway
|
|
|
|
import (
|
|
"github.com/krow/krow-backend/go-api/internal/config"
|
|
)
|
|
|
|
// ProviderOpenAI names the only wire protocol this platform speaks.
|
|
//
|
|
// One constant, not an enum, because there is one implementation. "openai" is
|
|
// the chat-completions shape — which is NOT only OpenAI. Groq, Gemini (through
|
|
// its compatible endpoint), OpenRouter, Together, vLLM and a local Ollama all
|
|
// serve it, and the difference between them is MODEL_BASE_URL and a model id,
|
|
// nothing more. Supporting six vendors is one implementation and six base URLs.
|
|
//
|
|
// The Anthropic path was removed deliberately, not lost. `MODEL_PROVIDER=anthropic`
|
|
// is now REFUSED at startup rather than ignored — see config.validateModel. A
|
|
// deployment carrying the old value must be told it moved, because the silent
|
|
// alternative is a stack that believes it is still on Claude while every run
|
|
// goes somewhere else.
|
|
const ProviderOpenAI = "openai"
|
|
|
|
// Effort is how hard a tier is allowed to think.
|
|
//
|
|
// PROVIDER-NEUTRAL ON PURPOSE, and the reason that mattered is now history
|
|
// worth keeping: this was a vendor SDK's own enum, baked into the routing table
|
|
// every provider has to read. Making it the platform's own vocabulary is what
|
|
// let that vendor be removed later without the routing table going with it —
|
|
// a one-line deletion instead of a re-typing of every tier.
|
|
//
|
|
// The three values are the platform's own vocabulary. Each implementation maps
|
|
// them onto whatever its API calls the same idea, and a provider with no such
|
|
// concept ignores them — the tier still selects the model, which is the larger
|
|
// lever anyway.
|
|
type Effort string
|
|
|
|
const (
|
|
EffortLow Effort = "low"
|
|
EffortHigh Effort = "high"
|
|
EffortXhigh Effort = "xhigh"
|
|
)
|
|
|
|
// Routing is how a tier becomes a model and an effort level.
|
|
//
|
|
// The model per tier is a deployment knob — a tenant on a different contract,
|
|
// or a deployment pinning a version through an incident, changes it without a
|
|
// spec edit. The *effort* per tier is not: "fast" and "deep" mean something
|
|
// specific about how much work an answer is worth, and letting a deployment
|
|
// redefine that would make the same spec behave differently in two places
|
|
// while claiming the same tier.
|
|
type Routing struct {
|
|
Model string
|
|
Effort Effort
|
|
}
|
|
|
|
// Config is the gateway's whole configuration surface.
|
|
//
|
|
// Built once at startup from the environment and passed in frozen, per §10.
|
|
// Nothing in this package reads the environment itself.
|
|
type Config struct {
|
|
// Provider selects the implementation. Empty means openai, which is now
|
|
// the only one; config.validateModel refuses any other value.
|
|
Provider string
|
|
|
|
APIKey string
|
|
|
|
// BaseURL points the OpenAI-compatible path at a specific service. Empty
|
|
// means OpenAI itself. This is the field that turns one implementation
|
|
// into a choice between Groq, Gemini, OpenRouter and a local Ollama.
|
|
BaseURL string
|
|
|
|
Fast Routing
|
|
Balanced Routing
|
|
Deep Routing
|
|
|
|
// Fallbacks are further providers to try, in order, when this one cannot
|
|
// answer. See failover.go for when that is sound and when it is not.
|
|
Fallbacks []Config
|
|
|
|
// MaxOutputTokens applies when a request does not set its own.
|
|
MaxOutputTokens int64
|
|
|
|
// SendReasoningEffort controls whether the OpenAI path transmits the
|
|
// effort level as `reasoning_effort`.
|
|
//
|
|
// OFF BY DEFAULT, and that default is the careful one. Reasoning models
|
|
// accept the field; most others reject the whole request with a 400 rather
|
|
// than ignoring an unknown key. A run that dies on a malformed request is
|
|
// worse than a run that thinks at the model's own default, so a deployment
|
|
// on a reasoning-capable model opts in rather than every other deployment
|
|
// opting out.
|
|
SendReasoningEffort bool
|
|
}
|
|
|
|
// FromConfig builds the gateway's routing table from validated settings.
|
|
//
|
|
// The effort per tier is fixed here rather than configured, and that is the
|
|
// point of the function existing at all: a deployment chooses *which model*
|
|
// answers a tier, and the platform chooses *how hard it thinks*. If a
|
|
// deployment could redefine effort, two installations running the same
|
|
// definition would disagree about what "deep" means while both reporting the
|
|
// tier as deep — and the tier is written into every trajectory.
|
|
//
|
|
// fast → low a lookup, a restatement, a short structured reading
|
|
// balanced → high the default, and what most turns should cost
|
|
// deep → xhigh a turn worth several tool calls and real deliberation
|
|
//
|
|
// `max` is deliberately not reachable from a spec. It is the setting for when
|
|
// correctness matters more than cost, which is a judgement an operator makes
|
|
// about a deployment, not one an agent author makes about a page.
|
|
func FromConfig(c config.ModelConfig) Config {
|
|
var fallbacks []Config
|
|
for _, f := range c.Fallbacks {
|
|
// Model ids default to the primary's. Usually wrong for a different
|
|
// vendor and deliberately not silently corrected: an id the endpoint
|
|
// does not serve answers invalid_request, which is a visible fault an
|
|
// operator can fix, where a guessed substitution would be an invisible
|
|
// one nobody asked for.
|
|
if f.Fast == "" {
|
|
f.Fast = c.Fast
|
|
}
|
|
if f.Balanced == "" {
|
|
f.Balanced = c.Balanced
|
|
}
|
|
if f.Deep == "" {
|
|
f.Deep = c.Deep
|
|
}
|
|
fallbacks = append(fallbacks, FromConfig(f))
|
|
}
|
|
return Config{
|
|
Fallbacks: fallbacks,
|
|
Provider: c.Provider,
|
|
APIKey: c.APIKey,
|
|
BaseURL: c.BaseURL,
|
|
Fast: Routing{Model: c.Fast, Effort: EffortLow},
|
|
Balanced: Routing{Model: c.Balanced, Effort: EffortHigh},
|
|
Deep: Routing{Model: c.Deep, Effort: EffortXhigh},
|
|
MaxOutputTokens: int64(c.MaxOutputTokens),
|
|
SendReasoningEffort: c.ReasoningEffort,
|
|
}
|
|
}
|
|
|
|
// New builds the gateway a deployment's configuration asks for.
|
|
//
|
|
// One provider, so this is a constructor rather than a choice. It survives the
|
|
// removal of the second implementation because the runtime wires itself through
|
|
// `gateway.New(gateway.FromConfig(...))` and should not learn a concrete type:
|
|
// the next provider is a change here and nowhere else.
|
|
func New(cfg Config) Gateway {
|
|
primary := NewOpenAI(cfg)
|
|
if len(cfg.Fallbacks) == 0 {
|
|
return primary
|
|
}
|
|
rest := make([]Gateway, 0, len(cfg.Fallbacks))
|
|
for _, f := range cfg.Fallbacks {
|
|
rest = append(rest, NewOpenAI(f))
|
|
}
|
|
return NewFailover(primary, rest...)
|
|
}
|
|
|
|
// routingFor resolves a tier against a table.
|
|
//
|
|
// An unknown tier has already been normalised by ParseTier, so the default arm
|
|
// is reached only by a zero value.
|
|
func (c Config) routingFor(t Tier) Routing {
|
|
switch t {
|
|
case TierFast:
|
|
return c.Fast
|
|
case TierDeep:
|
|
return c.Deep
|
|
default:
|
|
return c.Balanced
|
|
}
|
|
}
|