The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.
`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.
`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.
THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.
Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.
Three configuration mistakes are refused at startup rather than at runtime:
- MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
endpoint and ignores the field, so this is a deployment that believes it
switched providers and did not — every run still goes to Anthropic and is
still billed there, with nothing in the logs to say so. Cost is the whole
reason this change exists, and that is the one mistake that silently
defeats it.
- An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
- A production deployment with no credential — except against localhost,
which needs none, and demanding one would make the free local path
impossible to configure.
reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.
`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.
Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.
NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
144 lines
5.4 KiB
Go
144 lines
5.4 KiB
Go
package gateway
|
|
|
|
import (
|
|
"github.com/krow/krow-backend/go-api/internal/config"
|
|
)
|
|
|
|
// Provider names the wire protocol a deployment talks.
|
|
//
|
|
// Two, not two hundred: "anthropic" is the Claude API, and "openai" is the
|
|
// chat-completions shape that Groq, Gemini, OpenRouter, Together, vLLM and
|
|
// Ollama all serve. That second one is the reason this constant exists at all
|
|
// — supporting those five providers is one implementation and five different
|
|
// base URLs, and pretending otherwise would grow a package per vendor.
|
|
const (
|
|
ProviderAnthropic = "anthropic"
|
|
ProviderOpenAI = "openai"
|
|
)
|
|
|
|
// Effort is how hard a tier is allowed to think.
|
|
//
|
|
// PROVIDER-NEUTRAL ON PURPOSE. This was `anthropic.OutputConfigEffort` until a
|
|
// second provider existed, which meant the vendor's enum was baked into the
|
|
// routing table that every provider has to read. Nothing was wrong with it
|
|
// while there was one implementation; it became wrong the moment there were
|
|
// two, because the OpenAI path would have had to import the Anthropic SDK to
|
|
// learn how hard to think.
|
|
//
|
|
// The three values are the platform's own vocabulary. Each implementation maps
|
|
// them onto whatever its API calls the same idea, and a provider with no such
|
|
// concept ignores them — the tier still selects the model, which is the larger
|
|
// lever anyway.
|
|
type Effort string
|
|
|
|
const (
|
|
EffortLow Effort = "low"
|
|
EffortHigh Effort = "high"
|
|
EffortXhigh Effort = "xhigh"
|
|
)
|
|
|
|
// Routing is how a tier becomes a model and an effort level.
|
|
//
|
|
// The model per tier is a deployment knob — a tenant on a different contract,
|
|
// or a deployment pinning a version through an incident, changes it without a
|
|
// spec edit. The *effort* per tier is not: "fast" and "deep" mean something
|
|
// specific about how much work an answer is worth, and letting a deployment
|
|
// redefine that would make the same spec behave differently in two places
|
|
// while claiming the same tier.
|
|
type Routing struct {
|
|
Model string
|
|
Effort Effort
|
|
}
|
|
|
|
// Config is the gateway's whole configuration surface.
|
|
//
|
|
// Built once at startup from the environment and passed in frozen, per §10.
|
|
// Nothing in this package reads the environment itself.
|
|
type Config struct {
|
|
// Provider selects the implementation. Empty means anthropic, so a
|
|
// deployment that predates the second provider keeps working untouched.
|
|
Provider string
|
|
|
|
APIKey string
|
|
|
|
// BaseURL points the OpenAI-compatible path at a specific service. Empty
|
|
// means OpenAI itself. This is the field that turns one implementation
|
|
// into a choice between Groq, Gemini, OpenRouter and a local Ollama.
|
|
BaseURL string
|
|
|
|
Fast Routing
|
|
Balanced Routing
|
|
Deep Routing
|
|
|
|
// MaxOutputTokens applies when a request does not set its own.
|
|
MaxOutputTokens int64
|
|
|
|
// SendReasoningEffort controls whether the OpenAI path transmits the
|
|
// effort level as `reasoning_effort`.
|
|
//
|
|
// OFF BY DEFAULT, and that default is the careful one. Reasoning models
|
|
// accept the field; most others reject the whole request with a 400 rather
|
|
// than ignoring an unknown key. A run that dies on a malformed request is
|
|
// worse than a run that thinks at the model's own default, so a deployment
|
|
// on a reasoning-capable model opts in rather than every other deployment
|
|
// opting out.
|
|
SendReasoningEffort bool
|
|
}
|
|
|
|
// FromConfig builds the gateway's routing table from validated settings.
|
|
//
|
|
// The effort per tier is fixed here rather than configured, and that is the
|
|
// point of the function existing at all: a deployment chooses *which model*
|
|
// answers a tier, and the platform chooses *how hard it thinks*. If a
|
|
// deployment could redefine effort, two installations running the same
|
|
// definition would disagree about what "deep" means while both reporting the
|
|
// tier as deep — and the tier is written into every trajectory.
|
|
//
|
|
// fast → low a lookup, a restatement, a short structured reading
|
|
// balanced → high the default, and what most turns should cost
|
|
// deep → xhigh a turn worth several tool calls and real deliberation
|
|
//
|
|
// `max` is deliberately not reachable from a spec. It is the setting for when
|
|
// correctness matters more than cost, which is a judgement an operator makes
|
|
// about a deployment, not one an agent author makes about a page.
|
|
func FromConfig(c config.ModelConfig) Config {
|
|
return Config{
|
|
Provider: c.Provider,
|
|
APIKey: c.APIKey,
|
|
BaseURL: c.BaseURL,
|
|
Fast: Routing{Model: c.Fast, Effort: EffortLow},
|
|
Balanced: Routing{Model: c.Balanced, Effort: EffortHigh},
|
|
Deep: Routing{Model: c.Deep, Effort: EffortXhigh},
|
|
MaxOutputTokens: int64(c.MaxOutputTokens),
|
|
SendReasoningEffort: c.ReasoningEffort,
|
|
}
|
|
}
|
|
|
|
// New builds the gateway a deployment's configuration asks for.
|
|
//
|
|
// The one place that maps a provider name to an implementation, so a caller
|
|
// wires a gateway without knowing which vendor answers. An unrecognised
|
|
// provider cannot reach here — config.validate rejects it at startup, where a
|
|
// typo is one loud failure instead of one per run.
|
|
func New(cfg Config) Gateway {
|
|
if cfg.Provider == ProviderOpenAI {
|
|
return NewOpenAI(cfg)
|
|
}
|
|
return NewAnthropic(cfg)
|
|
}
|
|
|
|
// routingFor resolves a tier against a table.
|
|
//
|
|
// Shared by both implementations: an unknown tier has already been normalised
|
|
// by ParseTier, so the default arm is reached only by a zero value.
|
|
func (c Config) routingFor(t Tier) Routing {
|
|
switch t {
|
|
case TierFast:
|
|
return c.Fast
|
|
case TierDeep:
|
|
return c.Deep
|
|
default:
|
|
return c.Balanced
|
|
}
|
|
}
|