Add an OpenAI-compatible gateway, so the model provider is a config value
The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.
`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.
`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.
THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.
Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.
Three configuration mistakes are refused at startup rather than at runtime:
- MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
endpoint and ignores the field, so this is a deployment that believes it
switched providers and did not — every run still goes to Anthropic and is
still billed there, with nothing in the logs to say so. Cost is the whole
reason this change exists, and that is the one mistake that silently
defeats it.
- An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
- A production deployment with no credential — except against localhost,
which needs none, and demanding one would make the free local path
impossible to configure.
reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.
`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.
Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.
NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
This commit is contained in:
@@ -88,11 +88,30 @@ type KnowledgeConfig struct {
|
||||
// first model call, as a structured gateway.not_configured a run can end with,
|
||||
// not at startup as a refusal to boot.
|
||||
type ModelConfig struct {
|
||||
APIKey string
|
||||
// Provider names the wire protocol: "anthropic" or "openai". Empty means
|
||||
// anthropic, so a deployment that predates the second provider keeps
|
||||
// working with the environment it already has.
|
||||
//
|
||||
// "openai" is not only OpenAI. Groq, Gemini's compatibility endpoint,
|
||||
// OpenRouter, Together, vLLM and a local Ollama all serve that same shape,
|
||||
// and BaseURL is what chooses between them.
|
||||
Provider string
|
||||
|
||||
APIKey string
|
||||
|
||||
// BaseURL points the OpenAI-compatible provider at a specific service.
|
||||
// Ignored by the anthropic provider, which has one endpoint.
|
||||
BaseURL string
|
||||
|
||||
Fast string
|
||||
Balanced string
|
||||
Deep string
|
||||
MaxOutputTokens int
|
||||
|
||||
// ReasoningEffort opts into sending the tier's effort level on the
|
||||
// OpenAI-compatible wire. Off by default: reasoning models accept the
|
||||
// field and most others reject the entire request rather than ignoring it.
|
||||
ReasoningEffort bool
|
||||
}
|
||||
|
||||
// SeedConfig locates the demo fixture. The file is generated from the frontend
|
||||
@@ -245,7 +264,14 @@ func Load() (*Config, error) {
|
||||
UseLexicalEmbedder: boolDefault("EMBED_USE_LEXICAL", false),
|
||||
},
|
||||
Model: ModelConfig{
|
||||
APIKey: strings.TrimSpace(os.Getenv("ANTHROPIC_API_KEY")),
|
||||
Provider: strings.ToLower(strings.TrimSpace(os.Getenv("MODEL_PROVIDER"))),
|
||||
// MODEL_API_KEY first, then the Anthropic-specific name. Two
|
||||
// spellings because the second provider is not Anthropic and
|
||||
// ANTHROPIC_API_KEY=<a Groq key> would be a lie an operator has to
|
||||
// keep re-reading; the fallback keeps every existing deployment
|
||||
// working without an edit.
|
||||
APIKey: firstSet("MODEL_API_KEY", "ANTHROPIC_API_KEY"),
|
||||
BaseURL: strings.TrimSpace(os.Getenv("MODEL_BASE_URL")),
|
||||
Fast: withDefault("MODEL_FAST", defaultModel),
|
||||
Balanced: withDefault("MODEL_BALANCED", defaultModel),
|
||||
Deep: withDefault("MODEL_DEEP", defaultModel),
|
||||
@@ -254,6 +280,7 @@ func Load() (*Config, error) {
|
||||
// needs a long answer; this is the ceiling for a single
|
||||
// unstreamed call, not the run's budget.
|
||||
MaxOutputTokens: intDefault("MODEL_MAX_OUTPUT_TOKENS", 16000),
|
||||
ReasoningEffort: boolDefault("MODEL_REASONING_EFFORT", false),
|
||||
},
|
||||
DB: DBConfig{
|
||||
Host: required("DATABASE_HOST"),
|
||||
@@ -306,6 +333,45 @@ const DeepestAgentDeadline = 120 * time.Second
|
||||
// Streaming hides it, and that is the trap. The chat panel uses SSE and
|
||||
// survives, so the product looks healthy while every non-streaming caller — a
|
||||
// webhook, a script, an integration — gets 502 on a slow question.
|
||||
// validateModel refuses a model configuration that cannot work.
|
||||
//
|
||||
// Its own method for the same reason validateWriteTimeout is: these are the
|
||||
// mistakes that produce a *runtime* symptom far from their cause — a deployment
|
||||
// that believes it switched providers and is still being billed by the old one,
|
||||
// or a production install with no credential that fails one run at a time
|
||||
// instead of once at startup.
|
||||
func (c *Config) validateModel() error {
|
||||
switch c.Model.Provider {
|
||||
case "", "anthropic", "openai":
|
||||
default:
|
||||
return fmt.Errorf("MODEL_PROVIDER must be anthropic or openai, got %q", c.Model.Provider)
|
||||
}
|
||||
// A local model needs no credential, and demanding one would make the
|
||||
// zero-cost development path impossible to configure. Everything else does:
|
||||
// a production deployment without a key fails every run at the gateway,
|
||||
// which is a misconfiguration wearing a runtime error's clothes.
|
||||
if c.AppEnv == "production" && c.Model.APIKey == "" && !isLoopback(c.Model.BaseURL) {
|
||||
return fmt.Errorf("MODEL_API_KEY (or ANTHROPIC_API_KEY) is required when APP_ENV=production; " +
|
||||
"without it every agent run fails at the model gateway")
|
||||
}
|
||||
// A base URL is only read by the OpenAI-compatible provider. Setting one
|
||||
// while on anthropic is a deployment that believes it has switched
|
||||
// providers and has not — it would keep calling Claude and keep being
|
||||
// billed for it, with nothing in the logs to say so.
|
||||
if c.Model.BaseURL != "" && c.Model.Provider != "openai" {
|
||||
return fmt.Errorf("MODEL_BASE_URL only applies when MODEL_PROVIDER=openai; "+
|
||||
"it is set to %q but the provider is %q, so the base URL would be ignored "+
|
||||
"and every run would still go to Anthropic", c.Model.BaseURL, providerName(c.Model.Provider))
|
||||
}
|
||||
if c.Model.BaseURL != "" {
|
||||
u, err := url.Parse(c.Model.BaseURL)
|
||||
if err != nil || (u.Scheme != "http" && u.Scheme != "https") || u.Host == "" {
|
||||
return fmt.Errorf("MODEL_BASE_URL must be an http or https URL, got %q", c.Model.BaseURL)
|
||||
}
|
||||
}
|
||||
return nil
|
||||
}
|
||||
|
||||
func (c *Config) validateWriteTimeout() error {
|
||||
if c.HTTP.WriteTimeout <= 0 {
|
||||
return nil // no deadline set; the server will not cut anything off
|
||||
@@ -355,9 +421,8 @@ func (c *Config) validate() error {
|
||||
// misconfiguration wearing a runtime error's clothes, so it is caught here.
|
||||
// Development is left alone deliberately: working on migrations or the
|
||||
// definitions API must not require a key.
|
||||
if c.AppEnv == "production" && c.Model.APIKey == "" {
|
||||
return fmt.Errorf("ANTHROPIC_API_KEY is required when APP_ENV=production; " +
|
||||
"without it every agent run fails at the model gateway")
|
||||
if err := c.validateModel(); err != nil {
|
||||
return err
|
||||
}
|
||||
if c.Model.MaxOutputTokens < 1 {
|
||||
return fmt.Errorf("MODEL_MAX_OUTPUT_TOKENS must be at least 1, got %d", c.Model.MaxOutputTokens)
|
||||
@@ -482,6 +547,46 @@ func withDefault(key, fallback string) string {
|
||||
return fallback
|
||||
}
|
||||
|
||||
// firstSet returns the first of several environment variables that has a value.
|
||||
//
|
||||
// For settings that have more than one legitimate spelling — a generic name and
|
||||
// a provider-specific one — where the order expresses which wins rather than
|
||||
// leaving it to whichever happens to be read last.
|
||||
func firstSet(keys ...string) string {
|
||||
for _, k := range keys {
|
||||
if v := strings.TrimSpace(os.Getenv(k)); v != "" {
|
||||
return v
|
||||
}
|
||||
}
|
||||
return ""
|
||||
}
|
||||
|
||||
// isLoopback reports whether a base URL points at this machine.
|
||||
//
|
||||
// A model served from localhost needs no credential, and requiring one would
|
||||
// make the zero-cost local path impossible to configure. Host-only, so a
|
||||
// remote service that merely mentions "localhost" in a path does not qualify.
|
||||
func isLoopback(raw string) bool {
|
||||
if strings.TrimSpace(raw) == "" {
|
||||
return false
|
||||
}
|
||||
u, err := url.Parse(raw)
|
||||
if err != nil {
|
||||
return false
|
||||
}
|
||||
host := u.Hostname()
|
||||
return host == "localhost" || host == "127.0.0.1" || host == "::1"
|
||||
}
|
||||
|
||||
// providerName renders the provider for an error message, naming the default
|
||||
// rather than showing an empty string an operator then has to interpret.
|
||||
func providerName(p string) string {
|
||||
if p == "" {
|
||||
return "anthropic (the default)"
|
||||
}
|
||||
return p
|
||||
}
|
||||
|
||||
func intDefault(key string, fallback int) int {
|
||||
v := strings.TrimSpace(os.Getenv(key))
|
||||
if v == "" {
|
||||
|
||||
128
go-api/internal/config/model_test.go
Normal file
128
go-api/internal/config/model_test.go
Normal file
@@ -0,0 +1,128 @@
|
||||
package config
|
||||
|
||||
import (
|
||||
"strings"
|
||||
"testing"
|
||||
)
|
||||
|
||||
func modelCfg(env string, m ModelConfig) *Config {
|
||||
c := &Config{AppEnv: env}
|
||||
c.Model = m
|
||||
return c
|
||||
}
|
||||
|
||||
func TestValidateModelProvider(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
cfg *Config
|
||||
wantErr bool
|
||||
}{
|
||||
{
|
||||
"unset provider is anthropic, which is what every existing deployment has",
|
||||
modelCfg("development", ModelConfig{}), false,
|
||||
},
|
||||
{"anthropic named explicitly", modelCfg("development", ModelConfig{Provider: "anthropic"}), false},
|
||||
{"openai", modelCfg("development", ModelConfig{Provider: "openai"}), false},
|
||||
{"a typo is caught once at startup, not once per run",
|
||||
modelCfg("development", ModelConfig{Provider: "openal"}), true},
|
||||
{"a provider that does not exist", modelCfg("development", ModelConfig{Provider: "groq"}), true},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
err := tc.cfg.validateModel()
|
||||
if tc.wantErr != (err != nil) {
|
||||
t.Fatalf("validateModel() = %v, wantErr = %v", err, tc.wantErr)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
// THE EXPENSIVE MISTAKE.
|
||||
//
|
||||
// A deployment that sets MODEL_BASE_URL and forgets MODEL_PROVIDER believes it
|
||||
// has moved off Claude. It has not: the anthropic path has one endpoint and
|
||||
// ignores the field entirely, so every run keeps going to Anthropic and keeps
|
||||
// being billed there, with nothing in the logs to say so. The whole point of
|
||||
// this change is cost, and that is the one misconfiguration that silently
|
||||
// defeats it.
|
||||
func TestBaseURLWithoutOpenAIProviderIsRefused(t *testing.T) {
|
||||
err := modelCfg("development", ModelConfig{
|
||||
BaseURL: "https://api.groq.com/openai/v1",
|
||||
}).validateModel()
|
||||
if err == nil {
|
||||
t.Fatal("a base URL on the anthropic provider was accepted; every run would still go to Anthropic")
|
||||
}
|
||||
for _, want := range []string{"MODEL_BASE_URL", "MODEL_PROVIDER=openai", "Anthropic"} {
|
||||
if !strings.Contains(err.Error(), want) {
|
||||
t.Errorf("the message does not mention %q:\n %v", want, err)
|
||||
}
|
||||
}
|
||||
|
||||
// The same URL with the provider set is exactly the intended configuration.
|
||||
if err := modelCfg("development", ModelConfig{
|
||||
Provider: "openai", BaseURL: "https://api.groq.com/openai/v1",
|
||||
}).validateModel(); err != nil {
|
||||
t.Fatalf("the intended configuration was refused: %v", err)
|
||||
}
|
||||
}
|
||||
|
||||
func TestBaseURLMustBeAURL(t *testing.T) {
|
||||
for _, raw := range []string{"api.groq.com", "ftp://x.test", "not a url", "://broken"} {
|
||||
err := modelCfg("development", ModelConfig{Provider: "openai", BaseURL: raw}).validateModel()
|
||||
if err == nil {
|
||||
t.Errorf("MODEL_BASE_URL=%q was accepted", raw)
|
||||
}
|
||||
}
|
||||
for _, raw := range []string{"http://localhost:11434/v1", "https://api.groq.com/openai/v1"} {
|
||||
if err := modelCfg("development", ModelConfig{Provider: "openai", BaseURL: raw}).validateModel(); err != nil {
|
||||
t.Errorf("MODEL_BASE_URL=%q was refused: %v", raw, err)
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Production without a credential fails every run at the gateway, which is a
|
||||
// misconfiguration wearing a runtime error's clothes. A local model is the one
|
||||
// exception: it needs no key, and demanding one would make the zero-cost path
|
||||
// impossible to configure.
|
||||
func TestProductionCredentialRequirement(t *testing.T) {
|
||||
for _, tc := range []struct {
|
||||
name string
|
||||
cfg *Config
|
||||
wantErr bool
|
||||
}{
|
||||
{"production with no key", modelCfg("production", ModelConfig{}), true},
|
||||
{"production with a key", modelCfg("production", ModelConfig{APIKey: "k"}), false},
|
||||
{
|
||||
"production against a local model needs no key",
|
||||
modelCfg("production", ModelConfig{Provider: "openai", BaseURL: "http://localhost:11434/v1"}),
|
||||
false,
|
||||
},
|
||||
{
|
||||
"production against a hosted provider still does",
|
||||
modelCfg("production", ModelConfig{Provider: "openai", BaseURL: "https://api.groq.com/openai/v1"}),
|
||||
true,
|
||||
},
|
||||
{"development needs nothing", modelCfg("development", ModelConfig{}), false},
|
||||
} {
|
||||
t.Run(tc.name, func(t *testing.T) {
|
||||
err := tc.cfg.validateModel()
|
||||
if tc.wantErr != (err != nil) {
|
||||
t.Fatalf("validateModel() = %v, wantErr = %v", err, tc.wantErr)
|
||||
}
|
||||
})
|
||||
}
|
||||
}
|
||||
|
||||
func TestIsLoopback(t *testing.T) {
|
||||
for raw, want := range map[string]bool{
|
||||
"http://localhost:11434/v1": true,
|
||||
"http://127.0.0.1:11434/v1": true,
|
||||
"https://api.groq.com/v1": false,
|
||||
"": false,
|
||||
// A remote host that merely mentions localhost in its path is not local.
|
||||
"https://x.test/localhost/v1": false,
|
||||
} {
|
||||
if got := isLoopback(raw); got != want {
|
||||
t.Errorf("isLoopback(%q) = %v, want %v", raw, got, want)
|
||||
}
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user