Remove the Anthropic path; the gateway speaks one wire protocol
The platform now runs on Groq by default, through the OpenAI-compatible chat-completions shape. That shape is not one vendor — Gemini, OpenRouter, Together, vLLM and a local Ollama serve it too — so moving again stays configuration rather than code. Two things in the deleted file were not Anthropic's and would have gone with it silently: withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and CALLED BY openai.go. Deleting the file wholesale would have removed the retry policy of the provider that survived, and nothing in openai.go mentions it, so the loss would have been invisible until the next 429. The policy is a property of this platform's runs, not of a vendor's API; it now lives in retry.go where no provider can carry it off. StreamComplete had the same problem and moves to gateway.go, beside the Streamer interface whose comment already referenced it. Three stale-configuration failures are now refused at startup instead of being ignored. Each was verified firing through the real config.Load(): MODEL_PROVIDER=anthropic — named separately from every other wrong value because it used to be correct. Ignoring it gives a stack that believes it is on Claude while every run goes to Groq and is billed there. ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an operator did set is the worst version of this: they fail every run on a missing credential they are looking straight at. A leftover claude-* model id, naming the tier that carries it. This is the check the previous commit's error-detail work was diagnosing: such an id is accepted by this process, rejected by the provider, and 400s on EVERY run. "A model is wrong" does not say which of three lines to edit. Defaults ship as a matched pair. defaultBaseURL and the three tier ids are one decision, not four: an id is only meaningful against the service that serves it, and a Groq id on an OpenAI base URL is the same failure from the other side. The tiers also stop being one model — a tier whose cost does not differ is a distinction that buys nothing. Verified end to end against a stub of the wire, driving the real wiring (config.Load in production mode, gateway.New, StreamComplete): streamed deltas, tool-call decoding, the loopback credential exemption, and usage totalling 150 rather than 190 — the cached-prefix subtraction still holds. gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still needs a reachable database. NOT verified: the I7 planted-injection eval. Removing this path removed the only model whose refusal behaviour had been measured against it, so the new default is unproven there until `make eval-live` runs with a real key. The Groq model ids should also be confirmed against Groq's current lineup. Flagged in CLAUDE.md §12 and docs/handover.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
This commit is contained in:
@@ -184,19 +184,25 @@ configmap before shipping an image that contains the check.
|
||||
|
||||
## Changing model provider
|
||||
|
||||
The gateway speaks two wire protocols. `anthropic` is the Claude API.
|
||||
`openai` is the chat-completions shape — and that one is not only OpenAI:
|
||||
Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM and a local
|
||||
Ollama all serve it, so moving between them is configuration, not code.
|
||||
The gateway speaks one wire protocol: `openai`, the chat-completions shape.
|
||||
That is not the same as one vendor — Groq, Gemini's compatibility endpoint,
|
||||
OpenRouter, Together, vLLM and a local Ollama all serve it, so moving between
|
||||
them is configuration, not code.
|
||||
|
||||
**The Anthropic path was removed.** `MODEL_PROVIDER=anthropic` and a stale
|
||||
`ANTHROPIC_API_KEY` are both *refused at startup* rather than ignored, and so
|
||||
is a leftover `claude-*` model id. That is deliberate: each of those would
|
||||
otherwise produce a service that boots cleanly and fails every agent run.
|
||||
|
||||
The default with nothing set is Groq.
|
||||
|
||||
```bash
|
||||
# Groq
|
||||
MODEL_PROVIDER=openai
|
||||
# Groq (the default — base URL and ids below are what you get unset)
|
||||
MODEL_BASE_URL=https://api.groq.com/openai/v1
|
||||
MODEL_API_KEY=<key>
|
||||
MODEL_FAST=llama-3.1-8b-instant
|
||||
MODEL_BALANCED=openai/gpt-oss-120b
|
||||
MODEL_DEEP=openai/gpt-oss-120b
|
||||
MODEL_BALANCED=llama-3.3-70b-versatile
|
||||
MODEL_DEEP=llama-3.3-70b-versatile
|
||||
|
||||
# Gemini
|
||||
MODEL_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
|
||||
@@ -207,11 +213,10 @@ MODEL_BASE_URL=http://localhost:11434/v1
|
||||
|
||||
Four things worth knowing before you do it.
|
||||
|
||||
**Set `MODEL_PROVIDER`, not just the base URL.** The anthropic path has one
|
||||
endpoint and ignores `MODEL_BASE_URL` entirely, so setting the URL alone is a
|
||||
deployment that believes it has switched providers and has not — every run
|
||||
still goes to Anthropic and is still billed there. Config validation refuses
|
||||
that combination at startup rather than letting it run up a bill quietly.
|
||||
**A model id and a base URL are one decision, not two.** An id is only
|
||||
meaningful against the service that serves it, so changing the endpoint without
|
||||
changing the ids gives you a process that starts fine and 400s on every run.
|
||||
The defaults ship as a matched Groq pair for that reason.
|
||||
|
||||
**Leave `MODEL_REASONING_EFFORT` off unless every configured model is a
|
||||
reasoning model.** Reasoning models accept the field; most others reject the
|
||||
@@ -220,19 +225,23 @@ reasoning model.** Reasoning models accept the field; most others reject the
|
||||
**Run the evals before trusting it, and read the I7 case first.**
|
||||
|
||||
```bash
|
||||
MODEL_PROVIDER=openai MODEL_BASE_URL=… MODEL_API_KEY=… MODEL_BALANCED=… make eval-live
|
||||
MODEL_BASE_URL=… MODEL_API_KEY=… MODEL_BALANCED=… make eval-live
|
||||
```
|
||||
|
||||
`liveGateway` reads the same environment the service does and logs which
|
||||
provider and model answered. The handbook corpus contains a planted prompt
|
||||
injection; Claude refuses it and reports the document as tampered with. **A
|
||||
model that answers every other case well and follows that injection is not a
|
||||
cheaper option — it is a security regression.** That case is the gate, not the
|
||||
cost table.
|
||||
injection. A model worth running refuses it and reports the document as
|
||||
tampered with. **A model that answers every other case well and follows that
|
||||
injection is not a cheaper option — it is a security regression.** That case is
|
||||
the gate, not the cost table.
|
||||
|
||||
**Token accounting differs between the two wires and is already reconciled.**
|
||||
OpenAI reports `prompt_tokens` *inclusive* of the cached prefix; Anthropic
|
||||
reports input tokens *exclusive* of it. `oaiUsage.normalise` subtracts, because
|
||||
This one is not optional now: the removed provider was the one whose refusal
|
||||
behaviour had actually been measured here, so whatever replaces it is unproven
|
||||
against I7 until this suite says otherwise.
|
||||
|
||||
**Token accounting is already reconciled, and the subtraction is load-bearing.**
|
||||
This wire reports `prompt_tokens` *inclusive* of the cached prefix, while
|
||||
`Usage` carries the cached figure separately. `oaiUsage.normalise` subtracts, because
|
||||
`Usage.Total()` sums all four fields and copying both numbers across verbatim
|
||||
would bill the cached prefix twice — worst on long conversations, which is
|
||||
exactly where I3's budget matters most. Don't "simplify" that subtraction away;
|
||||
|
||||
Reference in New Issue
Block a user