The platform now runs on Groq by default, through the OpenAI-compatible
chat-completions shape. That shape is not one vendor — Gemini, OpenRouter,
Together, vLLM and a local Ollama serve it too — so moving again stays
configuration rather than code.
Two things in the deleted file were not Anthropic's and would have gone
with it silently:
withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and
CALLED BY openai.go. Deleting the file wholesale would have removed the
retry policy of the provider that survived, and nothing in openai.go
mentions it, so the loss would have been invisible until the next 429.
The policy is a property of this platform's runs, not of a vendor's API;
it now lives in retry.go where no provider can carry it off.
StreamComplete had the same problem and moves to gateway.go, beside the
Streamer interface whose comment already referenced it.
Three stale-configuration failures are now refused at startup instead of
being ignored. Each was verified firing through the real config.Load():
MODEL_PROVIDER=anthropic — named separately from every other wrong value
because it used to be correct. Ignoring it gives a stack that believes it
is on Claude while every run goes to Groq and is billed there.
ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an
operator did set is the worst version of this: they fail every run on a
missing credential they are looking straight at.
A leftover claude-* model id, naming the tier that carries it. This is
the check the previous commit's error-detail work was diagnosing: such an
id is accepted by this process, rejected by the provider, and 400s on
EVERY run. "A model is wrong" does not say which of three lines to edit.
Defaults ship as a matched pair. defaultBaseURL and the three tier ids are
one decision, not four: an id is only meaningful against the service that
serves it, and a Groq id on an OpenAI base URL is the same failure from the
other side. The tiers also stop being one model — a tier whose cost does
not differ is a distinction that buys nothing.
Verified end to end against a stub of the wire, driving the real wiring
(config.Load in production mode, gateway.New, StreamComplete): streamed
deltas, tool-call decoding, the loopback credential exemption, and usage
totalling 150 rather than 190 — the cached-prefix subtraction still holds.
gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still
needs a reachable database.
NOT verified: the I7 planted-injection eval. Removing this path removed the
only model whose refusal behaviour had been measured against it, so the new
default is unproven there until `make eval-live` runs with a real key. The
Groq model ids should also be confirmed against Groq's current lineup.
Flagged in CLAUDE.md §12 and docs/handover.md.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
§9 says no agent ships without evals. Eight of the nine had none: the two
other suites in evals/ are harness fixtures rather than agents in the
registry, so the rule was being met by one agent in nine.
Evals — 40 new cases, five per agent, every one carrying mustNotLeak:
- the agent is loaded from its real spec in agents/*.md rather than
written out again in Go. A hand-copied agent tests the copy: it keeps
passing after somebody edits the spec, which is the moment it most
needed to fail.
- callNamed calls the tool a case names. toolThenAnswer always called
tools[0], so seven of positions-agent's eight tools were unreachable,
and a boundary nothing calls is a boundary nothing tests.
- seedWorkspace fills BOTH tenants. A leak test against an empty second
tenant cannot fail.
Verified by breaking workersByScore's org predicate: six cases across four
agents fail with LEAKED "RIVAL".
Knowledge — six policy documents, taking the corpus from 2 to 8 (34
chunks). Three restricted to admin and employer, five tenant-wide. They
cover what the tools cannot: a tool reports how many shifts went unworked,
a policy says what cover costs inside 24 hours.
corpus_test.go treats those documents as product rather than fixtures. The
first version was tautological — it read audience: from a file and checked
that file's audience was enforced, so opening a restricted document passed.
mustNotBeTenantWide now holds that judgement apart from the files, with the
reason recorded for each.
CI — the checks this repository already had, made unskippable. testutil
calls t.Skipf on an unreachable database, so a dead service container would
produce a green build over a suite that ran almost nothing. Simulated: go
test exits 0 with 74 tests skipped, including every tenant-isolation test.
The guard exits 1 and names them, while still allowing TestLive* to skip
without a model key.
This CI tests; it does not deploy. The README's claim that migrations are
run by CI against the target database remains aspirational.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu