The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.
`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.
`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.
THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.
Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.
Three configuration mistakes are refused at startup rather than at runtime:
- MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
endpoint and ignores the field, so this is a deployment that believes it
switched providers and did not — every run still goes to Anthropic and is
still billed there, with nothing in the logs to say so. Cost is the whole
reason this change exists, and that is the one mistake that silently
defeats it.
- An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
- A production deployment with no credential — except against localhost,
which needs none, and demanding one would make the free local path
impossible to configure.
reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.
`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.
Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.
NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
Production had no EMBED_PROVIDER, so every knowledge_chunk carried a null
embedding and a question only matched documents that shared its words. A person
asking about a family emergency got nothing from a document titled "shift cover
and cancellation".
Ollama rather than Voyage: internal/knowledge/embed.go calls it "the default
worth reaching for" — real semantics, no credential, no per-token cost, and no
tenant text leaving the cluster. Voyage needs an API key nobody has issued.
Bounded deliberately. The API pods share this node, so an unbounded model
server is a way to evict them; the memory limit means the kubelet kills the
embedder and nothing else. The 1Gi request is also what keeps it off the second
node, which has 1.2Gi allocatable and could not hold it.
Applied in three stages so nothing was pointed at an embedder that had not
been proven: deploy and pull the model, run reembed with the settings passed as
exec environment — 34 chunks in 11s, which proves connectivity without touching
live config — and only then patch krow-config and restart. Rolling back is
removing four keys and restarting.
Verified after: 55/55 on verify-deploy, and a question with no literal keyword
overlap with the corpus returned the relevant policy documents.
This file is the record of what was applied. It was applied by hand, which is
the same gap the README already admits for migrations — there is no deploy
pipeline, so a manifest in the repository is a description of the cluster
rather than the thing that produces it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
The overlay had never been run against a fresh volume. Two faults, the first
hiding the second:
- postgres:16-alpine ships libssl but not the openssl CLI, so the first-boot
certificate generation exited 127 in a restart loop. It failed invisibly:
the 2>/dev/null on the openssl line swallowed sh's "not found" as well, so
`docker logs` was completely empty. openssl is now installed on the boot
that generates the certificate, inside the same guard, so a restart still
needs no network.
- the certificate was written into /var/lib/postgresql/data BEFORE initdb
ran, and initdb refuses to initialise a directory that is not empty. That
made a fresh volume unstartable regardless of the first fault. The
certificate now lives in its own volume, which keeps it persistent — the
reason it was put in the data directory — without touching the cluster's.
Separately, docker-compose.yml did not pass ANTHROPIC_API_KEY to the api
container, so a compose deployment could never register the agent run routes:
POST /agents/{id}/runs answered 404 and /version reported two endpoints fewer.
The model and embedder variables are now passed through, all defaulting to
empty so a deployment without them behaves exactly as it did.
Verified on a fresh volume: 56/56 verify-deploy checks against the resulting
stack, including a live agent run and 34 chunks embedded through Ollama.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
CORS
cors.go never set Access-Control-Allow-Credentials, so the
cookie-authenticated API was unreadable from any cross-origin frontend:
the server answered correctly and the browser blocked the page from
reading it. Set for allowlisted origins on both the preflight and the
actual response. Three tests added.
HTTP_COOKIE_SAMESITE (lax|none|strict, default lax) is new. CORS is only
half of what a cross-origin browser call needs; SameSite is judged on
registrable domain, so a frontend on an unrelated domain gets perfect CORS
headers and still no cookie. "none" is the only value that survives that,
and validate() refuses it without the Secure flag.
The "*" rejection now explains itself: browsers refuse Allow-Origin "*"
together with credentials, so it would break every authenticated call
rather than loosen anything.
Transactional endpoints (api-contract.md 12.1)
POST /api/v1/job-applications/{id}/hire
POST /api/v1/job-postings/{id}/assignments
Replaces two client-side loops that wrote several records with no
transaction and no rollback. Each is now one endpoint and one transaction,
built over repo.Repo so org scoping, derived columns, type casts and error
translation are not re-derived. Authorization reuses the existing policy
table rather than adding a parallel one: a workflow is exactly as
privileged as the writes it performs. 13 tests, including both rollback
paths.
Bug fix in the repository layer
repo.bindValue handled int64/int/float64/string but not int32, which is
what pgx returns for a PostgreSQL `int` column. Nothing previously read a
record and wrote one of its fields elsewhere, so it never surfaced; the
hire flow does exactly that and failed with "ai_score must be a number".
Both KindInt and KindFloat now accept the widths pgx actually produces.
Deployment
infrastructure/Dockerfile.api multi-stage, cross-compiling (BUILDPLATFORM
+ GOARCH) so linux/amd64 builds from arm64 are compiled rather than
emulated. Alpine runtime, non-root uid 10001, 22.1 MB. Ships api, seed,
setpassword and migrate, plus the migrations, so a Kubernetes
initContainer can apply the schema from the same image and tag as the
API. HEALTHCHECK keys on status code, not body, so a "degraded" instance
is not pulled from rotation during a migration window.
infrastructure/docker-compose.yml migrations run to completion before the
API starts. Assumes a managed PostgreSQL; the local-db overlay adds one
with TLS enabled so APP_ENV=production is met rather than dodged.
scripts/drop_public_tables.go the one-off used to clear an unrelated
schema from krowdb on 2026-08-24, kept for the record. Build-tagged
ignore and gated on CONFIRM_DROP=yes.
Verified against PostgreSQL: 16/16 new tests pass, and the image was built,
run and exercised end to end (login, CORS preflight, authenticated reads,
transaction rollback).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CmQiGq73Uyfq7J4yR8Vxxw