The API on krow-2 is down on a startup guard, and the fix is two environment variables. Written down rather than left in a chat log, in the same shape as deploy-b6f8655.md: what is being deployed, what was actually verified before claiming it works, the rollout, a smoke test, and rollback. The smoke test insists on one real agent run. Boot and /health both pass with a broken model configuration — that is precisely how the current outage stayed invisible until run time — so a health check alone is not evidence the deploy worked. Includes the symptom-to-cause table for the four failures this rollout can actually produce. Records what the preflight covered: the production image built and booted under APP_ENV=production against a TLS Postgres, 11 migrations applied, seed and 9 agents imported, login, a real Groq-backed run terminating Completed, and SSE streaming. 62 endpoints is written down as the expected number because two fewer is the signature of a missing credential. Not executed. No working SSH to that host from here, and a production rollout wants the operator watching. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
6.3 KiB
Deploying 9d3192a to krow-2
The API on krow-2 is currently down, and stayed down on purpose. It refuses to start with:
ERROR fatal error="ANTHROPIC_API_KEY is set but is no longer read, and
MODEL_API_KEY is empty: the Anthropic path was removed..."
That is a startup guard doing its job, not a crash. 34fa58a removed the
Anthropic path, and the deployment environment still describes the old one. The
fix is two environment variables; everything else here is the rollout around it.
Prepared and verified locally. Not executed — this machine has no working SSH to krow-2, and a production rollout is not something to do without the operator watching.
1. What is being deployed
9d3192a, which is HEAD and already origin/main. Nothing needs pushing.
Four commits since the last deploy point:
| Commit | What |
|---|---|
34fa58a |
Anthropic path removed; the gateway speaks one wire protocol |
bd9a8f9 |
The shipped example envs can actually start (see §5) |
7d83c16 |
Compose comments: a container has no keyless option |
9d3192a |
Model ids Groq actually serves; I7 measured |
No migrations. migrations/ is unchanged since 34fa58a, so migrate will
report nothing to apply and exit 0. This is a configuration rollout.
2. What was verified before writing this
The whole path, locally, in the production image against a TLS Postgres and the real Groq API:
| Step | Result |
|---|---|
docker build -f infrastructure/Dockerfile.api |
builds |
| 11 migrations against Postgres 18 over TLS | all apply |
Boot with APP_ENV=production, sslmode=require |
connects, 62 endpoints |
seed + importagents --org=krow-dev |
245 records, 9 agents, 24 skills |
POST /api/v1/auth/login |
session cookie issued |
POST /api/v1/agents/{id}/runs → Groq |
Completed, 2 model calls, 1707 tokens |
Same with Accept: text/event-stream |
data: {"delta":"…"} streams |
make eval-live ×2 |
all three cases pass, I7 included |
62 endpoints is the number to expect. Fewer by two means the agent routes did not register, which means the credential is missing — see §6.
3. The environment change
On krow-2, in the .env that docker compose reads (beside
docker-compose.yml):
MODEL_PROVIDER=openai
MODEL_BASE_URL=https://api.groq.com/openai/v1
MODEL_API_KEY=<the Groq key>
MODEL_FAST=openai/gpt-oss-20b
MODEL_BALANCED=openai/gpt-oss-120b
MODEL_DEEP=openai/gpt-oss-120b
And delete the ANTHROPIC_API_KEY line entirely. Commenting it out is
enough; leaving it set with an empty MODEL_API_KEY reproduces the failure.
Two things that look like details and are not:
- Replace the value, do not just rename the variable. An
sk-ant-…key under the nameMODEL_API_KEYpasses every startup check — the process cannot tell one opaque string from another — and then fails every run with 401. Startup validation catches the shape of a stale configuration, never a wrong secret. - The model ids are not interchangeable. Groq no longer serves the
llama-3.1-8b-instant/llama-3.3-70b-versatilepair that shipped in34fa58a; both were wrong the day they shipped and would have 400'd on every run. The ids above were checked against the live account.
While in the file, confirm HTTP_WRITE_TIMEOUT is above 2m — 180s is the
new default. Below that the container refuses to start. krow-2 only got past
this before because someone had already overridden it.
4. Rollout
cd <krow-backend checkout on krow-2>
git fetch origin && git checkout main && git pull --ff-only origin main
git log --oneline -1 # expect 9d3192a
cd infrastructure
# edit .env per §3, then:
docker compose up -d --build
migrate runs to completion before api starts and api will not start if it
fails. Expect migrate to exit 0 having applied nothing.
5. Smoke test
docker compose ps # api: Up (healthy)
docker compose logs api | tail -20 # no fatal; "listening" with endpoints=62
curl -s http://127.0.0.1:8080/health # {"status":"ok"}
"status":"degraded" means the schema is missing or dirty, not a model problem.
Then one real agent run, which is the only step that proves the model path — the earlier outage was invisible until run time:
# from the host, against the published port
curl -s -c /tmp/k.jar -X POST http://127.0.0.1:8080/api/v1/auth/login \
-H 'Content-Type: application/json' \
-d '{"email":"<an account on krow-2>","password":"<its password>"}' >/dev/null
AGENT=$(docker compose exec -T api sh -c 'true' >/dev/null 2>&1; \
psql "$DATABASE_URL" -tAc \
"select id from agent_definitions where definition_id='activity-agent' and status='published' limit 1;")
curl -s -b /tmp/k.jar -X POST "http://127.0.0.1:8080/api/v1/agents/$AGENT/runs" \
-H 'Content-Type: application/json' \
-d '{"input":"How many events happened in the last 7 days?"}'
Expect "termination":"Completed" and a non-zero usage.totalTokens.
| Symptom | Cause |
|---|---|
404 on the runs route |
no MODEL_API_KEY; the agent routes were never registered |
the model credentials were refused |
the key is wrong — an Anthropic key renamed, or a bad Groq key |
the model rejected the request naming a model id |
that id is not served; re-check §3 against GET /v1/models |
502 from the proxy on slow runs |
HTTP_WRITE_TIMEOUT below 2m |
6. Rollback
Nothing in the schema changed, so rollback is the previous image and the
previous .env:
cd infrastructure && git checkout <previous sha> && docker compose up -d --build
Restoring ANTHROPIC_API_KEY will not bring the old behaviour back at any
commit from 34fa58a onward — the provider is gone from the binary. Rolling
back past it means rolling back the key too.
7. Not covered here
- The frontend needs no redeploy.
krow-demois unchanged at6249e00; every change in this rollout is server-side. EMBED_PROVIDER=ollamaon krow-2 points athost.docker.internal:11434. Unrelated to this rollout, but if Ollama is not running on that host, dense retrieval degrades to keyword-only rather than failing loudly.- The seed fixture generator in
krow-demois a version behind the seeder and would delete theemployer@krow.appaccount if run. Do not runnpm run seed:fixtureas part of a deploy.