Files
krow_backend/docs/deploy-9d3192a.md
Suriyakumarvijayanayagam 74089eb3e9
Some checks failed
CI / test (push) Failing after 4m36s
CI / fixture (push) Failing after 8s
Add the krow-2 deploy runbook for 9d3192a
The API on krow-2 is down on a startup guard, and the fix is two environment
variables. Written down rather than left in a chat log, in the same shape as
deploy-b6f8655.md: what is being deployed, what was actually verified before
claiming it works, the rollout, a smoke test, and rollback.

The smoke test insists on one real agent run. Boot and /health both pass with a
broken model configuration — that is precisely how the current outage stayed
invisible until run time — so a health check alone is not evidence the deploy
worked. Includes the symptom-to-cause table for the four failures this rollout
can actually produce.

Records what the preflight covered: the production image built and booted under
APP_ENV=production against a TLS Postgres, 11 migrations applied, seed and 9
agents imported, login, a real Groq-backed run terminating Completed, and SSE
streaming. 62 endpoints is written down as the expected number because two
fewer is the signature of a missing credential.

Not executed. No working SSH to that host from here, and a production rollout
wants the operator watching.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-07 13:06:49 +05:30

6.3 KiB
Raw Permalink Blame History

Deploying 9d3192a to krow-2

The API on krow-2 is currently down, and stayed down on purpose. It refuses to start with:

ERROR fatal error="ANTHROPIC_API_KEY is set but is no longer read, and
MODEL_API_KEY is empty: the Anthropic path was removed..."

That is a startup guard doing its job, not a crash. 34fa58a removed the Anthropic path, and the deployment environment still describes the old one. The fix is two environment variables; everything else here is the rollout around it.

Prepared and verified locally. Not executed — this machine has no working SSH to krow-2, and a production rollout is not something to do without the operator watching.


1. What is being deployed

9d3192a, which is HEAD and already origin/main. Nothing needs pushing.

Four commits since the last deploy point:

Commit What
34fa58a Anthropic path removed; the gateway speaks one wire protocol
bd9a8f9 The shipped example envs can actually start (see §5)
7d83c16 Compose comments: a container has no keyless option
9d3192a Model ids Groq actually serves; I7 measured

No migrations. migrations/ is unchanged since 34fa58a, so migrate will report nothing to apply and exit 0. This is a configuration rollout.

2. What was verified before writing this

The whole path, locally, in the production image against a TLS Postgres and the real Groq API:

Step Result
docker build -f infrastructure/Dockerfile.api builds
11 migrations against Postgres 18 over TLS all apply
Boot with APP_ENV=production, sslmode=require connects, 62 endpoints
seed + importagents --org=krow-dev 245 records, 9 agents, 24 skills
POST /api/v1/auth/login session cookie issued
POST /api/v1/agents/{id}/runs → Groq Completed, 2 model calls, 1707 tokens
Same with Accept: text/event-stream data: {"delta":"…"} streams
make eval-live ×2 all three cases pass, I7 included

62 endpoints is the number to expect. Fewer by two means the agent routes did not register, which means the credential is missing — see §6.

3. The environment change

On krow-2, in the .env that docker compose reads (beside docker-compose.yml):

MODEL_PROVIDER=openai
MODEL_BASE_URL=https://api.groq.com/openai/v1
MODEL_API_KEY=<the Groq key>
MODEL_FAST=openai/gpt-oss-20b
MODEL_BALANCED=openai/gpt-oss-120b
MODEL_DEEP=openai/gpt-oss-120b

And delete the ANTHROPIC_API_KEY line entirely. Commenting it out is enough; leaving it set with an empty MODEL_API_KEY reproduces the failure.

Two things that look like details and are not:

  • Replace the value, do not just rename the variable. An sk-ant-… key under the name MODEL_API_KEY passes every startup check — the process cannot tell one opaque string from another — and then fails every run with 401. Startup validation catches the shape of a stale configuration, never a wrong secret.
  • The model ids are not interchangeable. Groq no longer serves the llama-3.1-8b-instant / llama-3.3-70b-versatile pair that shipped in 34fa58a; both were wrong the day they shipped and would have 400'd on every run. The ids above were checked against the live account.

While in the file, confirm HTTP_WRITE_TIMEOUT is above 2m — 180s is the new default. Below that the container refuses to start. krow-2 only got past this before because someone had already overridden it.

4. Rollout

cd <krow-backend checkout on krow-2>
git fetch origin && git checkout main && git pull --ff-only origin main
git log --oneline -1        # expect 9d3192a

cd infrastructure
# edit .env per §3, then:
docker compose up -d --build

migrate runs to completion before api starts and api will not start if it fails. Expect migrate to exit 0 having applied nothing.

5. Smoke test

docker compose ps                       # api: Up (healthy)
docker compose logs api | tail -20      # no fatal; "listening" with endpoints=62
curl -s http://127.0.0.1:8080/health    # {"status":"ok"}

"status":"degraded" means the schema is missing or dirty, not a model problem.

Then one real agent run, which is the only step that proves the model path — the earlier outage was invisible until run time:

# from the host, against the published port
curl -s -c /tmp/k.jar -X POST http://127.0.0.1:8080/api/v1/auth/login \
  -H 'Content-Type: application/json' \
  -d '{"email":"<an account on krow-2>","password":"<its password>"}' >/dev/null

AGENT=$(docker compose exec -T api sh -c 'true' >/dev/null 2>&1; \
  psql "$DATABASE_URL" -tAc \
  "select id from agent_definitions where definition_id='activity-agent' and status='published' limit 1;")

curl -s -b /tmp/k.jar -X POST "http://127.0.0.1:8080/api/v1/agents/$AGENT/runs" \
  -H 'Content-Type: application/json' \
  -d '{"input":"How many events happened in the last 7 days?"}'

Expect "termination":"Completed" and a non-zero usage.totalTokens.

Symptom Cause
404 on the runs route no MODEL_API_KEY; the agent routes were never registered
the model credentials were refused the key is wrong — an Anthropic key renamed, or a bad Groq key
the model rejected the request naming a model id that id is not served; re-check §3 against GET /v1/models
502 from the proxy on slow runs HTTP_WRITE_TIMEOUT below 2m

6. Rollback

Nothing in the schema changed, so rollback is the previous image and the previous .env:

cd infrastructure && git checkout <previous sha> && docker compose up -d --build

Restoring ANTHROPIC_API_KEY will not bring the old behaviour back at any commit from 34fa58a onward — the provider is gone from the binary. Rolling back past it means rolling back the key too.

7. Not covered here

  • The frontend needs no redeploy. krow-demo is unchanged at 6249e00; every change in this rollout is server-side.
  • EMBED_PROVIDER=ollama on krow-2 points at host.docker.internal:11434. Unrelated to this rollout, but if Ollama is not running on that host, dense retrieval degrades to keyword-only rather than failing loudly.
  • The seed fixture generator in krow-demo is a version behind the seeder and would delete the employer@krow.app account if run. Do not run npm run seed:fixture as part of a deploy.