Files
krow_backend/docs/deploy-db4803c.md
Suriya 939598a187
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 10s
Add the deploy runbook for db4803c and the Gemini switch
Image first, config second -- the old binary on Gemini config fails
every tool-using run on its second model call, and the runbook says
why, what was verified, and how to roll back either half.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-22 13:16:46 +05:30

7.4 KiB

Deploying db4803c and switching the model vendor to Gemini

Two things land together, and the order matters: the image must be running before the configuration switches vendor. Old binary on Gemini config = every tool-using run dies on its second model call (§2). New binary on Groq config = works exactly as today. So: image first, config second.

Live today: Groq free tier, 8000 TPM, 35% of runs since Sep 9 end in gateway.rate_limited. That number is why this deploy exists.


1. What is being deployed

db4803c, on origin/main. Five commits since 8e36faf:

Commit What
822b3b1 .env untracked again; ignore rules 3455ad0 deleted are back
b765495 Archiving an agent a published agent delegates to → 409
5166fde GatewayFailure termination; migration 000016
797ee5f Provider metadata round-tripped on tool calls — Gemini needs this
db4803c An unsaved trajectory is logged, not just noted in itself

One migration. 000016 widens agent_runs.termination_check to admit GatewayFailure. The migrate init container applies it before the API starts. Its down migration is verified (folds rows to ToolFailure before narrowing the CHECK), so a rollback of the image is safe.

2. Why the image must go first

Gemini 3 models attach a thought_signature to every function call and reject the follow-up request without it. 797ee5f teaches the gateway to carry it back. The image on the cluster today does not, and it was proved on 2026-09-22: pointed at Gemini, the first model call succeeded, the tool ran, the second call answered 400 Function call is missing a thought_signature. Rolled back to Groq within minutes.

3. What was verified before writing this

With the 797ee5f binary, locally, against the production database and the real Gemini API through a request-logging proxy:

Check Result
positions-agent, 2 model calls Completed, signature present on the echoed call
krow-workforce-agent, 3 tool-calling turns, 12,123 tokens Completed, correct answer. This exceeds Groq's entire per-minute ceiling
Streaming path carries the signature unit test + live run
GatewayFailure reaches the surface with its own wording seen live before rollback
Full Go suite against a real Postgres (the DB tests skip without one) 18/18 packages
Migration 000016 up → down → up on a scratch database clean

Model reliability, six bare calls each, 2026-09-22 ~13:00 IST:

Model HTTP codes
gemini-3.8-flash 503 503 200 200 503 503
gemini-3.5-flash 503 200 503 200 200 200
gemini-3.5-flash-lite 200 200 200 200 200 200

The gateway retries a 503 three times with backoff; at the rates above a three-call run on either larger model still fails often. All three tiers run gemini-3.5-flash-lite until the larger models stop shedding load or the key is on a paid tier. gemini-3.1-pro-preview answers 429 (pro is not on the free tier); the 2.5 family is listed but blocked for new keys.

4. Build and push the image (the other machine)

The Dockerfile cross-compiles, so any host with buildx and a Docker Hub login works:

git checkout db4803c
docker buildx build --platform linux/amd64 \
  -f infrastructure/Dockerfile.api \
  -t doormile/krowbackend:db4803c -t doormile/krowbackend:latest \
  --push .

Two tags on purpose: :latest is what the StatefulSet pulls; :db4803c is what you roll back to if you need to (§7). Confirm before touching the cluster:

docker buildx imagetools inspect doormile/krowbackend:latest | grep -E 'Platform|Digest' | head -3

5. Switch the cluster (on the server, as root)

The Gemini key is already staged in the Secret as MODEL_API_KEY_GEMINI; the Groq key stays as MODEL_API_KEY_GROQ. The manifests in /opt/kubernetes/manifests/krow already describe the Gemini configuration (committed pending), so the config half is an apply.

# 1. the credential the API reads becomes the Gemini one
kubectl -n krow patch secret krow-model --type=json \
  -p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]'

# 2. configmap → Gemini base URL and model ids
kubectl apply -k /opt/kubernetes/manifests/krow/

# 3. new pods: pull :latest, run migration 16, boot on the new config
kubectl -n krow rollout restart statefulset/krow
kubectl -n krow rollout status statefulset/krow --timeout=5m

rollout status waits for krow-2, then krow-1, each gated on readiness. If krow-2 does not come up, krow-1 is still serving on the old image and Groq.

6. Smoke test

# migration 16 applied?
kubectl -n krow logs krow-2 -c migrate | tail -2        # want: 16/u gateway_failure_termination

# boot on the right vendor?
kubectl -n krow logs krow-2 -c api | grep -m1 '"listening"' | grep -o '"endpoints":[0-9]*'   # want 71

# a real run, through the public URL
T=$(curl -s -D - -o /dev/null -X POST https://mcp.krowforce.com/api/v1/auth/login \
  -H 'Content-Type: application/json' \
  -d '{"email":"demo@krow.app","password":"<demo password>"}' \
  | sed -n 's/^[Ss]et-[Cc]ookie: krow_session=\([^;]*\).*/\1/p')
curl -s -X POST https://mcp.krowforce.com/api/v1/agents/65bfd77d-2f74-4548-ab52-4e720e153397/runs \
  -H "Cookie: krow_session=$T" -H 'Content-Type: application/json' \
  -d '{"input":"How many open positions are there?"}' | grep -E '"(termination|output)"'

Want "termination": "Completed" and a count. GatewayFailure with "usually it is busy" is Gemini shedding load — retry once. ToolFailure on the second model call means the old image is still running (§2).

Then check the trajectory landed with the right model:

# from a machine with psql / the postgres image; DATABASE_URL from secret/krow-db
psql "$DATABASE_URL" -Atc "SELECT model, termination FROM agent_runs ORDER BY started_at DESC LIMIT 1"

Want gemini-3.5-flash-lite | Completed.

7. Rollback

Config only (image stays — it works on Groq too):

kubectl -n krow patch secret krow-model --type=json \
  -p '[{"op":"copy","from":"/data/MODEL_API_KEY_GROQ","path":"/data/MODEL_API_KEY"}]'
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
  "MODEL_BASE_URL":"https://api.groq.com/openai/v1",
  "MODEL_FAST":"openai/gpt-oss-20b","MODEL_BALANCED":"openai/gpt-oss-120b","MODEL_DEEP":"openai/gpt-oss-120b"}}'
kubectl -n krow rollout restart statefulset/krow

Image too (only if db4803c itself misbehaves):

kubectl -n krow set image statefulset/krow api=doormile/krowbackend:<previous tag>

Migration 16 stays applied; the old binary never writes GatewayFailure, so the wider CHECK is harmless to it. Reverse it only if you must: migrate ... down 1 — it folds existing GatewayFailure rows to ToolFailure.

8. Not covered here

  • activity-agent is archived while krow-workforce-agent v2 delegates to it. b765495 prevents this happening again; it does not repair the existing case. Either unarchive activity-agent or publish workforce v3 without it — a product decision.
  • OAUTH_LOGIN_PATH (/login) 404s on mcp.krowforce.com. A signed-out MCP consent redirect goes nowhere. Signed-in users are unaffected.
  • Rotation. The Anthropic key in 3455ad0's history, the Groq key, the Gemini key (pasted in a chat), the DB admin password (8 chars, public IP, no TLS), the root SSH password, the demo login.