Image first, config second -- the old binary on Gemini config fails every tool-using run on its second model call, and the runbook says why, what was verified, and how to roll back either half. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
7.4 KiB
Deploying db4803c and switching the model vendor to Gemini
Two things land together, and the order matters: the image must be running before the configuration switches vendor. Old binary on Gemini config = every tool-using run dies on its second model call (§2). New binary on Groq config = works exactly as today. So: image first, config second.
Live today: Groq free tier, 8000 TPM, 35% of runs since Sep 9 end in
gateway.rate_limited. That number is why this deploy exists.
1. What is being deployed
db4803c, on origin/main. Five commits since 8e36faf:
| Commit | What |
|---|---|
822b3b1 |
.env untracked again; ignore rules 3455ad0 deleted are back |
b765495 |
Archiving an agent a published agent delegates to → 409 |
5166fde |
GatewayFailure termination; migration 000016 |
797ee5f |
Provider metadata round-tripped on tool calls — Gemini needs this |
db4803c |
An unsaved trajectory is logged, not just noted in itself |
One migration. 000016 widens agent_runs.termination_check to admit
GatewayFailure. The migrate init container applies it before the API
starts. Its down migration is verified (folds rows to ToolFailure before
narrowing the CHECK), so a rollback of the image is safe.
2. Why the image must go first
Gemini 3 models attach a thought_signature to every function call and
reject the follow-up request without it. 797ee5f teaches the gateway to
carry it back. The image on the cluster today does not, and it was proved on
2026-09-22: pointed at Gemini, the first model call succeeded, the tool ran,
the second call answered 400 Function call is missing a thought_signature.
Rolled back to Groq within minutes.
3. What was verified before writing this
With the 797ee5f binary, locally, against the production database and the
real Gemini API through a request-logging proxy:
| Check | Result |
|---|---|
positions-agent, 2 model calls |
Completed, signature present on the echoed call |
krow-workforce-agent, 3 tool-calling turns, 12,123 tokens |
Completed, correct answer. This exceeds Groq's entire per-minute ceiling |
| Streaming path carries the signature | unit test + live run |
GatewayFailure reaches the surface with its own wording |
seen live before rollback |
| Full Go suite against a real Postgres (the DB tests skip without one) | 18/18 packages |
| Migration 000016 up → down → up on a scratch database | clean |
Model reliability, six bare calls each, 2026-09-22 ~13:00 IST:
| Model | HTTP codes |
|---|---|
gemini-3.8-flash |
503 503 200 200 503 503 |
gemini-3.5-flash |
503 200 503 200 200 200 |
gemini-3.5-flash-lite |
200 200 200 200 200 200 |
The gateway retries a 503 three times with backoff; at the rates above a
three-call run on either larger model still fails often. All three tiers
run gemini-3.5-flash-lite until the larger models stop shedding load or
the key is on a paid tier. gemini-3.1-pro-preview answers 429 (pro is not
on the free tier); the 2.5 family is listed but blocked for new keys.
4. Build and push the image (the other machine)
The Dockerfile cross-compiles, so any host with buildx and a Docker Hub
login works:
git checkout db4803c
docker buildx build --platform linux/amd64 \
-f infrastructure/Dockerfile.api \
-t doormile/krowbackend:db4803c -t doormile/krowbackend:latest \
--push .
Two tags on purpose: :latest is what the StatefulSet pulls; :db4803c is
what you roll back to if you need to (§7). Confirm before touching the
cluster:
docker buildx imagetools inspect doormile/krowbackend:latest | grep -E 'Platform|Digest' | head -3
5. Switch the cluster (on the server, as root)
The Gemini key is already staged in the Secret as MODEL_API_KEY_GEMINI;
the Groq key stays as MODEL_API_KEY_GROQ. The manifests in
/opt/kubernetes/manifests/krow already describe the Gemini configuration
(committed pending), so the config half is an apply.
# 1. the credential the API reads becomes the Gemini one
kubectl -n krow patch secret krow-model --type=json \
-p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]'
# 2. configmap → Gemini base URL and model ids
kubectl apply -k /opt/kubernetes/manifests/krow/
# 3. new pods: pull :latest, run migration 16, boot on the new config
kubectl -n krow rollout restart statefulset/krow
kubectl -n krow rollout status statefulset/krow --timeout=5m
rollout status waits for krow-2, then krow-1, each gated on readiness. If
krow-2 does not come up, krow-1 is still serving on the old image and Groq.
6. Smoke test
# migration 16 applied?
kubectl -n krow logs krow-2 -c migrate | tail -2 # want: 16/u gateway_failure_termination
# boot on the right vendor?
kubectl -n krow logs krow-2 -c api | grep -m1 '"listening"' | grep -o '"endpoints":[0-9]*' # want 71
# a real run, through the public URL
T=$(curl -s -D - -o /dev/null -X POST https://mcp.krowforce.com/api/v1/auth/login \
-H 'Content-Type: application/json' \
-d '{"email":"demo@krow.app","password":"<demo password>"}' \
| sed -n 's/^[Ss]et-[Cc]ookie: krow_session=\([^;]*\).*/\1/p')
curl -s -X POST https://mcp.krowforce.com/api/v1/agents/65bfd77d-2f74-4548-ab52-4e720e153397/runs \
-H "Cookie: krow_session=$T" -H 'Content-Type: application/json' \
-d '{"input":"How many open positions are there?"}' | grep -E '"(termination|output)"'
Want "termination": "Completed" and a count. GatewayFailure with
"usually it is busy" is Gemini shedding load — retry once. ToolFailure on
the second model call means the old image is still running (§2).
Then check the trajectory landed with the right model:
# from a machine with psql / the postgres image; DATABASE_URL from secret/krow-db
psql "$DATABASE_URL" -Atc "SELECT model, termination FROM agent_runs ORDER BY started_at DESC LIMIT 1"
Want gemini-3.5-flash-lite | Completed.
7. Rollback
Config only (image stays — it works on Groq too):
kubectl -n krow patch secret krow-model --type=json \
-p '[{"op":"copy","from":"/data/MODEL_API_KEY_GROQ","path":"/data/MODEL_API_KEY"}]'
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
"MODEL_BASE_URL":"https://api.groq.com/openai/v1",
"MODEL_FAST":"openai/gpt-oss-20b","MODEL_BALANCED":"openai/gpt-oss-120b","MODEL_DEEP":"openai/gpt-oss-120b"}}'
kubectl -n krow rollout restart statefulset/krow
Image too (only if db4803c itself misbehaves):
kubectl -n krow set image statefulset/krow api=doormile/krowbackend:<previous tag>
Migration 16 stays applied; the old binary never writes GatewayFailure,
so the wider CHECK is harmless to it. Reverse it only if you must:
migrate ... down 1 — it folds existing GatewayFailure rows to ToolFailure.
8. Not covered here
activity-agentis archived whilekrow-workforce-agent v2delegates to it.b765495prevents this happening again; it does not repair the existing case. Either unarchiveactivity-agentor publish workforce v3 without it — a product decision.OAUTH_LOGIN_PATH(/login) 404s onmcp.krowforce.com. A signed-out MCP consent redirect goes nowhere. Signed-in users are unaffected.- Rotation. The Anthropic key in
3455ad0's history, the Groq key, the Gemini key (pasted in a chat), the DB admin password (8 chars, public IP, no TLS), the root SSH password, the demo login.