Add the deploy runbook for db4803c and the Gemini switch
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 10s

Image first, config second -- the old binary on Gemini config fails
every tool-using run on its second model call, and the runbook says
why, what was verified, and how to roll back either half.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
This commit is contained in:
2026-09-22 13:16:46 +05:30
parent db4803c557
commit 939598a187

176
docs/deploy-db4803c.md Normal file
View File

@@ -0,0 +1,176 @@
# Deploying `db4803c` and switching the model vendor to Gemini
Two things land together, and the order matters: the image **must** be
running before the configuration switches vendor. Old binary on Gemini
config = every tool-using run dies on its second model call (§2). New binary
on Groq config = works exactly as today. So: image first, config second.
Live today: Groq free tier, `8000 TPM`, **35% of runs since Sep 9 end in
`gateway.rate_limited`**. That number is why this deploy exists.
---
## 1. What is being deployed
`db4803c`, on `origin/main`. Five commits since `8e36faf`:
| Commit | What |
| --- | --- |
| `822b3b1` | `.env` untracked again; ignore rules `3455ad0` deleted are back |
| `b765495` | Archiving an agent a published agent delegates to → 409 |
| `5166fde` | `GatewayFailure` termination; **migration 000016** |
| `797ee5f` | Provider metadata round-tripped on tool calls — **Gemini needs this** |
| `db4803c` | An unsaved trajectory is logged, not just noted in itself |
**One migration.** `000016` widens `agent_runs.termination_check` to admit
`GatewayFailure`. The `migrate` init container applies it before the API
starts. Its down migration is verified (folds rows to `ToolFailure` before
narrowing the CHECK), so a rollback of the image is safe.
## 2. Why the image must go first
Gemini 3 models attach a `thought_signature` to every function call and
reject the follow-up request without it. `797ee5f` teaches the gateway to
carry it back. The image on the cluster today does not, and it was proved on
2026-09-22: pointed at Gemini, the first model call succeeded, the tool ran,
the second call answered `400 Function call is missing a thought_signature`.
Rolled back to Groq within minutes.
## 3. What was verified before writing this
With the `797ee5f` binary, locally, against the production database and the
real Gemini API through a request-logging proxy:
| Check | Result |
| --- | --- |
| `positions-agent`, 2 model calls | `Completed`, signature present on the echoed call |
| `krow-workforce-agent`, 3 tool-calling turns, **12,123 tokens** | `Completed`, correct answer. This exceeds Groq's entire per-minute ceiling |
| Streaming path carries the signature | unit test + live run |
| `GatewayFailure` reaches the surface with its own wording | seen live before rollback |
| Full Go suite against a real Postgres (the DB tests skip without one) | 18/18 packages |
| Migration 000016 up → down → up on a scratch database | clean |
Model reliability, six bare calls each, 2026-09-22 ~13:00 IST:
| Model | HTTP codes |
| --- | --- |
| `gemini-3.8-flash` | 503 503 200 200 503 503 |
| `gemini-3.5-flash` | 503 200 503 200 200 200 |
| `gemini-3.5-flash-lite` | 200 200 200 200 200 200 |
The gateway retries a 503 three times with backoff; at the rates above a
three-call run on either larger model still fails often. **All three tiers
run `gemini-3.5-flash-lite`** until the larger models stop shedding load or
the key is on a paid tier. `gemini-3.1-pro-preview` answers 429 (pro is not
on the free tier); the 2.5 family is listed but blocked for new keys.
## 4. Build and push the image (the other machine)
The Dockerfile cross-compiles, so any host with `buildx` and a Docker Hub
login works:
```bash
git checkout db4803c
docker buildx build --platform linux/amd64 \
-f infrastructure/Dockerfile.api \
-t doormile/krowbackend:db4803c -t doormile/krowbackend:latest \
--push .
```
Two tags on purpose: `:latest` is what the StatefulSet pulls; `:db4803c` is
what you roll back **to** if you need to (§7). Confirm before touching the
cluster:
```bash
docker buildx imagetools inspect doormile/krowbackend:latest | grep -E 'Platform|Digest' | head -3
```
## 5. Switch the cluster (on the server, as root)
The Gemini key is already staged in the Secret as `MODEL_API_KEY_GEMINI`;
the Groq key stays as `MODEL_API_KEY_GROQ`. The manifests in
`/opt/kubernetes/manifests/krow` already describe the Gemini configuration
(committed `pending`), so the config half is an `apply`.
```bash
# 1. the credential the API reads becomes the Gemini one
kubectl -n krow patch secret krow-model --type=json \
-p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]'
# 2. configmap → Gemini base URL and model ids
kubectl apply -k /opt/kubernetes/manifests/krow/
# 3. new pods: pull :latest, run migration 16, boot on the new config
kubectl -n krow rollout restart statefulset/krow
kubectl -n krow rollout status statefulset/krow --timeout=5m
```
`rollout status` waits for krow-2, then krow-1, each gated on readiness. If
krow-2 does not come up, krow-1 is still serving on the old image and Groq.
## 6. Smoke test
```bash
# migration 16 applied?
kubectl -n krow logs krow-2 -c migrate | tail -2 # want: 16/u gateway_failure_termination
# boot on the right vendor?
kubectl -n krow logs krow-2 -c api | grep -m1 '"listening"' | grep -o '"endpoints":[0-9]*' # want 71
# a real run, through the public URL
T=$(curl -s -D - -o /dev/null -X POST https://mcp.krowforce.com/api/v1/auth/login \
-H 'Content-Type: application/json' \
-d '{"email":"demo@krow.app","password":"<demo password>"}' \
| sed -n 's/^[Ss]et-[Cc]ookie: krow_session=\([^;]*\).*/\1/p')
curl -s -X POST https://mcp.krowforce.com/api/v1/agents/65bfd77d-2f74-4548-ab52-4e720e153397/runs \
-H "Cookie: krow_session=$T" -H 'Content-Type: application/json' \
-d '{"input":"How many open positions are there?"}' | grep -E '"(termination|output)"'
```
Want `"termination": "Completed"` and a count. `GatewayFailure` with
"usually it is busy" is Gemini shedding load — retry once. `ToolFailure` on
the **second** model call means the old image is still running (§2).
Then check the trajectory landed with the right model:
```bash
# from a machine with psql / the postgres image; DATABASE_URL from secret/krow-db
psql "$DATABASE_URL" -Atc "SELECT model, termination FROM agent_runs ORDER BY started_at DESC LIMIT 1"
```
Want `gemini-3.5-flash-lite | Completed`.
## 7. Rollback
Config only (image stays — it works on Groq too):
```bash
kubectl -n krow patch secret krow-model --type=json \
-p '[{"op":"copy","from":"/data/MODEL_API_KEY_GROQ","path":"/data/MODEL_API_KEY"}]'
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
"MODEL_BASE_URL":"https://api.groq.com/openai/v1",
"MODEL_FAST":"openai/gpt-oss-20b","MODEL_BALANCED":"openai/gpt-oss-120b","MODEL_DEEP":"openai/gpt-oss-120b"}}'
kubectl -n krow rollout restart statefulset/krow
```
Image too (only if `db4803c` itself misbehaves):
```bash
kubectl -n krow set image statefulset/krow api=doormile/krowbackend:<previous tag>
```
Migration 16 stays applied; the old binary never writes `GatewayFailure`,
so the wider CHECK is harmless to it. Reverse it only if you must:
`migrate ... down 1` — it folds existing `GatewayFailure` rows to `ToolFailure`.
## 8. Not covered here
- **`activity-agent` is archived while `krow-workforce-agent v2` delegates to
it.** `b765495` prevents this happening again; it does not repair the
existing case. Either unarchive `activity-agent` or publish workforce v3
without it — a product decision.
- **`OAUTH_LOGIN_PATH` (`/login`) 404s on `mcp.krowforce.com`.** A signed-out
MCP consent redirect goes nowhere. Signed-in users are unaffected.
- **Rotation.** The Anthropic key in `3455ad0`'s history, the Groq key, the
Gemini key (pasted in a chat), the DB admin password (8 chars, public IP,
no TLS), the root SSH password, the demo login.