# Deploying `db4803c` and switching the model vendor to Gemini Two things land together, and the order matters: the image **must** be running before the configuration switches vendor. Old binary on Gemini config = every tool-using run dies on its second model call (§2). New binary on Groq config = works exactly as today. So: image first, config second. Live today: Groq free tier, `8000 TPM`, **35% of runs since Sep 9 end in `gateway.rate_limited`**. That number is why this deploy exists. --- ## 1. What is being deployed `db4803c`, on `origin/main`. Five commits since `8e36faf`: | Commit | What | | --- | --- | | `822b3b1` | `.env` untracked again; ignore rules `3455ad0` deleted are back | | `b765495` | Archiving an agent a published agent delegates to → 409 | | `5166fde` | `GatewayFailure` termination; **migration 000016** | | `797ee5f` | Provider metadata round-tripped on tool calls — **Gemini needs this** | | `db4803c` | An unsaved trajectory is logged, not just noted in itself | **One migration.** `000016` widens `agent_runs.termination_check` to admit `GatewayFailure`. The `migrate` init container applies it before the API starts. Its down migration is verified (folds rows to `ToolFailure` before narrowing the CHECK), so a rollback of the image is safe. ## 2. Why the image must go first Gemini 3 models attach a `thought_signature` to every function call and reject the follow-up request without it. `797ee5f` teaches the gateway to carry it back. The image on the cluster today does not, and it was proved on 2026-09-22: pointed at Gemini, the first model call succeeded, the tool ran, the second call answered `400 Function call is missing a thought_signature`. Rolled back to Groq within minutes. ## 3. What was verified before writing this With the `797ee5f` binary, locally, against the production database and the real Gemini API through a request-logging proxy: | Check | Result | | --- | --- | | `positions-agent`, 2 model calls | `Completed`, signature present on the echoed call | | `krow-workforce-agent`, 3 tool-calling turns, **12,123 tokens** | `Completed`, correct answer. This exceeds Groq's entire per-minute ceiling | | Streaming path carries the signature | unit test + live run | | `GatewayFailure` reaches the surface with its own wording | seen live before rollback | | Full Go suite against a real Postgres (the DB tests skip without one) | 18/18 packages | | Migration 000016 up → down → up on a scratch database | clean | Model reliability, six bare calls each, 2026-09-22 ~13:00 IST: | Model | HTTP codes | | --- | --- | | `gemini-3.8-flash` | 503 503 200 200 503 503 | | `gemini-3.5-flash` | 503 200 503 200 200 200 | | `gemini-3.5-flash-lite` | 200 200 200 200 200 200 | The gateway retries a 503 three times with backoff; at the rates above a three-call run on either larger model still fails often. **All three tiers run `gemini-3.5-flash-lite`** until the larger models stop shedding load or the key is on a paid tier. `gemini-3.1-pro-preview` answers 429 (pro is not on the free tier); the 2.5 family is listed but blocked for new keys. ## 4. Build and push the image (the other machine) The Dockerfile cross-compiles, so any host with `buildx` and a Docker Hub login works: ```bash git checkout db4803c docker buildx build --platform linux/amd64 \ -f infrastructure/Dockerfile.api \ -t doormile/krowbackend:db4803c -t doormile/krowbackend:latest \ --push . ``` Two tags on purpose: `:latest` is what the StatefulSet pulls; `:db4803c` is what you roll back **to** if you need to (§7). Confirm before touching the cluster: ```bash docker buildx imagetools inspect doormile/krowbackend:latest | grep -E 'Platform|Digest' | head -3 ``` ## 5. Switch the cluster (on the server, as root) The Gemini key is already staged in the Secret as `MODEL_API_KEY_GEMINI`; the Groq key stays as `MODEL_API_KEY_GROQ`. The manifests in `/opt/kubernetes/manifests/krow` already describe the Gemini configuration (committed `pending`), so the config half is an `apply`. ```bash # 1. the credential the API reads becomes the Gemini one kubectl -n krow patch secret krow-model --type=json \ -p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]' # 2. configmap → Gemini base URL and model ids kubectl apply -k /opt/kubernetes/manifests/krow/ # 3. new pods: pull :latest, run migration 16, boot on the new config kubectl -n krow rollout restart statefulset/krow kubectl -n krow rollout status statefulset/krow --timeout=5m ``` `rollout status` waits for krow-2, then krow-1, each gated on readiness. If krow-2 does not come up, krow-1 is still serving on the old image and Groq. ## 6. Smoke test ```bash # migration 16 applied? kubectl -n krow logs krow-2 -c migrate | tail -2 # want: 16/u gateway_failure_termination # boot on the right vendor? kubectl -n krow logs krow-2 -c api | grep -m1 '"listening"' | grep -o '"endpoints":[0-9]*' # want 71 # a real run, through the public URL T=$(curl -s -D - -o /dev/null -X POST https://mcp.krowforce.com/api/v1/auth/login \ -H 'Content-Type: application/json' \ -d '{"email":"demo@krow.app","password":""}' \ | sed -n 's/^[Ss]et-[Cc]ookie: krow_session=\([^;]*\).*/\1/p') curl -s -X POST https://mcp.krowforce.com/api/v1/agents/65bfd77d-2f74-4548-ab52-4e720e153397/runs \ -H "Cookie: krow_session=$T" -H 'Content-Type: application/json' \ -d '{"input":"How many open positions are there?"}' | grep -E '"(termination|output)"' ``` Want `"termination": "Completed"` and a count. `GatewayFailure` with "usually it is busy" is Gemini shedding load — retry once. `ToolFailure` on the **second** model call means the old image is still running (§2). Then check the trajectory landed with the right model: ```bash # from a machine with psql / the postgres image; DATABASE_URL from secret/krow-db psql "$DATABASE_URL" -Atc "SELECT model, termination FROM agent_runs ORDER BY started_at DESC LIMIT 1" ``` Want `gemini-3.5-flash-lite | Completed`. ## 7. Rollback Config only (image stays — it works on Groq too): ```bash kubectl -n krow patch secret krow-model --type=json \ -p '[{"op":"copy","from":"/data/MODEL_API_KEY_GROQ","path":"/data/MODEL_API_KEY"}]' kubectl -n krow patch cm krow-config --type merge -p '{"data":{ "MODEL_BASE_URL":"https://api.groq.com/openai/v1", "MODEL_FAST":"openai/gpt-oss-20b","MODEL_BALANCED":"openai/gpt-oss-120b","MODEL_DEEP":"openai/gpt-oss-120b"}}' kubectl -n krow rollout restart statefulset/krow ``` Image too (only if `db4803c` itself misbehaves): ```bash kubectl -n krow set image statefulset/krow api=doormile/krowbackend: ``` Migration 16 stays applied; the old binary never writes `GatewayFailure`, so the wider CHECK is harmless to it. Reverse it only if you must: `migrate ... down 1` — it folds existing `GatewayFailure` rows to `ToolFailure`. ## 8. Not covered here - **`activity-agent` is archived while `krow-workforce-agent v2` delegates to it.** `b765495` prevents this happening again; it does not repair the existing case. Either unarchive `activity-agent` or publish workforce v3 without it — a product decision. - **`OAUTH_LOGIN_PATH` (`/login`) 404s on `mcp.krowforce.com`.** A signed-out MCP consent redirect goes nowhere. Signed-in users are unaffected. - **Rotation.** The Anthropic key in `3455ad0`'s history, the Groq key, the Gemini key (pasted in a chat), the DB admin password (8 chars, public IP, no TLS), the root SSH password, the demo login.