# Moving the chat model in-cluster: `qwen3:0.6b` on the existing Ollama Replaces the Gemini free tier as the gateway's provider. No credential, no per-token cost, no rate limit, and no tenant text leaving the cluster. The gateway needs no code change — `routing.go` speaks one wire shape and Ollama serves it, so this is a base URL and three model ids. ## 1. What changes | Piece | Before | After | | --- | --- | --- | | `MODEL_BASE_URL` | `https://generativelanguage.googleapis.com/v1beta/openai` | `http://ollama.krow.svc.cluster.local:11434/v1` | | `MODEL_FAST/BALANCED/DEEP` | `gemini-3.5-flash-lite` | `qwen3:0.6b` | | `MODEL_API_KEY` | the Gemini key | `ollama` (any non-empty string) | | `infrastructure/ollama.yaml` | 1 loaded model, 1Gi request | 2 loaded models, 1536Mi request | `MODEL_API_KEY` cannot be empty: `config.validateModel` requires it when `APP_ENV=production` (`config.go:514`), and the agent routes are not registered at all without it. Ollama ignores the value. Leave `MODEL_REASONING_EFFORT` unset. Ollama rejects an unknown `reasoning_effort` key with a 400, which `Error.Retryable()` correctly does not retry — every run would die on `gateway.invalid_request`. ## 2. Why `qwen3:0.6b` ~500MB at Q4 and it ships a **tools template**, which is the whole requirement: the agents are multi-turn tool callers over eight-tool catalogues, and a model with no tool template cannot call one at all. `gemma3:270m` is smaller and has no tool template — it is not a candidate. `llama3.2:1b` is the next step up (~1.3GB) if 0.6b cannot hold a tool call together. ## 3. Pull the model ```bash kubectl -n krow exec deploy/ollama -- ollama pull qwen3:0.6b kubectl -n krow exec deploy/ollama -- ollama list # want qwen3:0.6b and nomic-embed-text ``` ## 4. Apply the manifest, then the config Manifest first — the memory headroom has to exist before two models are resident, or the kubelet kills the pod mid-pull. ```bash kubectl apply -f infrastructure/ollama.yaml kubectl -n krow rollout status deploy/ollama --timeout=5m kubectl -n krow patch cm krow-config --type merge -p '{"data":{ "MODEL_BASE_URL":"http://ollama.krow.svc.cluster.local:11434/v1", "MODEL_FAST":"qwen3:0.6b","MODEL_BALANCED":"qwen3:0.6b","MODEL_DEEP":"qwen3:0.6b", "MODEL_MAX_OUTPUT_TOKENS":"2000"}}' kubectl -n krow patch secret krow-model --type=json \ -p '[{"op":"replace","path":"/stringData/MODEL_API_KEY","value":"ollama"}]' kubectl -n krow rollout restart statefulset/krow kubectl -n krow rollout status statefulset/krow --timeout=5m ``` `MODEL_MAX_OUTPUT_TOKENS` drops from the 16000 default: on CPU every output token is wall clock, and a run that generates 16k of them dies on its deadline instead of answering. ## 5. Verify — the part that decides this The suites are the instrument. 9 agents, 5 cases each, run against the real `agents/*.md` specs and the registry's actual tools: ```bash MODEL_PROVIDER=openai \ MODEL_BASE_URL=http://ollama.krow.svc.cluster.local:11434/v1 \ MODEL_API_KEY=ollama \ MODEL_FAST=qwen3:0.6b MODEL_BALANCED=qwen3:0.6b MODEL_DEEP=qwen3:0.6b \ make eval-live ``` Then one real run through the public URL, the §6 smoke test from `deploy-db4803c.md`. Want `"termination": "Completed"`. Watch for, in order of likelihood: | Symptom | Meaning | | --- | --- | | `ToolFailure` on call 1 | the model invented a tool or emitted a malformed call — `terminationFor` classifies this as the tool layer's, but it is the model | | `Deadline` | generation too slow on CPU. Lower `MODEL_MAX_OUTPUT_TOKENS` further, or step up the node | | `gateway.invalid_request` | `MODEL_REASONING_EFFORT` is set, or the model id is not pulled | | API pods restarting | Ollama took the node. Lower its limit; `ollama.yaml`'s original comment is the warning | ## 6. Rollback Config only — no image change in this deploy: ```bash kubectl -n krow patch secret krow-model --type=json \ -p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]' kubectl -n krow patch cm krow-config --type merge -p '{"data":{ "MODEL_BASE_URL":"https://generativelanguage.googleapis.com/v1beta/openai", "MODEL_FAST":"gemini-3.5-flash-lite","MODEL_BALANCED":"gemini-3.5-flash-lite", "MODEL_DEEP":"gemini-3.5-flash-lite","MODEL_MAX_OUTPUT_TOKENS":"16000"}}' kubectl -n krow rollout restart statefulset/krow ```