The first store whose contents are deliberately fed back into a prompt, which
makes it a different kind of table from everything around it. Org-scoped by
product decision: a memory written while one recruiter worked is available to
the next, because a workspace's view of its own hiring should not reset per
seat.
It remembers both kinds asked for — operational facts and observations about
named people — and the second is why most of this code is provenance rather
than payload. "This applicant seemed unreliable", stored automatically and
read into a later hiring answer, is profiling under GDPR and is the artefact an
employment claim is built on. The only thing that makes holding it defensible
is that it can be listed, shown and erased, so:
- subject_type and subject_id are mandatory for anything personal, refused at
the door rather than defaulted, because a memory about somebody that names
nobody cannot be shown to them or deleted for them;
- Held() answers a subject access request and Forget() answers an erasure,
each in one statement, and Forget is a soft delete so the erasure itself is
recorded;
- every memory carries its author and the run that wrote it, so "why did it
say that" survives memory entering the picture, and an inference is never
read back as if a person had written it;
- everything expires. Ninety days by default: a hiring workspace changes
shape over a quarter, and a stale fact read as a current one is worse than
no memory at all.
The block the model sees is fenced and labelled on the same terms as retrieved
documents, for a stronger reason — a memory is text this system wrote about its
own users, so a model that treated it as an instruction would let one run steer
every run after it. It states the origin of each line and says plainly that a
memory is never a reason on its own to accept or reject anybody. That sentence
is pinned by a test.
Recall is semantic where an embedder exists and newest-first where it does not,
and says which happened rather than quietly returning recency. Five memories by
default: this competes for the same prompt as the tool catalogue and the
retrieved block, against a ceiling of 8,000 tokens a minute.
Migration 000017 is WRITTEN AND NOT APPLIED. Nothing is wired into the runtime
yet — this is the store and its rules, reviewable on its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
4.3 KiB
Moving the chat model in-cluster: qwen3:0.6b on the existing Ollama
Replaces the Gemini free tier as the gateway's provider. No credential, no
per-token cost, no rate limit, and no tenant text leaving the cluster. The
gateway needs no code change — routing.go speaks one wire shape and Ollama
serves it, so this is a base URL and three model ids.
1. What changes
| Piece | Before | After |
|---|---|---|
MODEL_BASE_URL |
https://generativelanguage.googleapis.com/v1beta/openai |
http://ollama.krow.svc.cluster.local:11434/v1 |
MODEL_FAST/BALANCED/DEEP |
gemini-3.5-flash-lite |
qwen3:0.6b |
MODEL_API_KEY |
the Gemini key | ollama (any non-empty string) |
infrastructure/ollama.yaml |
1 loaded model, 1Gi request | 2 loaded models, 1536Mi request |
MODEL_API_KEY cannot be empty: config.validateModel requires it when
APP_ENV=production (config.go:514), and the agent routes are not registered
at all without it. Ollama ignores the value.
Leave MODEL_REASONING_EFFORT unset. Ollama rejects an unknown
reasoning_effort key with a 400, which Error.Retryable() correctly does not
retry — every run would die on gateway.invalid_request.
2. Why qwen3:0.6b
~500MB at Q4 and it ships a tools template, which is the whole requirement:
the agents are multi-turn tool callers over eight-tool catalogues, and a model
with no tool template cannot call one at all. gemma3:270m is smaller and has
no tool template — it is not a candidate. llama3.2:1b is the next step up
(~1.3GB) if 0.6b cannot hold a tool call together.
3. Pull the model
kubectl -n krow exec deploy/ollama -- ollama pull qwen3:0.6b
kubectl -n krow exec deploy/ollama -- ollama list # want qwen3:0.6b and nomic-embed-text
4. Apply the manifest, then the config
Manifest first — the memory headroom has to exist before two models are resident, or the kubelet kills the pod mid-pull.
kubectl apply -f infrastructure/ollama.yaml
kubectl -n krow rollout status deploy/ollama --timeout=5m
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
"MODEL_BASE_URL":"http://ollama.krow.svc.cluster.local:11434/v1",
"MODEL_FAST":"qwen3:0.6b","MODEL_BALANCED":"qwen3:0.6b","MODEL_DEEP":"qwen3:0.6b",
"MODEL_MAX_OUTPUT_TOKENS":"2000"}}'
kubectl -n krow patch secret krow-model --type=json \
-p '[{"op":"replace","path":"/stringData/MODEL_API_KEY","value":"ollama"}]'
kubectl -n krow rollout restart statefulset/krow
kubectl -n krow rollout status statefulset/krow --timeout=5m
MODEL_MAX_OUTPUT_TOKENS drops from the 16000 default: on CPU every output
token is wall clock, and a run that generates 16k of them dies on its deadline
instead of answering.
5. Verify — the part that decides this
The suites are the instrument. 9 agents, 5 cases each, run against the real
agents/*.md specs and the registry's actual tools:
MODEL_PROVIDER=openai \
MODEL_BASE_URL=http://ollama.krow.svc.cluster.local:11434/v1 \
MODEL_API_KEY=ollama \
MODEL_FAST=qwen3:0.6b MODEL_BALANCED=qwen3:0.6b MODEL_DEEP=qwen3:0.6b \
make eval-live
Then one real run through the public URL, the §6 smoke test from
deploy-db4803c.md. Want "termination": "Completed".
Watch for, in order of likelihood:
| Symptom | Meaning |
|---|---|
ToolFailure on call 1 |
the model invented a tool or emitted a malformed call — terminationFor classifies this as the tool layer's, but it is the model |
Deadline |
generation too slow on CPU. Lower MODEL_MAX_OUTPUT_TOKENS further, or step up the node |
gateway.invalid_request |
MODEL_REASONING_EFFORT is set, or the model id is not pulled |
| API pods restarting | Ollama took the node. Lower its limit; ollama.yaml's original comment is the warning |
6. Rollback
Config only — no image change in this deploy:
kubectl -n krow patch secret krow-model --type=json \
-p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]'
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
"MODEL_BASE_URL":"https://generativelanguage.googleapis.com/v1beta/openai",
"MODEL_FAST":"gemini-3.5-flash-lite","MODEL_BALANCED":"gemini-3.5-flash-lite",
"MODEL_DEEP":"gemini-3.5-flash-lite","MODEL_MAX_OUTPUT_TOKENS":"16000"}}'
kubectl -n krow rollout restart statefulset/krow