Files
krow_backend/docs/deploy-ollama-chat.md
Aravind 8c51c22c86 Add long-term memory: org-scoped, attributed to a subject, and expiring
The first store whose contents are deliberately fed back into a prompt, which
makes it a different kind of table from everything around it. Org-scoped by
product decision: a memory written while one recruiter worked is available to
the next, because a workspace's view of its own hiring should not reset per
seat.

It remembers both kinds asked for — operational facts and observations about
named people — and the second is why most of this code is provenance rather
than payload. "This applicant seemed unreliable", stored automatically and
read into a later hiring answer, is profiling under GDPR and is the artefact an
employment claim is built on. The only thing that makes holding it defensible
is that it can be listed, shown and erased, so:

  - subject_type and subject_id are mandatory for anything personal, refused at
    the door rather than defaulted, because a memory about somebody that names
    nobody cannot be shown to them or deleted for them;
  - Held() answers a subject access request and Forget() answers an erasure,
    each in one statement, and Forget is a soft delete so the erasure itself is
    recorded;
  - every memory carries its author and the run that wrote it, so "why did it
    say that" survives memory entering the picture, and an inference is never
    read back as if a person had written it;
  - everything expires. Ninety days by default: a hiring workspace changes
    shape over a quarter, and a stale fact read as a current one is worse than
    no memory at all.

The block the model sees is fenced and labelled on the same terms as retrieved
documents, for a stronger reason — a memory is text this system wrote about its
own users, so a model that treated it as an instruction would let one run steer
every run after it. It states the origin of each line and says plainly that a
memory is never a reason on its own to accept or reject anybody. That sentence
is pinned by a test.

Recall is semantic where an embedder exists and newest-first where it does not,
and says which happened rather than quietly returning recency. Five memories by
default: this competes for the same prompt as the tool catalogue and the
retrieved block, against a ceiling of 8,000 tokens a minute.

Migration 000017 is WRITTEN AND NOT APPLIED. Nothing is wired into the runtime
yet — this is the store and its rules, reviewable on its own.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-07 19:55:08 +05:30

4.3 KiB

Moving the chat model in-cluster: qwen3:0.6b on the existing Ollama

Replaces the Gemini free tier as the gateway's provider. No credential, no per-token cost, no rate limit, and no tenant text leaving the cluster. The gateway needs no code change — routing.go speaks one wire shape and Ollama serves it, so this is a base URL and three model ids.

1. What changes

Piece Before After
MODEL_BASE_URL https://generativelanguage.googleapis.com/v1beta/openai http://ollama.krow.svc.cluster.local:11434/v1
MODEL_FAST/BALANCED/DEEP gemini-3.5-flash-lite qwen3:0.6b
MODEL_API_KEY the Gemini key ollama (any non-empty string)
infrastructure/ollama.yaml 1 loaded model, 1Gi request 2 loaded models, 1536Mi request

MODEL_API_KEY cannot be empty: config.validateModel requires it when APP_ENV=production (config.go:514), and the agent routes are not registered at all without it. Ollama ignores the value.

Leave MODEL_REASONING_EFFORT unset. Ollama rejects an unknown reasoning_effort key with a 400, which Error.Retryable() correctly does not retry — every run would die on gateway.invalid_request.

2. Why qwen3:0.6b

~500MB at Q4 and it ships a tools template, which is the whole requirement: the agents are multi-turn tool callers over eight-tool catalogues, and a model with no tool template cannot call one at all. gemma3:270m is smaller and has no tool template — it is not a candidate. llama3.2:1b is the next step up (~1.3GB) if 0.6b cannot hold a tool call together.

3. Pull the model

kubectl -n krow exec deploy/ollama -- ollama pull qwen3:0.6b
kubectl -n krow exec deploy/ollama -- ollama list      # want qwen3:0.6b and nomic-embed-text

4. Apply the manifest, then the config

Manifest first — the memory headroom has to exist before two models are resident, or the kubelet kills the pod mid-pull.

kubectl apply -f infrastructure/ollama.yaml
kubectl -n krow rollout status deploy/ollama --timeout=5m

kubectl -n krow patch cm krow-config --type merge -p '{"data":{
  "MODEL_BASE_URL":"http://ollama.krow.svc.cluster.local:11434/v1",
  "MODEL_FAST":"qwen3:0.6b","MODEL_BALANCED":"qwen3:0.6b","MODEL_DEEP":"qwen3:0.6b",
  "MODEL_MAX_OUTPUT_TOKENS":"2000"}}'
kubectl -n krow patch secret krow-model --type=json \
  -p '[{"op":"replace","path":"/stringData/MODEL_API_KEY","value":"ollama"}]'

kubectl -n krow rollout restart statefulset/krow
kubectl -n krow rollout status statefulset/krow --timeout=5m

MODEL_MAX_OUTPUT_TOKENS drops from the 16000 default: on CPU every output token is wall clock, and a run that generates 16k of them dies on its deadline instead of answering.

5. Verify — the part that decides this

The suites are the instrument. 9 agents, 5 cases each, run against the real agents/*.md specs and the registry's actual tools:

MODEL_PROVIDER=openai \
MODEL_BASE_URL=http://ollama.krow.svc.cluster.local:11434/v1 \
MODEL_API_KEY=ollama \
MODEL_FAST=qwen3:0.6b MODEL_BALANCED=qwen3:0.6b MODEL_DEEP=qwen3:0.6b \
make eval-live

Then one real run through the public URL, the §6 smoke test from deploy-db4803c.md. Want "termination": "Completed".

Watch for, in order of likelihood:

Symptom Meaning
ToolFailure on call 1 the model invented a tool or emitted a malformed call — terminationFor classifies this as the tool layer's, but it is the model
Deadline generation too slow on CPU. Lower MODEL_MAX_OUTPUT_TOKENS further, or step up the node
gateway.invalid_request MODEL_REASONING_EFFORT is set, or the model id is not pulled
API pods restarting Ollama took the node. Lower its limit; ollama.yaml's original comment is the warning

6. Rollback

Config only — no image change in this deploy:

kubectl -n krow patch secret krow-model --type=json \
  -p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]'
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
  "MODEL_BASE_URL":"https://generativelanguage.googleapis.com/v1beta/openai",
  "MODEL_FAST":"gemini-3.5-flash-lite","MODEL_BALANCED":"gemini-3.5-flash-lite",
  "MODEL_DEEP":"gemini-3.5-flash-lite","MODEL_MAX_OUTPUT_TOKENS":"16000"}}'
kubectl -n krow rollout restart statefulset/krow