Add long-term memory: org-scoped, attributed to a subject, and expiring
The first store whose contents are deliberately fed back into a prompt, which
makes it a different kind of table from everything around it. Org-scoped by
product decision: a memory written while one recruiter worked is available to
the next, because a workspace's view of its own hiring should not reset per
seat.
It remembers both kinds asked for — operational facts and observations about
named people — and the second is why most of this code is provenance rather
than payload. "This applicant seemed unreliable", stored automatically and
read into a later hiring answer, is profiling under GDPR and is the artefact an
employment claim is built on. The only thing that makes holding it defensible
is that it can be listed, shown and erased, so:
- subject_type and subject_id are mandatory for anything personal, refused at
the door rather than defaulted, because a memory about somebody that names
nobody cannot be shown to them or deleted for them;
- Held() answers a subject access request and Forget() answers an erasure,
each in one statement, and Forget is a soft delete so the erasure itself is
recorded;
- every memory carries its author and the run that wrote it, so "why did it
say that" survives memory entering the picture, and an inference is never
read back as if a person had written it;
- everything expires. Ninety days by default: a hiring workspace changes
shape over a quarter, and a stale fact read as a current one is worse than
no memory at all.
The block the model sees is fenced and labelled on the same terms as retrieved
documents, for a stronger reason — a memory is text this system wrote about its
own users, so a model that treated it as an instruction would let one run steer
every run after it. It states the origin of each line and says plainly that a
memory is never a reason on its own to accept or reject anybody. That sentence
is pinned by a test.
Recall is semantic where an embedder exists and newest-first where it does not,
and says which happened rather than quietly returning recency. Five memories by
default: this competes for the same prompt as the tool catalogue and the
retrieved block, against a ceiling of 8,000 tokens a minute.
Migration 000017 is WRITTEN AND NOT APPLIED. Nothing is wired into the runtime
yet — this is the store and its rules, reviewable on its own.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
101
docs/deploy-ollama-chat.md
Normal file
101
docs/deploy-ollama-chat.md
Normal file
@@ -0,0 +1,101 @@
|
||||
# Moving the chat model in-cluster: `qwen3:0.6b` on the existing Ollama
|
||||
|
||||
Replaces the Gemini free tier as the gateway's provider. No credential, no
|
||||
per-token cost, no rate limit, and no tenant text leaving the cluster. The
|
||||
gateway needs no code change — `routing.go` speaks one wire shape and Ollama
|
||||
serves it, so this is a base URL and three model ids.
|
||||
|
||||
## 1. What changes
|
||||
|
||||
| Piece | Before | After |
|
||||
| --- | --- | --- |
|
||||
| `MODEL_BASE_URL` | `https://generativelanguage.googleapis.com/v1beta/openai` | `http://ollama.krow.svc.cluster.local:11434/v1` |
|
||||
| `MODEL_FAST/BALANCED/DEEP` | `gemini-3.5-flash-lite` | `qwen3:0.6b` |
|
||||
| `MODEL_API_KEY` | the Gemini key | `ollama` (any non-empty string) |
|
||||
| `infrastructure/ollama.yaml` | 1 loaded model, 1Gi request | 2 loaded models, 1536Mi request |
|
||||
|
||||
`MODEL_API_KEY` cannot be empty: `config.validateModel` requires it when
|
||||
`APP_ENV=production` (`config.go:514`), and the agent routes are not registered
|
||||
at all without it. Ollama ignores the value.
|
||||
|
||||
Leave `MODEL_REASONING_EFFORT` unset. Ollama rejects an unknown
|
||||
`reasoning_effort` key with a 400, which `Error.Retryable()` correctly does not
|
||||
retry — every run would die on `gateway.invalid_request`.
|
||||
|
||||
## 2. Why `qwen3:0.6b`
|
||||
|
||||
~500MB at Q4 and it ships a **tools template**, which is the whole requirement:
|
||||
the agents are multi-turn tool callers over eight-tool catalogues, and a model
|
||||
with no tool template cannot call one at all. `gemma3:270m` is smaller and has
|
||||
no tool template — it is not a candidate. `llama3.2:1b` is the next step up
|
||||
(~1.3GB) if 0.6b cannot hold a tool call together.
|
||||
|
||||
## 3. Pull the model
|
||||
|
||||
```bash
|
||||
kubectl -n krow exec deploy/ollama -- ollama pull qwen3:0.6b
|
||||
kubectl -n krow exec deploy/ollama -- ollama list # want qwen3:0.6b and nomic-embed-text
|
||||
```
|
||||
|
||||
## 4. Apply the manifest, then the config
|
||||
|
||||
Manifest first — the memory headroom has to exist before two models are
|
||||
resident, or the kubelet kills the pod mid-pull.
|
||||
|
||||
```bash
|
||||
kubectl apply -f infrastructure/ollama.yaml
|
||||
kubectl -n krow rollout status deploy/ollama --timeout=5m
|
||||
|
||||
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
|
||||
"MODEL_BASE_URL":"http://ollama.krow.svc.cluster.local:11434/v1",
|
||||
"MODEL_FAST":"qwen3:0.6b","MODEL_BALANCED":"qwen3:0.6b","MODEL_DEEP":"qwen3:0.6b",
|
||||
"MODEL_MAX_OUTPUT_TOKENS":"2000"}}'
|
||||
kubectl -n krow patch secret krow-model --type=json \
|
||||
-p '[{"op":"replace","path":"/stringData/MODEL_API_KEY","value":"ollama"}]'
|
||||
|
||||
kubectl -n krow rollout restart statefulset/krow
|
||||
kubectl -n krow rollout status statefulset/krow --timeout=5m
|
||||
```
|
||||
|
||||
`MODEL_MAX_OUTPUT_TOKENS` drops from the 16000 default: on CPU every output
|
||||
token is wall clock, and a run that generates 16k of them dies on its deadline
|
||||
instead of answering.
|
||||
|
||||
## 5. Verify — the part that decides this
|
||||
|
||||
The suites are the instrument. 9 agents, 5 cases each, run against the real
|
||||
`agents/*.md` specs and the registry's actual tools:
|
||||
|
||||
```bash
|
||||
MODEL_PROVIDER=openai \
|
||||
MODEL_BASE_URL=http://ollama.krow.svc.cluster.local:11434/v1 \
|
||||
MODEL_API_KEY=ollama \
|
||||
MODEL_FAST=qwen3:0.6b MODEL_BALANCED=qwen3:0.6b MODEL_DEEP=qwen3:0.6b \
|
||||
make eval-live
|
||||
```
|
||||
|
||||
Then one real run through the public URL, the §6 smoke test from
|
||||
`deploy-db4803c.md`. Want `"termination": "Completed"`.
|
||||
|
||||
Watch for, in order of likelihood:
|
||||
|
||||
| Symptom | Meaning |
|
||||
| --- | --- |
|
||||
| `ToolFailure` on call 1 | the model invented a tool or emitted a malformed call — `terminationFor` classifies this as the tool layer's, but it is the model |
|
||||
| `Deadline` | generation too slow on CPU. Lower `MODEL_MAX_OUTPUT_TOKENS` further, or step up the node |
|
||||
| `gateway.invalid_request` | `MODEL_REASONING_EFFORT` is set, or the model id is not pulled |
|
||||
| API pods restarting | Ollama took the node. Lower its limit; `ollama.yaml`'s original comment is the warning |
|
||||
|
||||
## 6. Rollback
|
||||
|
||||
Config only — no image change in this deploy:
|
||||
|
||||
```bash
|
||||
kubectl -n krow patch secret krow-model --type=json \
|
||||
-p '[{"op":"copy","from":"/data/MODEL_API_KEY_GEMINI","path":"/data/MODEL_API_KEY"}]'
|
||||
kubectl -n krow patch cm krow-config --type merge -p '{"data":{
|
||||
"MODEL_BASE_URL":"https://generativelanguage.googleapis.com/v1beta/openai",
|
||||
"MODEL_FAST":"gemini-3.5-flash-lite","MODEL_BALANCED":"gemini-3.5-flash-lite",
|
||||
"MODEL_DEEP":"gemini-3.5-flash-lite","MODEL_MAX_OUTPUT_TOKENS":"16000"}}'
|
||||
kubectl -n krow rollout restart statefulset/krow
|
||||
```
|
||||
Reference in New Issue
Block a user