diff --git a/docs/architecture-rag-and-memory.md b/docs/architecture-rag-and-memory.md index 5b97808..6e3eac8 100644 --- a/docs/architecture-rag-and-memory.md +++ b/docs/architecture-rag-and-memory.md @@ -1,6 +1,7 @@ # Retrieval and memory, as actually built -Written 2026-10-07, from a read of the code rather than from intent. It records +Written 2026-10-07 and updated the same day when the memory store landed, from +a read of the code rather than from intent. It records what is there, what is deliberately absent, and the reasoning for each — so the next person does not have to re-derive it, and so nobody claims more than the system does. @@ -13,7 +14,7 @@ system does. | Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents | | Working memory (within one answer) | **built** | | Conversation memory (across turns) | **built**, browser-side, token-budgeted | -| Long-term / semantic memory | **not built**, deliberately | +| Long-term / semantic memory | **store built and migrated; not wired** | ## 1. The model layer @@ -84,20 +85,60 @@ exchange drops the one the follow-up is about. The transcript is fenced and labelled as data: an earlier answer is the model's own words, but an earlier QUESTION is the reader's, and a reader can type anything. -**Long-term and semantic memory are absent, and that is a decision.** The -schema says so: +**Long-term memory: the store exists, the behaviour does not.** + +`internal/memory` and migration `000017_agent_memories` are built and applied +in production. Nothing in `loop.go` reads or writes a memory yet, so no answer +has ever been shaped by one. State it that way: the foundation is deployed, the +feature is not switched on. "We have long-term memory" is not yet true of +anything a user would experience. + +What the store is, and why it is mostly provenance: + +- **Org-scoped.** A memory written while one recruiter worked is available to + the next, because a workspace's view of its own hiring should not reset per + seat. `org_id` is NOT NULL and the predicate is in every read. +- **It remembers both kinds** — operational facts and observations about named + people. The second is why the rest of this list exists. "This applicant + seemed unreliable", written automatically and read into a later hiring + answer, is profiling under GDPR and is the artefact an employment claim would + be built on. +- **Every personal memory names its subject**, refused at the door rather than + defaulted. A memory about somebody that names nobody cannot be shown to them + or erased for them, which is the single property that makes holding it + defensible. `Held()` answers a subject access request and `Forget()` an + erasure, each in one statement; `Forget` is a soft delete, so the erasure is + itself on the record. +- **Provenance on every row** — the author (an agent's inference or a person's + note) and the run that wrote it, so "why did it say that" survives memory + entering the picture, and an inference is never read back as though a person + had written it. +- **Everything expires**, ninety days by default. A hiring workspace changes + shape over a quarter, and a stale fact read as a current one is worse than no + memory at all. +- **A memory cannot decide.** The block reaching the model is fenced and + labelled on the same terms as a retrieved document, states the origin of each + line, and says plainly that a memory is never a reason on its own to accept + or reject anybody. That sentence is pinned by a test, so an edit cannot + quietly drop it. + +Recall is semantic where an embedder exists and newest-first where it does not, +and says which happened rather than silently returning recency. Five memories +by default: this competes for the same prompt as the tool catalogue and the +retrieved block, against a ceiling of 8,000 tokens a minute. + +**What is left, and it is the hard part.** Wiring the read into `loop.go` is +small. The write trigger is not: automatic means something judges what is worth +remembering, and a bad judge fills the table with noise that then shapes every +answer after it. That decision is open. + +**Conversation threading is still absent**, separately, and the schema still +says why: > `conversations` — a run is one turn. Threading runs into a conversation is a > surface-layer concern and no surface asks for it yet; adding the column later > is trivial, and inventing the semantics now is not. -Nothing is learned across runs. `agent_runs` is an audit record and is never -replayed into a prompt. In a hiring product that is the conservative choice: -memory that influences a hiring answer is retained profiling of named -candidates, and it has to be auditable and explainable before it touches a -record. The retrieval and ACL layers it would reuse already exist, so building -it is a focused piece of work rather than a rebuild. - ## 4. What to verify before trusting retrieval in a demo Configuration being present does not mean a corpus was ingested: @@ -114,7 +155,15 @@ SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks; Hybrid retrieval with permission pre-filtering, multi-provider failover, per-run budgets, a full trajectory per answer, human confirmation before any -write, conversation memory with a token budget, and a bilingual interface. +write, conversation memory with a token budget, a bilingual interface, and a +long-term memory store that is built, migrated and auditable by subject. -It does not support "complete memory" or "everything is retrieval-backed". -Both are false, and the second is false on purpose. +It does NOT support: + +- "the agent remembers across sessions" — the store is not wired, so nothing + does yet; +- "complete memory"; +- "everything is retrieval-backed" — two of nine agents are, on purpose. + +The first of those is the one most likely to be said by accident, because the +code and the table both exist. Built is not the same as switched on.