diff --git a/docs/architecture-rag-and-memory.md b/docs/architecture-rag-and-memory.md index 92c6965..4a4d45c 100644 --- a/docs/architecture-rag-and-memory.md +++ b/docs/architecture-rag-and-memory.md @@ -1,7 +1,9 @@ # Retrieval and memory, as actually built -Written 2026-10-07 and updated the same day when the memory store landed, from -a read of the code rather than from intent. It records +Written 2026-10-07 and updated the same day as memory went from absent to +deployed, from a read of the code rather than from intent. Where it says +something is live, that was checked against the running binary in production, +not against a test run. It records what is there, what is deliberately absent, and the reasoning for each — so the next person does not have to re-derive it, and so nobody claims more than the system does. @@ -14,7 +16,7 @@ system does. | Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents | | Working memory (within one answer) | **built** | | Conversation memory (across turns) | **built**, browser-side, token-budgeted | -| Long-term / semantic memory | **store built and migrated; Wiring in Progress** | +| Long-term / semantic memory | **built, wired and deployed** — empty until used | ## 1. The model layer @@ -85,13 +87,15 @@ exchange drops the one the follow-up is about. The transcript is fenced and labelled as data: an earlier answer is the model's own words, but an earlier QUESTION is the reader's, and a reader can type anything. -**Long-term memory: the store exists, the behaviour does not.** +**Long-term memory is live.** `internal/memory`, migration `000017`, the +`remember` tool and the recall path in `loop.go` are all deployed and verified +in the running binary. -`internal/memory` and migration `000017_agent_memories` are built and applied -in production. Nothing in `loop.go` reads or writes a memory yet, so no answer -has ever been shaped by one. State it that way: the foundation is deployed, the -feature is not switched on. "We have long-term memory" is not yet true of -anything a user would experience. +It is EMPTY until a workspace uses it. Memory accumulates from what agents are +told; it does not arrive populated, and a memory only enters a prompt on a +LATER run — so the conversation that creates one shows no difference. That is +the design, not a fault, and it is the thing most likely to be mistaken for the +feature not working. What the store is, and why it is mostly provenance: @@ -127,10 +131,36 @@ and says which happened rather than silently returning recency. Five memories by default: this competes for the same prompt as the tool catalogue and the retrieved block, against a ceiling of 8,000 tokens a minute. -**What is left, and it is the hard part.** Wiring the read into `loop.go` is -small. The write trigger is not: automatic means something judges what is worth -remembering, and a bad judge fills the table with noise that then shapes every -answer after it. That decision is open. +**The write trigger: a tool, not an extraction pass.** Three designs were +available. A second model call after each run judges well and costs a whole +extra call against a ceiling of 8,000 tokens a minute, on every run, most of +which have nothing worth keeping. A heuristic in the loop is cheap and +remembers the wrong things, because the shape of a run says nothing about +whether a fact outlives it. So the agent gets a `remember` tool: it costs +nothing extra, it is automatic in the sense that matters — nobody types +"remember this" — and it is visible in the trajectory, which an extraction pass +would not be. + +**It is a confirmed write**, because `EffectWrite` forces it and that is the +invariant working rather than an obstacle: this stores personal data that will +shape later hiring answers. A person sees the sentence, who it is about, that +an agent and not a person decided it, and when it expires. If workspace facts +should later be kept without asking, the honest change is a SECOND tool scoped +to workspace subjects — loosening this one would quietly make personal +memories unconfirmed too. + +**The read path.** Memories are recalled before retrieval and placed before it: +a standing preference frames how documents should be read, where a document +does not frame a preference. The question stays last, because a model reads the +last thing and answers it. Skipped for smalltalk on the same terms as +retrieval — nobody needs remembering to say good morning, and paying for it is +how "hi" came to cost six thousand tokens. + +**It fails quiet and is recorded loudly.** A memory store that is unreachable +does not take the run with it: an answer without memory is worse, not wrong, +and the alternative is an outage in the knowledge layer becoming an outage in +the product. The trajectory records the failure, and records separately when +the store returned recency instead of relevance. **Conversation threading is still absent**, separately, and the schema still says why: @@ -160,10 +190,13 @@ long-term memory store that is built, migrated and auditable by subject. It does NOT support: -- "the agent remembers across sessions" — the store is not wired, so nothing - does yet; -- "complete memory"; +- "it already knows your workspace" — memory is live and starts empty. It + accumulates from use, and on a fresh deployment there is nothing in it; +- "it remembers everything" — five memories per run, ninety-day expiry, and a + person confirms each one before it is kept; +- "complete memory" — conversation threading is still a browser-side window, + not a server-side thread; - "everything is retrieval-backed" — two of nine agents are, on purpose. -The first of those is the one most likely to be said by accident, because the -code and the table both exist. Built is not the same as switched on. +The first is now the one most likely to be said by accident. The feature works; +the table is empty until somebody uses it.