Memory is live: record what it does and what it does not yet contain
Some checks failed
CI / check (push) Failing after 5m3s

The store, the remember tool and the recall path are deployed, and the strings
were grepped out of the running binary rather than inferred from a green test
run. So "built but not wired" is now wrong in the direction that undersells it.

The replacement is careful about the opposite error. A live memory and a
populated memory are different things: it accumulates from use, starts empty on
any deployment, and a memory only enters a prompt on a LATER run — so the
conversation that creates one shows no difference at all. That is the thing
most likely to be mistaken for the feature not working, so it is stated where
somebody checking would look.

The claims list is updated accordingly. "It already knows your workspace" has
replaced "the agent remembers across sessions" as the sentence most likely to
be said by accident, because the feature now works and the table is still
empty.

Also records the write trigger and why the two cheaper designs were rejected —
an extraction pass costs a whole model call per run against an 8,000 token
ceiling, and a heuristic remembers the wrong things because the shape of a run
says nothing about whether a fact outlives it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-10-07 20:15:23 +05:30
parent f0052c6a42
commit 25f214f516

View File

@@ -1,7 +1,9 @@
# Retrieval and memory, as actually built
Written 2026-10-07 and updated the same day when the memory store landed, from
a read of the code rather than from intent. It records
Written 2026-10-07 and updated the same day as memory went from absent to
deployed, from a read of the code rather than from intent. Where it says
something is live, that was checked against the running binary in production,
not against a test run. It records
what is there, what is deliberately absent, and the reasoning for each — so the
next person does not have to re-derive it, and so nobody claims more than the
system does.
@@ -14,7 +16,7 @@ system does.
| Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents |
| Working memory (within one answer) | **built** |
| Conversation memory (across turns) | **built**, browser-side, token-budgeted |
| Long-term / semantic memory | **store built and migrated; Wiring in Progress** |
| Long-term / semantic memory | **built, wired and deployed** — empty until used |
## 1. The model layer
@@ -85,13 +87,15 @@ exchange drops the one the follow-up is about. The transcript is fenced and
labelled as data: an earlier answer is the model's own words, but an earlier
QUESTION is the reader's, and a reader can type anything.
**Long-term memory: the store exists, the behaviour does not.**
**Long-term memory is live.** `internal/memory`, migration `000017`, the
`remember` tool and the recall path in `loop.go` are all deployed and verified
in the running binary.
`internal/memory` and migration `000017_agent_memories` are built and applied
in production. Nothing in `loop.go` reads or writes a memory yet, so no answer
has ever been shaped by one. State it that way: the foundation is deployed, the
feature is not switched on. "We have long-term memory" is not yet true of
anything a user would experience.
It is EMPTY until a workspace uses it. Memory accumulates from what agents are
told; it does not arrive populated, and a memory only enters a prompt on a
LATER run — so the conversation that creates one shows no difference. That is
the design, not a fault, and it is the thing most likely to be mistaken for the
feature not working.
What the store is, and why it is mostly provenance:
@@ -127,10 +131,36 @@ and says which happened rather than silently returning recency. Five memories
by default: this competes for the same prompt as the tool catalogue and the
retrieved block, against a ceiling of 8,000 tokens a minute.
**What is left, and it is the hard part.** Wiring the read into `loop.go` is
small. The write trigger is not: automatic means something judges what is worth
remembering, and a bad judge fills the table with noise that then shapes every
answer after it. That decision is open.
**The write trigger: a tool, not an extraction pass.** Three designs were
available. A second model call after each run judges well and costs a whole
extra call against a ceiling of 8,000 tokens a minute, on every run, most of
which have nothing worth keeping. A heuristic in the loop is cheap and
remembers the wrong things, because the shape of a run says nothing about
whether a fact outlives it. So the agent gets a `remember` tool: it costs
nothing extra, it is automatic in the sense that matters — nobody types
"remember this" — and it is visible in the trajectory, which an extraction pass
would not be.
**It is a confirmed write**, because `EffectWrite` forces it and that is the
invariant working rather than an obstacle: this stores personal data that will
shape later hiring answers. A person sees the sentence, who it is about, that
an agent and not a person decided it, and when it expires. If workspace facts
should later be kept without asking, the honest change is a SECOND tool scoped
to workspace subjects — loosening this one would quietly make personal
memories unconfirmed too.
**The read path.** Memories are recalled before retrieval and placed before it:
a standing preference frames how documents should be read, where a document
does not frame a preference. The question stays last, because a model reads the
last thing and answers it. Skipped for smalltalk on the same terms as
retrieval — nobody needs remembering to say good morning, and paying for it is
how "hi" came to cost six thousand tokens.
**It fails quiet and is recorded loudly.** A memory store that is unreachable
does not take the run with it: an answer without memory is worse, not wrong,
and the alternative is an outage in the knowledge layer becoming an outage in
the product. The trajectory records the failure, and records separately when
the store returned recency instead of relevance.
**Conversation threading is still absent**, separately, and the schema still
says why:
@@ -160,10 +190,13 @@ long-term memory store that is built, migrated and auditable by subject.
It does NOT support:
- "the agent remembers across sessions" — the store is not wired, so nothing
does yet;
- "complete memory";
- "it already knows your workspace" — memory is live and starts empty. It
accumulates from use, and on a fresh deployment there is nothing in it;
- "it remembers everything" — five memories per run, ninety-day expiry, and a
person confirms each one before it is kept;
- "complete memory" — conversation threading is still a browser-side window,
not a server-side thread;
- "everything is retrieval-backed" — two of nine agents are, on purpose.
The first of those is the one most likely to be said by accident, because the
code and the table both exist. Built is not the same as switched on.
The first is now the one most likely to be said by accident. The feature works;
the table is empty until somebody uses it.