Memory is live: record what it does and what it does not yet contain
Some checks failed
CI / check (push) Failing after 5m3s
Some checks failed
CI / check (push) Failing after 5m3s
The store, the remember tool and the recall path are deployed, and the strings were grepped out of the running binary rather than inferred from a green test run. So "built but not wired" is now wrong in the direction that undersells it. The replacement is careful about the opposite error. A live memory and a populated memory are different things: it accumulates from use, starts empty on any deployment, and a memory only enters a prompt on a LATER run — so the conversation that creates one shows no difference at all. That is the thing most likely to be mistaken for the feature not working, so it is stated where somebody checking would look. The claims list is updated accordingly. "It already knows your workspace" has replaced "the agent remembers across sessions" as the sentence most likely to be said by accident, because the feature now works and the table is still empty. Also records the write trigger and why the two cheaper designs were rejected — an extraction pass costs a whole model call per run against an 8,000 token ceiling, and a heuristic remembers the wrong things because the shape of a run says nothing about whether a fact outlives it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -1,7 +1,9 @@
|
||||
# Retrieval and memory, as actually built
|
||||
|
||||
Written 2026-10-07 and updated the same day when the memory store landed, from
|
||||
a read of the code rather than from intent. It records
|
||||
Written 2026-10-07 and updated the same day as memory went from absent to
|
||||
deployed, from a read of the code rather than from intent. Where it says
|
||||
something is live, that was checked against the running binary in production,
|
||||
not against a test run. It records
|
||||
what is there, what is deliberately absent, and the reasoning for each — so the
|
||||
next person does not have to re-derive it, and so nobody claims more than the
|
||||
system does.
|
||||
@@ -14,7 +16,7 @@ system does.
|
||||
| Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents |
|
||||
| Working memory (within one answer) | **built** |
|
||||
| Conversation memory (across turns) | **built**, browser-side, token-budgeted |
|
||||
| Long-term / semantic memory | **store built and migrated; Wiring in Progress** |
|
||||
| Long-term / semantic memory | **built, wired and deployed** — empty until used |
|
||||
|
||||
## 1. The model layer
|
||||
|
||||
@@ -85,13 +87,15 @@ exchange drops the one the follow-up is about. The transcript is fenced and
|
||||
labelled as data: an earlier answer is the model's own words, but an earlier
|
||||
QUESTION is the reader's, and a reader can type anything.
|
||||
|
||||
**Long-term memory: the store exists, the behaviour does not.**
|
||||
**Long-term memory is live.** `internal/memory`, migration `000017`, the
|
||||
`remember` tool and the recall path in `loop.go` are all deployed and verified
|
||||
in the running binary.
|
||||
|
||||
`internal/memory` and migration `000017_agent_memories` are built and applied
|
||||
in production. Nothing in `loop.go` reads or writes a memory yet, so no answer
|
||||
has ever been shaped by one. State it that way: the foundation is deployed, the
|
||||
feature is not switched on. "We have long-term memory" is not yet true of
|
||||
anything a user would experience.
|
||||
It is EMPTY until a workspace uses it. Memory accumulates from what agents are
|
||||
told; it does not arrive populated, and a memory only enters a prompt on a
|
||||
LATER run — so the conversation that creates one shows no difference. That is
|
||||
the design, not a fault, and it is the thing most likely to be mistaken for the
|
||||
feature not working.
|
||||
|
||||
What the store is, and why it is mostly provenance:
|
||||
|
||||
@@ -127,10 +131,36 @@ and says which happened rather than silently returning recency. Five memories
|
||||
by default: this competes for the same prompt as the tool catalogue and the
|
||||
retrieved block, against a ceiling of 8,000 tokens a minute.
|
||||
|
||||
**What is left, and it is the hard part.** Wiring the read into `loop.go` is
|
||||
small. The write trigger is not: automatic means something judges what is worth
|
||||
remembering, and a bad judge fills the table with noise that then shapes every
|
||||
answer after it. That decision is open.
|
||||
**The write trigger: a tool, not an extraction pass.** Three designs were
|
||||
available. A second model call after each run judges well and costs a whole
|
||||
extra call against a ceiling of 8,000 tokens a minute, on every run, most of
|
||||
which have nothing worth keeping. A heuristic in the loop is cheap and
|
||||
remembers the wrong things, because the shape of a run says nothing about
|
||||
whether a fact outlives it. So the agent gets a `remember` tool: it costs
|
||||
nothing extra, it is automatic in the sense that matters — nobody types
|
||||
"remember this" — and it is visible in the trajectory, which an extraction pass
|
||||
would not be.
|
||||
|
||||
**It is a confirmed write**, because `EffectWrite` forces it and that is the
|
||||
invariant working rather than an obstacle: this stores personal data that will
|
||||
shape later hiring answers. A person sees the sentence, who it is about, that
|
||||
an agent and not a person decided it, and when it expires. If workspace facts
|
||||
should later be kept without asking, the honest change is a SECOND tool scoped
|
||||
to workspace subjects — loosening this one would quietly make personal
|
||||
memories unconfirmed too.
|
||||
|
||||
**The read path.** Memories are recalled before retrieval and placed before it:
|
||||
a standing preference frames how documents should be read, where a document
|
||||
does not frame a preference. The question stays last, because a model reads the
|
||||
last thing and answers it. Skipped for smalltalk on the same terms as
|
||||
retrieval — nobody needs remembering to say good morning, and paying for it is
|
||||
how "hi" came to cost six thousand tokens.
|
||||
|
||||
**It fails quiet and is recorded loudly.** A memory store that is unreachable
|
||||
does not take the run with it: an answer without memory is worse, not wrong,
|
||||
and the alternative is an outage in the knowledge layer becoming an outage in
|
||||
the product. The trajectory records the failure, and records separately when
|
||||
the store returned recency instead of relevance.
|
||||
|
||||
**Conversation threading is still absent**, separately, and the schema still
|
||||
says why:
|
||||
@@ -160,10 +190,13 @@ long-term memory store that is built, migrated and auditable by subject.
|
||||
|
||||
It does NOT support:
|
||||
|
||||
- "the agent remembers across sessions" — the store is not wired, so nothing
|
||||
does yet;
|
||||
- "complete memory";
|
||||
- "it already knows your workspace" — memory is live and starts empty. It
|
||||
accumulates from use, and on a fresh deployment there is nothing in it;
|
||||
- "it remembers everything" — five memories per run, ninety-day expiry, and a
|
||||
person confirms each one before it is kept;
|
||||
- "complete memory" — conversation threading is still a browser-side window,
|
||||
not a server-side thread;
|
||||
- "everything is retrieval-backed" — two of nine agents are, on purpose.
|
||||
|
||||
The first of those is the one most likely to be said by accident, because the
|
||||
code and the table both exist. Built is not the same as switched on.
|
||||
The first is now the one most likely to be said by accident. The feature works;
|
||||
the table is empty until somebody uses it.
|
||||
|
||||
Reference in New Issue
Block a user