The ids always travelled TO the model; nothing a reader could look at ever
travelled back. So the panel stripped the citations the model wrote — there was
nowhere to put them — and a grounded answer became indistinguishable from an
invented one, which is the opposite of what citing is for.
A run now carries its sources: the id the model was told to cite, the document
title, the heading, and the opening of the passage. Both the JSON and the
streamed paths return them, because both build the same response.
Source is NOT knowledge.Result. That type carries ranks, scores and the whole
chunk, which exist to debug a retrieval rather than to be shown: an RRF score
is a rank and would be read as a percentage, and the full text would make the
response larger than the answer. This is the subset a citation needs.
Snippets are capped at 240 runes. A reader checking "the policy says X" needs
to recognise the passage, not to receive the corpus one answer at a time.
A run that retrieved nothing carries nothing rather than an empty list, so the
panel has no heading to draw for absent evidence — which is most runs, since
seven of nine agents answer from tools.
Next, and only now possible: the panel can render these and stop stripping the
citations that point at them.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The store and the write trigger existed and nothing used either. This wires
both ends, so "we have long-term memory" stops being a statement about code
that exists and becomes one about behaviour.
READ. Memories are recalled before retrieval and placed before it in the
prompt: it is the smaller block and the more general one — a standing
preference frames how the documents should be read, where a document does not
frame a preference. The question stays last, because a model reads the last
thing and answers it, and evidence after the question becomes the prompt.
Skipped for smalltalk on the same terms as retrieval. Nobody needs remembering
to say good morning, and paying for it is how "hi" came to cost six thousand
tokens.
FAILS QUIET, RECORDED LOUDLY. A memory store that is unreachable must not take
the run with it: an answer without memory is worse, not wrong, and the
alternative is an outage in the knowledge layer becoming an outage in the
product. The trajectory records the failure, and records separately when the
store returned recency instead of relevance — a reader comparing two answers
needs to know which one got which.
WRITE. tools.Remember is registered only where there is somewhere to put it,
through DefaultToolsWithMemory rather than a nil check inside the old
constructor: a deployment that has not migrated 000017 must not offer a tool
whose every call fails against a table that is not there. Passing nil
registers exactly the catalogue that was there before.
Memory shares retrieval's embedder, and memory.Embedder is knowledge.Embedder
by structure so it cannot be given a different one. Two embedding models in
one deployment produce vectors that cannot be compared, and the failure is
silent: a recall that returns nothing rather than an error.
Five new tests on the loop, including the two that matter — a failing store
still answers, and a greeting carries nothing.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>