Switch memory on: the agent may keep one, and carries them into the next run
The store and the write trigger existed and nothing used either. This wires both ends, so "we have long-term memory" stops being a statement about code that exists and becomes one about behaviour. READ. Memories are recalled before retrieval and placed before it in the prompt: it is the smaller block and the more general one — a standing preference frames how the documents should be read, where a document does not frame a preference. The question stays last, because a model reads the last thing and answers it, and evidence after the question becomes the prompt. Skipped for smalltalk on the same terms as retrieval. Nobody needs remembering to say good morning, and paying for it is how "hi" came to cost six thousand tokens. FAILS QUIET, RECORDED LOUDLY. A memory store that is unreachable must not take the run with it: an answer without memory is worse, not wrong, and the alternative is an outage in the knowledge layer becoming an outage in the product. The trajectory records the failure, and records separately when the store returned recency instead of relevance — a reader comparing two answers needs to know which one got which. WRITE. tools.Remember is registered only where there is somewhere to put it, through DefaultToolsWithMemory rather than a nil check inside the old constructor: a deployment that has not migrated 000017 must not offer a tool whose every call fails against a table that is not there. Passing nil registers exactly the catalogue that was there before. Memory shares retrieval's embedder, and memory.Embedder is knowledge.Embedder by structure so it cannot be given a different one. Two embedding models in one deployment produce vectors that cannot be compared, and the failure is silent: a recall that returns nothing rather than an error. Five new tests on the loop, including the two that matter — a failing store still answers, and a greeting carries nothing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -51,6 +51,11 @@ type ModelExecutor struct {
|
||||
// agent built so far answers from the operational tables through tools, and
|
||||
// none of them needs a corpus.
|
||||
retriever Retriever
|
||||
|
||||
// memory is the long-term store, or nil for a deployment that does not
|
||||
// remember. Nil is a supported state: the loop behaves exactly as it did
|
||||
// before memory existed.
|
||||
memory Recaller
|
||||
}
|
||||
|
||||
// Retriever is what the loop needs from the knowledge layer.
|
||||
@@ -333,19 +338,37 @@ func (m *ModelExecutor) executeRun(
|
||||
// never on the question, so without this a greeting was handed eight policy
|
||||
// chunks AHEAD of the word "hi" — which is both the dominant cost of the
|
||||
// turn and the reason the answer came back as an operational briefing.
|
||||
conversation := []gateway.Message{{Role: gateway.RoleUser, Text: question}}
|
||||
/* Evidence blocks, in the order the model should meet them: what this
|
||||
workspace remembered, then what the corpus says, then the question.
|
||||
Both are fenced and labelled as data; neither is an instruction. */
|
||||
var blocks []string
|
||||
|
||||
// Memory, before retrieval. It is the smaller block and the more general
|
||||
// one — a standing preference frames how the documents should be read, and
|
||||
// a reader meeting it first is not being told a conclusion, only a
|
||||
// context. Skipped for smalltalk on the same terms as retrieval: a
|
||||
// greeting does not need remembering, and paying for it is how "hi" came
|
||||
// to cost six thousand tokens.
|
||||
if !smalltalk {
|
||||
if block := m.recall(runCtx, rec, input, question); block != "" {
|
||||
blocks = append(blocks, block)
|
||||
}
|
||||
}
|
||||
|
||||
if !smalltalk {
|
||||
if block, retrieved := m.retrieve(runCtx, rec, agent, input, question); block != "" {
|
||||
conversation = []gateway.Message{{
|
||||
Role: gateway.RoleUser,
|
||||
// Context first, question second. A model reads the question last
|
||||
// and answers it, rather than treating the evidence as the prompt.
|
||||
Text: block + "\n\n" + question,
|
||||
}}
|
||||
blocks = append(blocks, block)
|
||||
rec.Retrieval(retrieved)
|
||||
}
|
||||
}
|
||||
|
||||
// Context first, question second. A model reads the question last and
|
||||
// answers it, rather than treating the evidence as the prompt.
|
||||
conversation := []gateway.Message{{
|
||||
Role: gateway.RoleUser,
|
||||
Text: strings.TrimSpace(strings.Join(append(blocks, question), "\n\n")),
|
||||
}}
|
||||
|
||||
// Recorded once rather than on each remaining step, so a long run does not
|
||||
// fill its trajectory with the same note.
|
||||
var toolsWithheld bool
|
||||
|
||||
Reference in New Issue
Block a user