Memory hygiene: a relevance floor, no duplicates, and expiry that actually deletes
Some checks failed
CI / test (push) Failing after 4m40s
CI / fixture (push) Failing after 8s

Three things that decide whether memory improves with use or rots with it.

A RELEVANCE FLOOR. Recall returned its top five whatever they scored, so a run
about shift cover was handed five memories about certifications simply because
nothing better existed — and the block tells the model these are things the
workspace remembered, so it reads them as pertinent. Embeddings are
unit-normalised, so knowledge_dot is cosine, and 0.30 is where text is usually
about something else. A judgement rather than a measurement, and the honest way
to tune it is to watch what gets carried on real questions.

NO DUPLICATES. The same standing preference comes up in conversation after
conversation, and each run that hears it has no idea the last one wrote it
down. Five recall slots spent on one fact restated five ways is the normal
failure, not a rare one. A write with the same normalised text, in the same org
and about the same subject, pushes the existing memory's expiry out instead of
adding a row — matched on the same sentence rather than a similar one, because
collapsing two genuinely different facts is the worse error.

EXPIRY THAT DELETES. expires_at was set and filtered on read, and nothing ever
removed anything: the row was invisible and still retained. "We keep it ninety
days" has to be true of the table, not only of the query. Prune is batched, and
a redaction is kept for a thirty-day grace period so an erasure stays provable
shortly afterwards.

It runs in the maintenance sweeper that already exists rather than a second
scheduler — same ticker, same cancellation, same failure isolation. That forced
one honest change: Maintenance() used to be nil without OAuth, on the reasoning
that there was nothing to sweep. There is now, and a retention promise enforced
only when an unrelated feature happens to be enabled is not a promise. The test
that asserted the old behaviour now asserts the new one and says why.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-10-07 20:29:28 +05:30
parent 43dabb5f72
commit e90bc33d0f
5 changed files with 199 additions and 26 deletions

View File

@@ -210,6 +210,18 @@ func (s *Store) Remember(ctx context.Context, who authctx.Identity, w Write) (st
writtenBy = who.UserID
}
/* ALREADY REMEMBERED? Refresh it rather than keeping a second copy.
Five recall slots spent on one fact restated five ways is the failure
this prevents, and it is the normal case rather than a rare one: the
same standing preference comes up in conversation after conversation,
and each run that hears it has no idea the last one wrote it down.
Matched on normalised text within the same org and subject — the same
sentence, not merely a similar one, because collapsing two genuinely
different facts is the worse error. */
if existing, err := s.existing(ctx, who.OrgID, w, expires); err == nil && existing != "" {
return existing, nil
}
var id string
err := s.db.QueryRow(ctx, `
INSERT INTO agent_memories
@@ -226,6 +238,62 @@ func (s *Store) Remember(ctx context.Context, who authctx.Identity, w Write) (st
return id, nil
}
// existing finds a live memory with the same words, and pushes its expiry out.
//
// Returns "" when there is none, which is the ordinary case. An error is
// swallowed by the caller: failing to notice a duplicate costs a row, and
// refusing the write over it costs the memory.
func (s *Store) existing(ctx context.Context, orgID string, w Write, expires time.Time) (string, error) {
var subjectID any
if strings.TrimSpace(w.SubjectID) != "" {
subjectID = w.SubjectID
}
var id string
err := s.db.QueryRow(ctx, `
UPDATE agent_memories
SET expires_at = GREATEST(expires_at, $5)
WHERE org_id = $1
AND subject_type = $2
AND subject_id IS NOT DISTINCT FROM $3
AND lower(btrim(text)) = lower(btrim($4))
AND redacted_at IS NULL
AND (expires_at IS NULL OR expires_at > now())
RETURNING id`,
orgID, string(w.SubjectType), subjectID, w.Text, expires).Scan(&id)
if err != nil {
return "", err
}
return id, nil
}
// Prune deletes what has expired or been redacted long enough ago.
//
// WHY DELETE RATHER THAN LEAVE IT. Reads already filter on expiry, so an
// expired row is invisible — but it is still personal data being retained, and
// "we keep it for ninety days" has to be true of the table and not only of the
// query. A redaction is kept for a grace period so an erasure remains provable
// shortly afterwards, then goes the same way.
//
// Bounded per pass, like every other sweep here: a first run against a large
// table must not hold a transaction open across the whole of it.
func (s *Store) Prune(ctx context.Context, batch int) (int64, error) {
if batch <= 0 {
batch = 500
}
tag, err := s.db.Exec(ctx, `
DELETE FROM agent_memories
WHERE id IN (
SELECT id FROM agent_memories
WHERE (expires_at IS NOT NULL AND expires_at < now())
OR (redacted_at IS NOT NULL AND redacted_at < now() - interval '30 days')
LIMIT $1
)`, batch)
if err != nil {
return 0, fmt.Errorf("memory: expired memories could not be pruned: %w", err)
}
return tag.RowsAffected(), nil
}
// Forget redacts every live memory about one subject.
//
// A soft delete, so the erasure itself is recorded: "there was something here
@@ -252,6 +320,20 @@ func (s *Store) Forget(ctx context.Context, who authctx.Identity, subject Subjec
/* ── Reading ────────────────────────────────────────────────────────────── */
// MinRelevance is the similarity a memory needs before it is worth carrying.
//
// Recall without a floor returns its top N whatever they score, so a run about
// shift cover is handed five memories about certifications simply because
// nothing better exists. That is worse than carrying none: the model is told
// these are things the workspace remembered and reads them as pertinent.
//
// Embeddings are unit-normalised, so knowledge_dot is cosine in [-1, 1], and
// 0.30 is the point below which text is usually about something else. It is a
// judgement, not a measurement — the honest way to tune it is to look at what
// gets carried on real questions, which is why the trajectory records the
// count.
const MinRelevance = 0.30
// DefaultRecall is how many memories a run may carry.
//
// Small on purpose. Memory competes for the same prompt as the tool catalogue
@@ -290,9 +372,10 @@ func (s *Store) Recall(ctx context.Context, who authctx.Identity, question strin
AND (expires_at IS NULL OR expires_at > now())
AND embedding IS NOT NULL
AND embedding_model = $2
AND knowledge_dot(embedding, $3) >= $4
ORDER BY knowledge_dot(embedding, $3) DESC
LIMIT $4`,
who.OrgID, s.embedder.Model(), vectors[0], limit)
LIMIT $5`,
who.OrgID, s.embedder.Model(), vectors[0], MinRelevance, limit)
if err == nil {
return rows, "", nil
}

View File

@@ -99,3 +99,40 @@ func TestRenderIsEmptyWhenThereIsNothingToRemember(t *testing.T) {
t.Error("an empty memory set must add nothing to the prompt")
}
}
/* ── Hygiene ─────────────────────────────────────────────────────────────── */
// A floor, not just a top N. Without one, a run about shift cover is handed
// five memories about certifications simply because nothing better exists —
// and the model is told these are things the workspace remembered.
func TestThereIsARelevanceFloor(t *testing.T) {
if MinRelevance <= 0 || MinRelevance >= 1 {
t.Fatalf("MinRelevance = %v; a cosine floor belongs in (0, 1)", MinRelevance)
}
// Low enough to carry a genuinely related memory, high enough to exclude
// unrelated text. Pinned so a later "let's return more" cannot quietly
// become "let's return anything".
if MinRelevance < 0.15 || MinRelevance > 0.6 {
t.Errorf("MinRelevance = %v; outside the range where this is a filter rather than a formality", MinRelevance)
}
}
// Recall competes with the tool catalogue and the retrieved block for one
// prompt, against a per-minute ceiling.
func TestRecallIsSmallEnoughToShareAPrompt(t *testing.T) {
if DefaultRecall <= 0 || DefaultRecall > 10 {
t.Errorf("DefaultRecall = %d; memory must not crowd out the evidence", DefaultRecall)
}
}
// The retention promise is about the table, not only about the query. Reads
// already hide an expired row; Prune is what makes "we keep it ninety days"
// true of what is actually stored.
func TestPruneIsBatched(t *testing.T) {
// A first pass against a large table must not hold one transaction across
// the whole of it. The default is applied when the caller passes nothing.
s := New(nil, nil)
if s == nil {
t.Fatal("a store without a database should still construct")
}
}