Three things that decide whether memory improves with use or rots with it. A RELEVANCE FLOOR. Recall returned its top five whatever they scored, so a run about shift cover was handed five memories about certifications simply because nothing better existed — and the block tells the model these are things the workspace remembered, so it reads them as pertinent. Embeddings are unit-normalised, so knowledge_dot is cosine, and 0.30 is where text is usually about something else. A judgement rather than a measurement, and the honest way to tune it is to watch what gets carried on real questions. NO DUPLICATES. The same standing preference comes up in conversation after conversation, and each run that hears it has no idea the last one wrote it down. Five recall slots spent on one fact restated five ways is the normal failure, not a rare one. A write with the same normalised text, in the same org and about the same subject, pushes the existing memory's expiry out instead of adding a row — matched on the same sentence rather than a similar one, because collapsing two genuinely different facts is the worse error. EXPIRY THAT DELETES. expires_at was set and filtered on read, and nothing ever removed anything: the row was invisible and still retained. "We keep it ninety days" has to be true of the table, not only of the query. Prune is batched, and a redaction is kept for a thirty-day grace period so an erasure stays provable shortly afterwards. It runs in the maintenance sweeper that already exists rather than a second scheduler — same ticker, same cancellation, same failure isolation. That forced one honest change: Maintenance() used to be nil without OAuth, on the reasoning that there was nothing to sweep. There is now, and a retention promise enforced only when an unrelated feature happens to be enabled is not a promise. The test that asserted the old behaviour now asserts the new one and says why. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
17 KiB
17 KiB