Memory hygiene: a relevance floor, no duplicates, and expiry that actually deletes
Some checks failed
CI / test (push) Failing after 4m40s
CI / fixture (push) Failing after 8s

Three things that decide whether memory improves with use or rots with it.

A RELEVANCE FLOOR. Recall returned its top five whatever they scored, so a run
about shift cover was handed five memories about certifications simply because
nothing better existed — and the block tells the model these are things the
workspace remembered, so it reads them as pertinent. Embeddings are
unit-normalised, so knowledge_dot is cosine, and 0.30 is where text is usually
about something else. A judgement rather than a measurement, and the honest way
to tune it is to watch what gets carried on real questions.

NO DUPLICATES. The same standing preference comes up in conversation after
conversation, and each run that hears it has no idea the last one wrote it
down. Five recall slots spent on one fact restated five ways is the normal
failure, not a rare one. A write with the same normalised text, in the same org
and about the same subject, pushes the existing memory's expiry out instead of
adding a row — matched on the same sentence rather than a similar one, because
collapsing two genuinely different facts is the worse error.

EXPIRY THAT DELETES. expires_at was set and filtered on read, and nothing ever
removed anything: the row was invisible and still retained. "We keep it ninety
days" has to be true of the table, not only of the query. Prune is batched, and
a redaction is kept for a thirty-day grace period so an erasure stays provable
shortly afterwards.

It runs in the maintenance sweeper that already exists rather than a second
scheduler — same ticker, same cancellation, same failure isolation. That forced
one honest change: Maintenance() used to be nil without OAuth, on the reasoning
that there was nothing to sweep. There is now, and a retention promise enforced
only when an unrelated feature happens to be enabled is not a promise. The test
that asserted the old behaviour now asserts the new one and says why.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-10-07 20:29:28 +05:30
parent 43dabb5f72
commit e90bc33d0f
5 changed files with 199 additions and 26 deletions

View File

@@ -37,6 +37,7 @@ import (
"github.com/krow/krow-backend/go-api/internal/db"
"github.com/krow/krow-backend/go-api/internal/definition"
"github.com/krow/krow-backend/go-api/internal/knowledge"
"github.com/krow/krow-backend/go-api/internal/memory"
"github.com/krow/krow-backend/go-api/internal/ratelimit"
"github.com/krow/krow-backend/go-api/internal/runtime"
"github.com/krow/krow-backend/go-api/internal/service"
@@ -60,6 +61,11 @@ type Server struct {
runs *runtime.RunReader
version string
// memories is the long-term memory store, for the scheduled sweep that
// enforces its retention promise. The runtime builds its own; this one is
// here so maintenance can prune without reaching through the engine.
memories *memory.Store
// toolCatalogue is the tool set an agent author may choose from.
//
// Built whether or not a model credential exists: the catalogue describes
@@ -254,6 +260,11 @@ func New(cfg *config.Config, database *db.DB, log *slog.Logger, opts ...Option)
s.runs = runtime.NewRunReader(database.Pool)
}
// Memories are swept wherever the database is, agents or not: a deployment
// that stops serving agents still holds what earlier ones remembered, and
// retention is a promise about the table rather than about the feature.
s.memories = memory.New(database.Pool, runtime.NewEmbedder(*cfg))
// Built the same way the runtime builds its own, so the list an author is
// offered is the list their agent will actually have.
toolRegistry := runtime.DefaultTools(