Some checks failed
CI / check (push) Failing after 5m3s
The store, the remember tool and the recall path are deployed, and the strings were grepped out of the running binary rather than inferred from a green test run. So "built but not wired" is now wrong in the direction that undersells it. The replacement is careful about the opposite error. A live memory and a populated memory are different things: it accumulates from use, starts empty on any deployment, and a memory only enters a prompt on a LATER run — so the conversation that creates one shows no difference at all. That is the thing most likely to be mistaken for the feature not working, so it is stated where somebody checking would look. The claims list is updated accordingly. "It already knows your workspace" has replaced "the agent remembers across sessions" as the sentence most likely to be said by accident, because the feature now works and the table is still empty. Also records the write trigger and why the two cheaper designs were rejected — an extraction pass costs a whole model call per run against an 8,000 token ceiling, and a heuristic remembers the wrong things because the shape of a run says nothing about whether a fact outlives it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
203 lines
10 KiB
Markdown
203 lines
10 KiB
Markdown
# Retrieval and memory, as actually built
|
|
|
|
Written 2026-10-07 and updated the same day as memory went from absent to
|
|
deployed, from a read of the code rather than from intent. Where it says
|
|
something is live, that was checked against the running binary in production,
|
|
not against a test run. It records
|
|
what is there, what is deliberately absent, and the reasoning for each — so the
|
|
next person does not have to re-derive it, and so nobody claims more than the
|
|
system does.
|
|
|
|
## The short version
|
|
|
|
| Layer | State |
|
|
| --- | --- |
|
|
| LLM gateway, multi-provider | **built** |
|
|
| Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents |
|
|
| Working memory (within one answer) | **built** |
|
|
| Conversation memory (across turns) | **built**, browser-side, token-budgeted |
|
|
| Long-term / semantic memory | **built, wired and deployed** — empty until used |
|
|
|
|
## 1. The model layer
|
|
|
|
`internal/gateway/` speaks one wire protocol — OpenAI chat-completions — which
|
|
is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM
|
|
and Ollama serve. Supporting six vendors is one implementation and six base
|
|
URLs.
|
|
|
|
Three tiers (`fast` / `balanced` / `deep`) selected per agent by its
|
|
`reasoning:` value. Tier → model is a deployment knob; tier → effort is not,
|
|
because "deep" must mean the same thing in every deployment.
|
|
|
|
Failover is per provider, because a free tier's ceiling is per provider: a
|
|
second key is a second budget. It is refused on terminal errors (a rejected
|
|
credential fails the same way everywhere) and on a conversation that has already
|
|
called a tool, because provider-specific metadata on that call cannot travel.
|
|
When a rate limit lands mid-run, the loop re-runs the whole turn on the next
|
|
provider from the original question — unless the run carried a confirmation, in
|
|
which case it is never replayed, because re-running re-runs its tools and a
|
|
write twice is two shifts assigned.
|
|
|
|
## 2. Retrieval
|
|
|
|
**Used by `control-center-agent` and `krow-workforce-agent` only.** The other
|
|
seven answer from SQL tools. That split is the design: "how many open positions"
|
|
is a query, not a retrieval problem, and routing it through a corpus would make
|
|
a precise answer approximate.
|
|
|
|
What makes it more than a vector lookup:
|
|
|
|
- **Hybrid.** Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only:
|
|
semantic search is weak on exact terms and this corpus is full of them — shift
|
|
codes, certification names, venues. Not keyword-only either, which is the
|
|
failure a deployment with no embedding credential ships by accident.
|
|
- **Permission as a PRE-filter.** The same predicate goes into both queries'
|
|
`WHERE` clauses. Rank first and drop afterwards and forbidden rows leak
|
|
through the shape of what is left: a short result set, a top-3 with a hole in
|
|
it, a confidence that tracks documents the caller cannot see.
|
|
- **Honest degradation.** With no embedder, results come back marked
|
|
`"no embedder is configured; these results are keyword-only"` rather than
|
|
quietly worse.
|
|
- **Citable.** Every chunk carries the ids to point back at it.
|
|
- **Injection boundary.** Chunks go into a delimited block in a *user* message,
|
|
never the system prompt, and document text cannot close its own fence.
|
|
|
|
`DefaultK` is 4. It was 8; a retrieval block is re-sent on every model call of a
|
|
run, and the deployment's ceiling is 8,000 tokens a minute.
|
|
|
|
## 3. Memory
|
|
|
|
**Working memory** — within one run the loop accumulates tool calls and results
|
|
across model calls. Discarded when the run ends.
|
|
|
|
**Conversation memory** — `components/ai-assistant/recall.ts`. The last few
|
|
exchanges travel with the question.
|
|
|
|
It lives in the browser because the API takes an `input` and no message list,
|
|
and `agent_runs` records each run independently. That is right for an API and
|
|
wrong for a panel that reads as a conversation. The proper fix is a `messages`
|
|
array on the run request; `recall.ts` is shaped like that future field so the
|
|
swap is a deletion rather than a rewrite.
|
|
|
|
**Bounded in tokens, not turns.** Turns are not a unit of cost — three short
|
|
exchanges are nothing, three carrying a table each is a question that no longer
|
|
fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped
|
|
to 400 characters, eviction oldest-first because dropping the most recent
|
|
exchange drops the one the follow-up is about. The transcript is fenced and
|
|
labelled as data: an earlier answer is the model's own words, but an earlier
|
|
QUESTION is the reader's, and a reader can type anything.
|
|
|
|
**Long-term memory is live.** `internal/memory`, migration `000017`, the
|
|
`remember` tool and the recall path in `loop.go` are all deployed and verified
|
|
in the running binary.
|
|
|
|
It is EMPTY until a workspace uses it. Memory accumulates from what agents are
|
|
told; it does not arrive populated, and a memory only enters a prompt on a
|
|
LATER run — so the conversation that creates one shows no difference. That is
|
|
the design, not a fault, and it is the thing most likely to be mistaken for the
|
|
feature not working.
|
|
|
|
What the store is, and why it is mostly provenance:
|
|
|
|
- **Org-scoped.** A memory written while one recruiter worked is available to
|
|
the next, because a workspace's view of its own hiring should not reset per
|
|
seat. `org_id` is NOT NULL and the predicate is in every read.
|
|
- **It remembers both kinds** — operational facts and observations about named
|
|
people. The second is why the rest of this list exists. "This applicant
|
|
seemed unreliable", written automatically and read into a later hiring
|
|
answer, is profiling under GDPR and is the artefact an employment claim would
|
|
be built on.
|
|
- **Every personal memory names its subject**, refused at the door rather than
|
|
defaulted. A memory about somebody that names nobody cannot be shown to them
|
|
or erased for them, which is the single property that makes holding it
|
|
defensible. `Held()` answers a subject access request and `Forget()` an
|
|
erasure, each in one statement; `Forget` is a soft delete, so the erasure is
|
|
itself on the record.
|
|
- **Provenance on every row** — the author (an agent's inference or a person's
|
|
note) and the run that wrote it, so "why did it say that" survives memory
|
|
entering the picture, and an inference is never read back as though a person
|
|
had written it.
|
|
- **Everything expires**, ninety days by default. A hiring workspace changes
|
|
shape over a quarter, and a stale fact read as a current one is worse than no
|
|
memory at all.
|
|
- **A memory cannot decide.** The block reaching the model is fenced and
|
|
labelled on the same terms as a retrieved document, states the origin of each
|
|
line, and says plainly that a memory is never a reason on its own to accept
|
|
or reject anybody. That sentence is pinned by a test, so an edit cannot
|
|
quietly drop it.
|
|
|
|
Recall is semantic where an embedder exists and newest-first where it does not,
|
|
and says which happened rather than silently returning recency. Five memories
|
|
by default: this competes for the same prompt as the tool catalogue and the
|
|
retrieved block, against a ceiling of 8,000 tokens a minute.
|
|
|
|
**The write trigger: a tool, not an extraction pass.** Three designs were
|
|
available. A second model call after each run judges well and costs a whole
|
|
extra call against a ceiling of 8,000 tokens a minute, on every run, most of
|
|
which have nothing worth keeping. A heuristic in the loop is cheap and
|
|
remembers the wrong things, because the shape of a run says nothing about
|
|
whether a fact outlives it. So the agent gets a `remember` tool: it costs
|
|
nothing extra, it is automatic in the sense that matters — nobody types
|
|
"remember this" — and it is visible in the trajectory, which an extraction pass
|
|
would not be.
|
|
|
|
**It is a confirmed write**, because `EffectWrite` forces it and that is the
|
|
invariant working rather than an obstacle: this stores personal data that will
|
|
shape later hiring answers. A person sees the sentence, who it is about, that
|
|
an agent and not a person decided it, and when it expires. If workspace facts
|
|
should later be kept without asking, the honest change is a SECOND tool scoped
|
|
to workspace subjects — loosening this one would quietly make personal
|
|
memories unconfirmed too.
|
|
|
|
**The read path.** Memories are recalled before retrieval and placed before it:
|
|
a standing preference frames how documents should be read, where a document
|
|
does not frame a preference. The question stays last, because a model reads the
|
|
last thing and answers it. Skipped for smalltalk on the same terms as
|
|
retrieval — nobody needs remembering to say good morning, and paying for it is
|
|
how "hi" came to cost six thousand tokens.
|
|
|
|
**It fails quiet and is recorded loudly.** A memory store that is unreachable
|
|
does not take the run with it: an answer without memory is worse, not wrong,
|
|
and the alternative is an outage in the knowledge layer becoming an outage in
|
|
the product. The trajectory records the failure, and records separately when
|
|
the store returned recency instead of relevance.
|
|
|
|
**Conversation threading is still absent**, separately, and the schema still
|
|
says why:
|
|
|
|
> `conversations` — a run is one turn. Threading runs into a conversation is a
|
|
> surface-layer concern and no surface asks for it yet; adding the column later
|
|
> is trivial, and inventing the semantics now is not.
|
|
|
|
## 4. What to verify before trusting retrieval in a demo
|
|
|
|
Configuration being present does not mean a corpus was ingested:
|
|
|
|
```sql
|
|
SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks;
|
|
```
|
|
|
|
- `0` — RAG is wired and has nothing to retrieve
|
|
- chunks but no embeddings — keyword-only; the dense half never runs
|
|
- both non-zero — working as designed
|
|
|
|
## 5. Claims this document supports
|
|
|
|
Hybrid retrieval with permission pre-filtering, multi-provider failover,
|
|
per-run budgets, a full trajectory per answer, human confirmation before any
|
|
write, conversation memory with a token budget, a bilingual interface, and a
|
|
long-term memory store that is built, migrated and auditable by subject.
|
|
|
|
It does NOT support:
|
|
|
|
- "it already knows your workspace" — memory is live and starts empty. It
|
|
accumulates from use, and on a fresh deployment there is nothing in it;
|
|
- "it remembers everything" — five memories per run, ninety-day expiry, and a
|
|
person confirms each one before it is kept;
|
|
- "complete memory" — conversation threading is still a browser-side window,
|
|
not a server-side thread;
|
|
- "everything is retrieval-backed" — two of nine agents are, on purpose.
|
|
|
|
The first is now the one most likely to be said by accident. The feature works;
|
|
the table is empty until somebody uses it.
|