170 lines
8.3 KiB
Markdown
170 lines
8.3 KiB
Markdown
# Retrieval and memory, as actually built
|
|
|
|
Written 2026-10-07 and updated the same day when the memory store landed, from
|
|
a read of the code rather than from intent. It records
|
|
what is there, what is deliberately absent, and the reasoning for each — so the
|
|
next person does not have to re-derive it, and so nobody claims more than the
|
|
system does.
|
|
|
|
## The short version
|
|
|
|
| Layer | State |
|
|
| --- | --- |
|
|
| LLM gateway, multi-provider | **built** |
|
|
| Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents |
|
|
| Working memory (within one answer) | **built** |
|
|
| Conversation memory (across turns) | **built**, browser-side, token-budgeted |
|
|
| Long-term / semantic memory | **store built and migrated; Wiring in Progress** |
|
|
|
|
## 1. The model layer
|
|
|
|
`internal/gateway/` speaks one wire protocol — OpenAI chat-completions — which
|
|
is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM
|
|
and Ollama serve. Supporting six vendors is one implementation and six base
|
|
URLs.
|
|
|
|
Three tiers (`fast` / `balanced` / `deep`) selected per agent by its
|
|
`reasoning:` value. Tier → model is a deployment knob; tier → effort is not,
|
|
because "deep" must mean the same thing in every deployment.
|
|
|
|
Failover is per provider, because a free tier's ceiling is per provider: a
|
|
second key is a second budget. It is refused on terminal errors (a rejected
|
|
credential fails the same way everywhere) and on a conversation that has already
|
|
called a tool, because provider-specific metadata on that call cannot travel.
|
|
When a rate limit lands mid-run, the loop re-runs the whole turn on the next
|
|
provider from the original question — unless the run carried a confirmation, in
|
|
which case it is never replayed, because re-running re-runs its tools and a
|
|
write twice is two shifts assigned.
|
|
|
|
## 2. Retrieval
|
|
|
|
**Used by `control-center-agent` and `krow-workforce-agent` only.** The other
|
|
seven answer from SQL tools. That split is the design: "how many open positions"
|
|
is a query, not a retrieval problem, and routing it through a corpus would make
|
|
a precise answer approximate.
|
|
|
|
What makes it more than a vector lookup:
|
|
|
|
- **Hybrid.** Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only:
|
|
semantic search is weak on exact terms and this corpus is full of them — shift
|
|
codes, certification names, venues. Not keyword-only either, which is the
|
|
failure a deployment with no embedding credential ships by accident.
|
|
- **Permission as a PRE-filter.** The same predicate goes into both queries'
|
|
`WHERE` clauses. Rank first and drop afterwards and forbidden rows leak
|
|
through the shape of what is left: a short result set, a top-3 with a hole in
|
|
it, a confidence that tracks documents the caller cannot see.
|
|
- **Honest degradation.** With no embedder, results come back marked
|
|
`"no embedder is configured; these results are keyword-only"` rather than
|
|
quietly worse.
|
|
- **Citable.** Every chunk carries the ids to point back at it.
|
|
- **Injection boundary.** Chunks go into a delimited block in a *user* message,
|
|
never the system prompt, and document text cannot close its own fence.
|
|
|
|
`DefaultK` is 4. It was 8; a retrieval block is re-sent on every model call of a
|
|
run, and the deployment's ceiling is 8,000 tokens a minute.
|
|
|
|
## 3. Memory
|
|
|
|
**Working memory** — within one run the loop accumulates tool calls and results
|
|
across model calls. Discarded when the run ends.
|
|
|
|
**Conversation memory** — `components/ai-assistant/recall.ts`. The last few
|
|
exchanges travel with the question.
|
|
|
|
It lives in the browser because the API takes an `input` and no message list,
|
|
and `agent_runs` records each run independently. That is right for an API and
|
|
wrong for a panel that reads as a conversation. The proper fix is a `messages`
|
|
array on the run request; `recall.ts` is shaped like that future field so the
|
|
swap is a deletion rather than a rewrite.
|
|
|
|
**Bounded in tokens, not turns.** Turns are not a unit of cost — three short
|
|
exchanges are nothing, three carrying a table each is a question that no longer
|
|
fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped
|
|
to 400 characters, eviction oldest-first because dropping the most recent
|
|
exchange drops the one the follow-up is about. The transcript is fenced and
|
|
labelled as data: an earlier answer is the model's own words, but an earlier
|
|
QUESTION is the reader's, and a reader can type anything.
|
|
|
|
**Long-term memory: the store exists, the behaviour does not.**
|
|
|
|
`internal/memory` and migration `000017_agent_memories` are built and applied
|
|
in production. Nothing in `loop.go` reads or writes a memory yet, so no answer
|
|
has ever been shaped by one. State it that way: the foundation is deployed, the
|
|
feature is not switched on. "We have long-term memory" is not yet true of
|
|
anything a user would experience.
|
|
|
|
What the store is, and why it is mostly provenance:
|
|
|
|
- **Org-scoped.** A memory written while one recruiter worked is available to
|
|
the next, because a workspace's view of its own hiring should not reset per
|
|
seat. `org_id` is NOT NULL and the predicate is in every read.
|
|
- **It remembers both kinds** — operational facts and observations about named
|
|
people. The second is why the rest of this list exists. "This applicant
|
|
seemed unreliable", written automatically and read into a later hiring
|
|
answer, is profiling under GDPR and is the artefact an employment claim would
|
|
be built on.
|
|
- **Every personal memory names its subject**, refused at the door rather than
|
|
defaulted. A memory about somebody that names nobody cannot be shown to them
|
|
or erased for them, which is the single property that makes holding it
|
|
defensible. `Held()` answers a subject access request and `Forget()` an
|
|
erasure, each in one statement; `Forget` is a soft delete, so the erasure is
|
|
itself on the record.
|
|
- **Provenance on every row** — the author (an agent's inference or a person's
|
|
note) and the run that wrote it, so "why did it say that" survives memory
|
|
entering the picture, and an inference is never read back as though a person
|
|
had written it.
|
|
- **Everything expires**, ninety days by default. A hiring workspace changes
|
|
shape over a quarter, and a stale fact read as a current one is worse than no
|
|
memory at all.
|
|
- **A memory cannot decide.** The block reaching the model is fenced and
|
|
labelled on the same terms as a retrieved document, states the origin of each
|
|
line, and says plainly that a memory is never a reason on its own to accept
|
|
or reject anybody. That sentence is pinned by a test, so an edit cannot
|
|
quietly drop it.
|
|
|
|
Recall is semantic where an embedder exists and newest-first where it does not,
|
|
and says which happened rather than silently returning recency. Five memories
|
|
by default: this competes for the same prompt as the tool catalogue and the
|
|
retrieved block, against a ceiling of 8,000 tokens a minute.
|
|
|
|
**What is left, and it is the hard part.** Wiring the read into `loop.go` is
|
|
small. The write trigger is not: automatic means something judges what is worth
|
|
remembering, and a bad judge fills the table with noise that then shapes every
|
|
answer after it. That decision is open.
|
|
|
|
**Conversation threading is still absent**, separately, and the schema still
|
|
says why:
|
|
|
|
> `conversations` — a run is one turn. Threading runs into a conversation is a
|
|
> surface-layer concern and no surface asks for it yet; adding the column later
|
|
> is trivial, and inventing the semantics now is not.
|
|
|
|
## 4. What to verify before trusting retrieval in a demo
|
|
|
|
Configuration being present does not mean a corpus was ingested:
|
|
|
|
```sql
|
|
SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks;
|
|
```
|
|
|
|
- `0` — RAG is wired and has nothing to retrieve
|
|
- chunks but no embeddings — keyword-only; the dense half never runs
|
|
- both non-zero — working as designed
|
|
|
|
## 5. Claims this document supports
|
|
|
|
Hybrid retrieval with permission pre-filtering, multi-provider failover,
|
|
per-run budgets, a full trajectory per answer, human confirmation before any
|
|
write, conversation memory with a token budget, a bilingual interface, and a
|
|
long-term memory store that is built, migrated and auditable by subject.
|
|
|
|
It does NOT support:
|
|
|
|
- "the agent remembers across sessions" — the store is not wired, so nothing
|
|
does yet;
|
|
- "complete memory";
|
|
- "everything is retrieval-backed" — two of nine agents are, on purpose.
|
|
|
|
The first of those is the one most likely to be said by accident, because the
|
|
code and the table both exist. Built is not the same as switched on.
|