Files
krow_talent_app/docs/architecture-rag-and-memory.md
Aravind 76f168673a
Some checks failed
CI / check (push) Failing after 5m6s
Write down what retrieval and memory actually are
Read from the code rather than from intent, after a conversation in which the
honest answer to "are we using RAG and memory correctly" took an hour of
digging that nobody should have to repeat.

It records three things that are easy to get wrong when describing this system
to somebody else: that retrieval is hybrid with the permission predicate as a
PRE-filter rather than a post-filter, that only two of nine agents use it and
that is the design, and that long-term memory is absent on purpose rather than
by omission — with the schema comment that says so quoted in place.

It also states what the document does NOT support claiming, because the failure
this guards against is overclaiming to a client, not under-documenting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-07 19:49:04 +05:30

121 lines
5.7 KiB
Markdown

# Retrieval and memory, as actually built
Written 2026-10-07, from a read of the code rather than from intent. It records
what is there, what is deliberately absent, and the reasoning for each — so the
next person does not have to re-derive it, and so nobody claims more than the
system does.
## The short version
| Layer | State |
| --- | --- |
| LLM gateway, multi-provider | **built** |
| Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents |
| Working memory (within one answer) | **built** |
| Conversation memory (across turns) | **built**, browser-side, token-budgeted |
| Long-term / semantic memory | **not built**, deliberately |
## 1. The model layer
`internal/gateway/` speaks one wire protocol — OpenAI chat-completions — which
is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM
and Ollama serve. Supporting six vendors is one implementation and six base
URLs.
Three tiers (`fast` / `balanced` / `deep`) selected per agent by its
`reasoning:` value. Tier → model is a deployment knob; tier → effort is not,
because "deep" must mean the same thing in every deployment.
Failover is per provider, because a free tier's ceiling is per provider: a
second key is a second budget. It is refused on terminal errors (a rejected
credential fails the same way everywhere) and on a conversation that has already
called a tool, because provider-specific metadata on that call cannot travel.
When a rate limit lands mid-run, the loop re-runs the whole turn on the next
provider from the original question — unless the run carried a confirmation, in
which case it is never replayed, because re-running re-runs its tools and a
write twice is two shifts assigned.
## 2. Retrieval
**Used by `control-center-agent` and `krow-workforce-agent` only.** The other
seven answer from SQL tools. That split is the design: "how many open positions"
is a query, not a retrieval problem, and routing it through a corpus would make
a precise answer approximate.
What makes it more than a vector lookup:
- **Hybrid.** Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only:
semantic search is weak on exact terms and this corpus is full of them — shift
codes, certification names, venues. Not keyword-only either, which is the
failure a deployment with no embedding credential ships by accident.
- **Permission as a PRE-filter.** The same predicate goes into both queries'
`WHERE` clauses. Rank first and drop afterwards and forbidden rows leak
through the shape of what is left: a short result set, a top-3 with a hole in
it, a confidence that tracks documents the caller cannot see.
- **Honest degradation.** With no embedder, results come back marked
`"no embedder is configured; these results are keyword-only"` rather than
quietly worse.
- **Citable.** Every chunk carries the ids to point back at it.
- **Injection boundary.** Chunks go into a delimited block in a *user* message,
never the system prompt, and document text cannot close its own fence.
`DefaultK` is 4. It was 8; a retrieval block is re-sent on every model call of a
run, and the deployment's ceiling is 8,000 tokens a minute.
## 3. Memory
**Working memory** — within one run the loop accumulates tool calls and results
across model calls. Discarded when the run ends.
**Conversation memory** — `components/ai-assistant/recall.ts`. The last few
exchanges travel with the question.
It lives in the browser because the API takes an `input` and no message list,
and `agent_runs` records each run independently. That is right for an API and
wrong for a panel that reads as a conversation. The proper fix is a `messages`
array on the run request; `recall.ts` is shaped like that future field so the
swap is a deletion rather than a rewrite.
**Bounded in tokens, not turns.** Turns are not a unit of cost — three short
exchanges are nothing, three carrying a table each is a question that no longer
fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped
to 400 characters, eviction oldest-first because dropping the most recent
exchange drops the one the follow-up is about. The transcript is fenced and
labelled as data: an earlier answer is the model's own words, but an earlier
QUESTION is the reader's, and a reader can type anything.
**Long-term and semantic memory are absent, and that is a decision.** The
schema says so:
> `conversations` — a run is one turn. Threading runs into a conversation is a
> surface-layer concern and no surface asks for it yet; adding the column later
> is trivial, and inventing the semantics now is not.
Nothing is learned across runs. `agent_runs` is an audit record and is never
replayed into a prompt. In a hiring product that is the conservative choice:
memory that influences a hiring answer is retained profiling of named
candidates, and it has to be auditable and explainable before it touches a
record. The retrieval and ACL layers it would reuse already exist, so building
it is a focused piece of work rather than a rebuild.
## 4. What to verify before trusting retrieval in a demo
Configuration being present does not mean a corpus was ingested:
```sql
SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks;
```
- `0` — RAG is wired and has nothing to retrieve
- chunks but no embeddings — keyword-only; the dense half never runs
- both non-zero — working as designed
## 5. Claims this document supports
Hybrid retrieval with permission pre-filtering, multi-provider failover,
per-run budgets, a full trajectory per answer, human confirmation before any
write, conversation memory with a token budget, and a bilingual interface.
It does not support "complete memory" or "everything is retrieval-backed".
Both are false, and the second is false on purpose.