Some checks failed
CI / check (push) Failing after 5m6s
Read from the code rather than from intent, after a conversation in which the honest answer to "are we using RAG and memory correctly" took an hour of digging that nobody should have to repeat. It records three things that are easy to get wrong when describing this system to somebody else: that retrieval is hybrid with the permission predicate as a PRE-filter rather than a post-filter, that only two of nine agents use it and that is the design, and that long-term memory is absent on purpose rather than by omission — with the schema comment that says so quoted in place. It also states what the document does NOT support claiming, because the failure this guards against is overclaiming to a client, not under-documenting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
121 lines
5.7 KiB
Markdown
121 lines
5.7 KiB
Markdown
# Retrieval and memory, as actually built
|
|
|
|
Written 2026-10-07, from a read of the code rather than from intent. It records
|
|
what is there, what is deliberately absent, and the reasoning for each — so the
|
|
next person does not have to re-derive it, and so nobody claims more than the
|
|
system does.
|
|
|
|
## The short version
|
|
|
|
| Layer | State |
|
|
| --- | --- |
|
|
| LLM gateway, multi-provider | **built** |
|
|
| Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents |
|
|
| Working memory (within one answer) | **built** |
|
|
| Conversation memory (across turns) | **built**, browser-side, token-budgeted |
|
|
| Long-term / semantic memory | **not built**, deliberately |
|
|
|
|
## 1. The model layer
|
|
|
|
`internal/gateway/` speaks one wire protocol — OpenAI chat-completions — which
|
|
is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM
|
|
and Ollama serve. Supporting six vendors is one implementation and six base
|
|
URLs.
|
|
|
|
Three tiers (`fast` / `balanced` / `deep`) selected per agent by its
|
|
`reasoning:` value. Tier → model is a deployment knob; tier → effort is not,
|
|
because "deep" must mean the same thing in every deployment.
|
|
|
|
Failover is per provider, because a free tier's ceiling is per provider: a
|
|
second key is a second budget. It is refused on terminal errors (a rejected
|
|
credential fails the same way everywhere) and on a conversation that has already
|
|
called a tool, because provider-specific metadata on that call cannot travel.
|
|
When a rate limit lands mid-run, the loop re-runs the whole turn on the next
|
|
provider from the original question — unless the run carried a confirmation, in
|
|
which case it is never replayed, because re-running re-runs its tools and a
|
|
write twice is two shifts assigned.
|
|
|
|
## 2. Retrieval
|
|
|
|
**Used by `control-center-agent` and `krow-workforce-agent` only.** The other
|
|
seven answer from SQL tools. That split is the design: "how many open positions"
|
|
is a query, not a retrieval problem, and routing it through a corpus would make
|
|
a precise answer approximate.
|
|
|
|
What makes it more than a vector lookup:
|
|
|
|
- **Hybrid.** Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only:
|
|
semantic search is weak on exact terms and this corpus is full of them — shift
|
|
codes, certification names, venues. Not keyword-only either, which is the
|
|
failure a deployment with no embedding credential ships by accident.
|
|
- **Permission as a PRE-filter.** The same predicate goes into both queries'
|
|
`WHERE` clauses. Rank first and drop afterwards and forbidden rows leak
|
|
through the shape of what is left: a short result set, a top-3 with a hole in
|
|
it, a confidence that tracks documents the caller cannot see.
|
|
- **Honest degradation.** With no embedder, results come back marked
|
|
`"no embedder is configured; these results are keyword-only"` rather than
|
|
quietly worse.
|
|
- **Citable.** Every chunk carries the ids to point back at it.
|
|
- **Injection boundary.** Chunks go into a delimited block in a *user* message,
|
|
never the system prompt, and document text cannot close its own fence.
|
|
|
|
`DefaultK` is 4. It was 8; a retrieval block is re-sent on every model call of a
|
|
run, and the deployment's ceiling is 8,000 tokens a minute.
|
|
|
|
## 3. Memory
|
|
|
|
**Working memory** — within one run the loop accumulates tool calls and results
|
|
across model calls. Discarded when the run ends.
|
|
|
|
**Conversation memory** — `components/ai-assistant/recall.ts`. The last few
|
|
exchanges travel with the question.
|
|
|
|
It lives in the browser because the API takes an `input` and no message list,
|
|
and `agent_runs` records each run independently. That is right for an API and
|
|
wrong for a panel that reads as a conversation. The proper fix is a `messages`
|
|
array on the run request; `recall.ts` is shaped like that future field so the
|
|
swap is a deletion rather than a rewrite.
|
|
|
|
**Bounded in tokens, not turns.** Turns are not a unit of cost — three short
|
|
exchanges are nothing, three carrying a table each is a question that no longer
|
|
fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped
|
|
to 400 characters, eviction oldest-first because dropping the most recent
|
|
exchange drops the one the follow-up is about. The transcript is fenced and
|
|
labelled as data: an earlier answer is the model's own words, but an earlier
|
|
QUESTION is the reader's, and a reader can type anything.
|
|
|
|
**Long-term and semantic memory are absent, and that is a decision.** The
|
|
schema says so:
|
|
|
|
> `conversations` — a run is one turn. Threading runs into a conversation is a
|
|
> surface-layer concern and no surface asks for it yet; adding the column later
|
|
> is trivial, and inventing the semantics now is not.
|
|
|
|
Nothing is learned across runs. `agent_runs` is an audit record and is never
|
|
replayed into a prompt. In a hiring product that is the conservative choice:
|
|
memory that influences a hiring answer is retained profiling of named
|
|
candidates, and it has to be auditable and explainable before it touches a
|
|
record. The retrieval and ACL layers it would reuse already exist, so building
|
|
it is a focused piece of work rather than a rebuild.
|
|
|
|
## 4. What to verify before trusting retrieval in a demo
|
|
|
|
Configuration being present does not mean a corpus was ingested:
|
|
|
|
```sql
|
|
SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks;
|
|
```
|
|
|
|
- `0` — RAG is wired and has nothing to retrieve
|
|
- chunks but no embeddings — keyword-only; the dense half never runs
|
|
- both non-zero — working as designed
|
|
|
|
## 5. Claims this document supports
|
|
|
|
Hybrid retrieval with permission pre-filtering, multi-provider failover,
|
|
per-run budgets, a full trajectory per answer, human confirmation before any
|
|
write, conversation memory with a token budget, and a bilingual interface.
|
|
|
|
It does not support "complete memory" or "everything is retrieval-backed".
|
|
Both are false, and the second is false on purpose.
|