Read from the code rather than from intent, after a conversation in which the honest answer to "are we using RAG and memory correctly" took an hour of digging that nobody should have to repeat. It records three things that are easy to get wrong when describing this system to somebody else: that retrieval is hybrid with the permission predicate as a PRE-filter rather than a post-filter, that only two of nine agents use it and that is the design, and that long-term memory is absent on purpose rather than by omission — with the schema comment that says so quoted in place. It also states what the document does NOT support claiming, because the failure this guards against is overclaiming to a client, not under-documenting. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
5.7 KiB
Retrieval and memory, as actually built
Written 2026-10-07, from a read of the code rather than from intent. It records what is there, what is deliberately absent, and the reasoning for each — so the next person does not have to re-derive it, and so nobody claims more than the system does.
The short version
| Layer | State |
|---|---|
| LLM gateway, multi-provider | built |
| Hybrid RAG with ACL pre-filter | built, used by 2 of 9 agents |
| Working memory (within one answer) | built |
| Conversation memory (across turns) | built, browser-side, token-budgeted |
| Long-term / semantic memory | not built, deliberately |
1. The model layer
internal/gateway/ speaks one wire protocol — OpenAI chat-completions — which
is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM
and Ollama serve. Supporting six vendors is one implementation and six base
URLs.
Three tiers (fast / balanced / deep) selected per agent by its
reasoning: value. Tier → model is a deployment knob; tier → effort is not,
because "deep" must mean the same thing in every deployment.
Failover is per provider, because a free tier's ceiling is per provider: a second key is a second budget. It is refused on terminal errors (a rejected credential fails the same way everywhere) and on a conversation that has already called a tool, because provider-specific metadata on that call cannot travel. When a rate limit lands mid-run, the loop re-runs the whole turn on the next provider from the original question — unless the run carried a confirmation, in which case it is never replayed, because re-running re-runs its tools and a write twice is two shifts assigned.
2. Retrieval
Used by control-center-agent and krow-workforce-agent only. The other
seven answer from SQL tools. That split is the design: "how many open positions"
is a query, not a retrieval problem, and routing it through a corpus would make
a precise answer approximate.
What makes it more than a vector lookup:
- Hybrid. Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only: semantic search is weak on exact terms and this corpus is full of them — shift codes, certification names, venues. Not keyword-only either, which is the failure a deployment with no embedding credential ships by accident.
- Permission as a PRE-filter. The same predicate goes into both queries'
WHEREclauses. Rank first and drop afterwards and forbidden rows leak through the shape of what is left: a short result set, a top-3 with a hole in it, a confidence that tracks documents the caller cannot see. - Honest degradation. With no embedder, results come back marked
"no embedder is configured; these results are keyword-only"rather than quietly worse. - Citable. Every chunk carries the ids to point back at it.
- Injection boundary. Chunks go into a delimited block in a user message, never the system prompt, and document text cannot close its own fence.
DefaultK is 4. It was 8; a retrieval block is re-sent on every model call of a
run, and the deployment's ceiling is 8,000 tokens a minute.
3. Memory
Working memory — within one run the loop accumulates tool calls and results across model calls. Discarded when the run ends.
Conversation memory — components/ai-assistant/recall.ts. The last few
exchanges travel with the question.
It lives in the browser because the API takes an input and no message list,
and agent_runs records each run independently. That is right for an API and
wrong for a panel that reads as a conversation. The proper fix is a messages
array on the run request; recall.ts is shaped like that future field so the
swap is a deletion rather than a rewrite.
Bounded in tokens, not turns. Turns are not a unit of cost — three short exchanges are nothing, three carrying a table each is a question that no longer fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped to 400 characters, eviction oldest-first because dropping the most recent exchange drops the one the follow-up is about. The transcript is fenced and labelled as data: an earlier answer is the model's own words, but an earlier QUESTION is the reader's, and a reader can type anything.
Long-term and semantic memory are absent, and that is a decision. The schema says so:
conversations— a run is one turn. Threading runs into a conversation is a surface-layer concern and no surface asks for it yet; adding the column later is trivial, and inventing the semantics now is not.
Nothing is learned across runs. agent_runs is an audit record and is never
replayed into a prompt. In a hiring product that is the conservative choice:
memory that influences a hiring answer is retained profiling of named
candidates, and it has to be auditable and explainable before it touches a
record. The retrieval and ACL layers it would reuse already exist, so building
it is a focused piece of work rather than a rebuild.
4. What to verify before trusting retrieval in a demo
Configuration being present does not mean a corpus was ingested:
SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks;
0— RAG is wired and has nothing to retrieve- chunks but no embeddings — keyword-only; the dense half never runs
- both non-zero — working as designed
5. Claims this document supports
Hybrid retrieval with permission pre-filtering, multi-provider failover, per-run budgets, a full trajectory per answer, human confirmation before any write, conversation memory with a token budget, and a bilingual interface.
It does not support "complete memory" or "everything is retrieval-backed". Both are false, and the second is false on purpose.