diff --git a/docs/architecture-rag-and-memory.md b/docs/architecture-rag-and-memory.md new file mode 100644 index 0000000..5b97808 --- /dev/null +++ b/docs/architecture-rag-and-memory.md @@ -0,0 +1,120 @@ +# Retrieval and memory, as actually built + +Written 2026-10-07, from a read of the code rather than from intent. It records +what is there, what is deliberately absent, and the reasoning for each — so the +next person does not have to re-derive it, and so nobody claims more than the +system does. + +## The short version + +| Layer | State | +| --- | --- | +| LLM gateway, multi-provider | **built** | +| Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents | +| Working memory (within one answer) | **built** | +| Conversation memory (across turns) | **built**, browser-side, token-budgeted | +| Long-term / semantic memory | **not built**, deliberately | + +## 1. The model layer + +`internal/gateway/` speaks one wire protocol — OpenAI chat-completions — which +is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM +and Ollama serve. Supporting six vendors is one implementation and six base +URLs. + +Three tiers (`fast` / `balanced` / `deep`) selected per agent by its +`reasoning:` value. Tier → model is a deployment knob; tier → effort is not, +because "deep" must mean the same thing in every deployment. + +Failover is per provider, because a free tier's ceiling is per provider: a +second key is a second budget. It is refused on terminal errors (a rejected +credential fails the same way everywhere) and on a conversation that has already +called a tool, because provider-specific metadata on that call cannot travel. +When a rate limit lands mid-run, the loop re-runs the whole turn on the next +provider from the original question — unless the run carried a confirmation, in +which case it is never replayed, because re-running re-runs its tools and a +write twice is two shifts assigned. + +## 2. Retrieval + +**Used by `control-center-agent` and `krow-workforce-agent` only.** The other +seven answer from SQL tools. That split is the design: "how many open positions" +is a query, not a retrieval problem, and routing it through a corpus would make +a precise answer approximate. + +What makes it more than a vector lookup: + +- **Hybrid.** Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only: + semantic search is weak on exact terms and this corpus is full of them — shift + codes, certification names, venues. Not keyword-only either, which is the + failure a deployment with no embedding credential ships by accident. +- **Permission as a PRE-filter.** The same predicate goes into both queries' + `WHERE` clauses. Rank first and drop afterwards and forbidden rows leak + through the shape of what is left: a short result set, a top-3 with a hole in + it, a confidence that tracks documents the caller cannot see. +- **Honest degradation.** With no embedder, results come back marked + `"no embedder is configured; these results are keyword-only"` rather than + quietly worse. +- **Citable.** Every chunk carries the ids to point back at it. +- **Injection boundary.** Chunks go into a delimited block in a *user* message, + never the system prompt, and document text cannot close its own fence. + +`DefaultK` is 4. It was 8; a retrieval block is re-sent on every model call of a +run, and the deployment's ceiling is 8,000 tokens a minute. + +## 3. Memory + +**Working memory** — within one run the loop accumulates tool calls and results +across model calls. Discarded when the run ends. + +**Conversation memory** — `components/ai-assistant/recall.ts`. The last few +exchanges travel with the question. + +It lives in the browser because the API takes an `input` and no message list, +and `agent_runs` records each run independently. That is right for an API and +wrong for a panel that reads as a conversation. The proper fix is a `messages` +array on the run request; `recall.ts` is shaped like that future field so the +swap is a deletion rather than a rewrite. + +**Bounded in tokens, not turns.** Turns are not a unit of cost — three short +exchanges are nothing, three carrying a table each is a question that no longer +fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped +to 400 characters, eviction oldest-first because dropping the most recent +exchange drops the one the follow-up is about. The transcript is fenced and +labelled as data: an earlier answer is the model's own words, but an earlier +QUESTION is the reader's, and a reader can type anything. + +**Long-term and semantic memory are absent, and that is a decision.** The +schema says so: + +> `conversations` — a run is one turn. Threading runs into a conversation is a +> surface-layer concern and no surface asks for it yet; adding the column later +> is trivial, and inventing the semantics now is not. + +Nothing is learned across runs. `agent_runs` is an audit record and is never +replayed into a prompt. In a hiring product that is the conservative choice: +memory that influences a hiring answer is retained profiling of named +candidates, and it has to be auditable and explainable before it touches a +record. The retrieval and ACL layers it would reuse already exist, so building +it is a focused piece of work rather than a rebuild. + +## 4. What to verify before trusting retrieval in a demo + +Configuration being present does not mean a corpus was ingested: + +```sql +SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks; +``` + +- `0` — RAG is wired and has nothing to retrieve +- chunks but no embeddings — keyword-only; the dense half never runs +- both non-zero — working as designed + +## 5. Claims this document supports + +Hybrid retrieval with permission pre-filtering, multi-provider failover, +per-run budgets, a full trajectory per answer, human confirmation before any +write, conversation memory with a token budget, and a bilingual interface. + +It does not support "complete memory" or "everything is retrieval-backed". +Both are false, and the second is false on purpose.