# Retrieval and memory, as actually built Written 2026-10-07, from a read of the code rather than from intent. It records what is there, what is deliberately absent, and the reasoning for each — so the next person does not have to re-derive it, and so nobody claims more than the system does. ## The short version | Layer | State | | --- | --- | | LLM gateway, multi-provider | **built** | | Hybrid RAG with ACL pre-filter | **built**, used by 2 of 9 agents | | Working memory (within one answer) | **built** | | Conversation memory (across turns) | **built**, browser-side, token-budgeted | | Long-term / semantic memory | **not built**, deliberately | ## 1. The model layer `internal/gateway/` speaks one wire protocol — OpenAI chat-completions — which is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM and Ollama serve. Supporting six vendors is one implementation and six base URLs. Three tiers (`fast` / `balanced` / `deep`) selected per agent by its `reasoning:` value. Tier → model is a deployment knob; tier → effort is not, because "deep" must mean the same thing in every deployment. Failover is per provider, because a free tier's ceiling is per provider: a second key is a second budget. It is refused on terminal errors (a rejected credential fails the same way everywhere) and on a conversation that has already called a tool, because provider-specific metadata on that call cannot travel. When a rate limit lands mid-run, the loop re-runs the whole turn on the next provider from the original question — unless the run carried a confirmation, in which case it is never replayed, because re-running re-runs its tools and a write twice is two shifts assigned. ## 2. Retrieval **Used by `control-center-agent` and `krow-workforce-agent` only.** The other seven answer from SQL tools. That split is the design: "how many open positions" is a query, not a retrieval problem, and routing it through a corpus would make a precise answer approximate. What makes it more than a vector lookup: - **Hybrid.** Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only: semantic search is weak on exact terms and this corpus is full of them — shift codes, certification names, venues. Not keyword-only either, which is the failure a deployment with no embedding credential ships by accident. - **Permission as a PRE-filter.** The same predicate goes into both queries' `WHERE` clauses. Rank first and drop afterwards and forbidden rows leak through the shape of what is left: a short result set, a top-3 with a hole in it, a confidence that tracks documents the caller cannot see. - **Honest degradation.** With no embedder, results come back marked `"no embedder is configured; these results are keyword-only"` rather than quietly worse. - **Citable.** Every chunk carries the ids to point back at it. - **Injection boundary.** Chunks go into a delimited block in a *user* message, never the system prompt, and document text cannot close its own fence. `DefaultK` is 4. It was 8; a retrieval block is re-sent on every model call of a run, and the deployment's ceiling is 8,000 tokens a minute. ## 3. Memory **Working memory** — within one run the loop accumulates tool calls and results across model calls. Discarded when the run ends. **Conversation memory** — `components/ai-assistant/recall.ts`. The last few exchanges travel with the question. It lives in the browser because the API takes an `input` and no message list, and `agent_runs` records each run independently. That is right for an API and wrong for a panel that reads as a conversation. The proper fix is a `messages` array on the run request; `recall.ts` is shaped like that future field so the swap is a deletion rather than a rewrite. **Bounded in tokens, not turns.** Turns are not a unit of cost — three short exchanges are nothing, three carrying a table each is a question that no longer fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped to 400 characters, eviction oldest-first because dropping the most recent exchange drops the one the follow-up is about. The transcript is fenced and labelled as data: an earlier answer is the model's own words, but an earlier QUESTION is the reader's, and a reader can type anything. **Long-term and semantic memory are absent, and that is a decision.** The schema says so: > `conversations` — a run is one turn. Threading runs into a conversation is a > surface-layer concern and no surface asks for it yet; adding the column later > is trivial, and inventing the semantics now is not. Nothing is learned across runs. `agent_runs` is an audit record and is never replayed into a prompt. In a hiring product that is the conservative choice: memory that influences a hiring answer is retained profiling of named candidates, and it has to be auditable and explainable before it touches a record. The retrieval and ACL layers it would reuse already exist, so building it is a focused piece of work rather than a rebuild. ## 4. What to verify before trusting retrieval in a demo Configuration being present does not mean a corpus was ingested: ```sql SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks; ``` - `0` — RAG is wired and has nothing to retrieve - chunks but no embeddings — keyword-only; the dense half never runs - both non-zero — working as designed ## 5. Claims this document supports Hybrid retrieval with permission pre-filtering, multi-provider failover, per-run budgets, a full trajectory per answer, human confirmation before any write, conversation memory with a token budget, and a bilingual interface. It does not support "complete memory" or "everything is retrieval-backed". Both are false, and the second is false on purpose.