Files
krow_talent_app/docs/architecture-rag-and-memory.md
Aravind 76f168673a
Some checks failed
CI / check (push) Failing after 5m6s
Write down what retrieval and memory actually are
Read from the code rather than from intent, after a conversation in which the
honest answer to "are we using RAG and memory correctly" took an hour of
digging that nobody should have to repeat.

It records three things that are easy to get wrong when describing this system
to somebody else: that retrieval is hybrid with the permission predicate as a
PRE-filter rather than a post-filter, that only two of nine agents use it and
that is the design, and that long-term memory is absent on purpose rather than
by omission — with the schema comment that says so quoted in place.

It also states what the document does NOT support claiming, because the failure
this guards against is overclaiming to a client, not under-documenting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-07 19:49:04 +05:30

5.7 KiB

Retrieval and memory, as actually built

Written 2026-10-07, from a read of the code rather than from intent. It records what is there, what is deliberately absent, and the reasoning for each — so the next person does not have to re-derive it, and so nobody claims more than the system does.

The short version

Layer State
LLM gateway, multi-provider built
Hybrid RAG with ACL pre-filter built, used by 2 of 9 agents
Working memory (within one answer) built
Conversation memory (across turns) built, browser-side, token-budgeted
Long-term / semantic memory not built, deliberately

1. The model layer

internal/gateway/ speaks one wire protocol — OpenAI chat-completions — which is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM and Ollama serve. Supporting six vendors is one implementation and six base URLs.

Three tiers (fast / balanced / deep) selected per agent by its reasoning: value. Tier → model is a deployment knob; tier → effort is not, because "deep" must mean the same thing in every deployment.

Failover is per provider, because a free tier's ceiling is per provider: a second key is a second budget. It is refused on terminal errors (a rejected credential fails the same way everywhere) and on a conversation that has already called a tool, because provider-specific metadata on that call cannot travel. When a rate limit lands mid-run, the loop re-runs the whole turn on the next provider from the original question — unless the run carried a confirmation, in which case it is never replayed, because re-running re-runs its tools and a write twice is two shifts assigned.

2. Retrieval

Used by control-center-agent and krow-workforce-agent only. The other seven answer from SQL tools. That split is the design: "how many open positions" is a query, not a retrieval problem, and routing it through a corpus would make a precise answer approximate.

What makes it more than a vector lookup:

  • Hybrid. Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only: semantic search is weak on exact terms and this corpus is full of them — shift codes, certification names, venues. Not keyword-only either, which is the failure a deployment with no embedding credential ships by accident.
  • Permission as a PRE-filter. The same predicate goes into both queries' WHERE clauses. Rank first and drop afterwards and forbidden rows leak through the shape of what is left: a short result set, a top-3 with a hole in it, a confidence that tracks documents the caller cannot see.
  • Honest degradation. With no embedder, results come back marked "no embedder is configured; these results are keyword-only" rather than quietly worse.
  • Citable. Every chunk carries the ids to point back at it.
  • Injection boundary. Chunks go into a delimited block in a user message, never the system prompt, and document text cannot close its own fence.

DefaultK is 4. It was 8; a retrieval block is re-sent on every model call of a run, and the deployment's ceiling is 8,000 tokens a minute.

3. Memory

Working memory — within one run the loop accumulates tool calls and results across model calls. Discarded when the run ends.

Conversation memory — components/ai-assistant/recall.ts. The last few exchanges travel with the question.

It lives in the browser because the API takes an input and no message list, and agent_runs records each run independently. That is right for an API and wrong for a panel that reads as a conversation. The proper fix is a messages array on the run request; recall.ts is shaped like that future field so the swap is a deletion rather than a rewrite.

Bounded in tokens, not turns. Turns are not a unit of cost — three short exchanges are nothing, three carrying a table each is a question that no longer fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped to 400 characters, eviction oldest-first because dropping the most recent exchange drops the one the follow-up is about. The transcript is fenced and labelled as data: an earlier answer is the model's own words, but an earlier QUESTION is the reader's, and a reader can type anything.

Long-term and semantic memory are absent, and that is a decision. The schema says so:

conversations — a run is one turn. Threading runs into a conversation is a surface-layer concern and no surface asks for it yet; adding the column later is trivial, and inventing the semantics now is not.

Nothing is learned across runs. agent_runs is an audit record and is never replayed into a prompt. In a hiring product that is the conservative choice: memory that influences a hiring answer is retained profiling of named candidates, and it has to be auditable and explainable before it touches a record. The retrieval and ACL layers it would reuse already exist, so building it is a focused piece of work rather than a rebuild.

4. What to verify before trusting retrieval in a demo

Configuration being present does not mean a corpus was ingested:

SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks;
  • 0 — RAG is wired and has nothing to retrieve
  • chunks but no embeddings — keyword-only; the dense half never runs
  • both non-zero — working as designed

5. Claims this document supports

Hybrid retrieval with permission pre-filtering, multi-provider failover, per-run budgets, a full trajectory per answer, human confirmation before any write, conversation memory with a token budget, and a bilingual interface.

It does not support "complete memory" or "everything is retrieval-backed". Both are false, and the second is false on purpose.