Files
krow_talent_app/docs/architecture-rag-and-memory.md
Aravind f0052c6a42
Some checks failed
CI / check (push) Failing after 5m3s
Update docs/architecture-rag-and-memory.md
2026-10-07 14:30:26 +00:00

8.3 KiB

Retrieval and memory, as actually built

Written 2026-10-07 and updated the same day when the memory store landed, from a read of the code rather than from intent. It records what is there, what is deliberately absent, and the reasoning for each — so the next person does not have to re-derive it, and so nobody claims more than the system does.

The short version

Layer State
LLM gateway, multi-provider built
Hybrid RAG with ACL pre-filter built, used by 2 of 9 agents
Working memory (within one answer) built
Conversation memory (across turns) built, browser-side, token-budgeted
Long-term / semantic memory store built and migrated; Wiring in Progress

1. The model layer

internal/gateway/ speaks one wire protocol — OpenAI chat-completions — which is also what Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM and Ollama serve. Supporting six vendors is one implementation and six base URLs.

Three tiers (fast / balanced / deep) selected per agent by its reasoning: value. Tier → model is a deployment knob; tier → effort is not, because "deep" must mean the same thing in every deployment.

Failover is per provider, because a free tier's ceiling is per provider: a second key is a second budget. It is refused on terminal errors (a rejected credential fails the same way everywhere) and on a conversation that has already called a tool, because provider-specific metadata on that call cannot travel. When a rate limit lands mid-run, the loop re-runs the whole turn on the next provider from the original question — unless the run carried a confirmation, in which case it is never replayed, because re-running re-runs its tools and a write twice is two shifts assigned.

2. Retrieval

Used by control-center-agent and krow-workforce-agent only. The other seven answer from SQL tools. That split is the design: "how many open positions" is a query, not a retrieval problem, and routing it through a corpus would make a precise answer approximate.

What makes it more than a vector lookup:

  • Hybrid. Dense and BM25, fused with Reciprocal Rank Fusion. Not dense-only: semantic search is weak on exact terms and this corpus is full of them — shift codes, certification names, venues. Not keyword-only either, which is the failure a deployment with no embedding credential ships by accident.
  • Permission as a PRE-filter. The same predicate goes into both queries' WHERE clauses. Rank first and drop afterwards and forbidden rows leak through the shape of what is left: a short result set, a top-3 with a hole in it, a confidence that tracks documents the caller cannot see.
  • Honest degradation. With no embedder, results come back marked "no embedder is configured; these results are keyword-only" rather than quietly worse.
  • Citable. Every chunk carries the ids to point back at it.
  • Injection boundary. Chunks go into a delimited block in a user message, never the system prompt, and document text cannot close its own fence.

DefaultK is 4. It was 8; a retrieval block is re-sent on every model call of a run, and the deployment's ceiling is 8,000 tokens a minute.

3. Memory

Working memory — within one run the loop accumulates tool calls and results across model calls. Discarded when the run ends.

Conversation memory — components/ai-assistant/recall.ts. The last few exchanges travel with the question.

It lives in the browser because the API takes an input and no message list, and agent_runs records each run independently. That is right for an API and wrong for a panel that reads as a conversation. The proper fix is a messages array on the run request; recall.ts is shaped like that future field so the swap is a deletion rather than a rewrite.

Bounded in tokens, not turns. Turns are not a unit of cost — three short exchanges are nothing, three carrying a table each is a question that no longer fits. So: a 600-token ceiling (about 7% of a minute's budget), each turn clipped to 400 characters, eviction oldest-first because dropping the most recent exchange drops the one the follow-up is about. The transcript is fenced and labelled as data: an earlier answer is the model's own words, but an earlier QUESTION is the reader's, and a reader can type anything.

Long-term memory: the store exists, the behaviour does not.

internal/memory and migration 000017_agent_memories are built and applied in production. Nothing in loop.go reads or writes a memory yet, so no answer has ever been shaped by one. State it that way: the foundation is deployed, the feature is not switched on. "We have long-term memory" is not yet true of anything a user would experience.

What the store is, and why it is mostly provenance:

  • Org-scoped. A memory written while one recruiter worked is available to the next, because a workspace's view of its own hiring should not reset per seat. org_id is NOT NULL and the predicate is in every read.
  • It remembers both kinds — operational facts and observations about named people. The second is why the rest of this list exists. "This applicant seemed unreliable", written automatically and read into a later hiring answer, is profiling under GDPR and is the artefact an employment claim would be built on.
  • Every personal memory names its subject, refused at the door rather than defaulted. A memory about somebody that names nobody cannot be shown to them or erased for them, which is the single property that makes holding it defensible. Held() answers a subject access request and Forget() an erasure, each in one statement; Forget is a soft delete, so the erasure is itself on the record.
  • Provenance on every row — the author (an agent's inference or a person's note) and the run that wrote it, so "why did it say that" survives memory entering the picture, and an inference is never read back as though a person had written it.
  • Everything expires, ninety days by default. A hiring workspace changes shape over a quarter, and a stale fact read as a current one is worse than no memory at all.
  • A memory cannot decide. The block reaching the model is fenced and labelled on the same terms as a retrieved document, states the origin of each line, and says plainly that a memory is never a reason on its own to accept or reject anybody. That sentence is pinned by a test, so an edit cannot quietly drop it.

Recall is semantic where an embedder exists and newest-first where it does not, and says which happened rather than silently returning recency. Five memories by default: this competes for the same prompt as the tool catalogue and the retrieved block, against a ceiling of 8,000 tokens a minute.

What is left, and it is the hard part. Wiring the read into loop.go is small. The write trigger is not: automatic means something judges what is worth remembering, and a bad judge fills the table with noise that then shapes every answer after it. That decision is open.

Conversation threading is still absent, separately, and the schema still says why:

conversations — a run is one turn. Threading runs into a conversation is a surface-layer concern and no surface asks for it yet; adding the column later is trivial, and inventing the semantics now is not.

4. What to verify before trusting retrieval in a demo

Configuration being present does not mean a corpus was ingested:

SELECT count(*) AS chunks, count(embedding) AS embedded FROM knowledge_chunks;
  • 0 — RAG is wired and has nothing to retrieve
  • chunks but no embeddings — keyword-only; the dense half never runs
  • both non-zero — working as designed

5. Claims this document supports

Hybrid retrieval with permission pre-filtering, multi-provider failover, per-run budgets, a full trajectory per answer, human confirmation before any write, conversation memory with a token budget, a bilingual interface, and a long-term memory store that is built, migrated and auditable by subject.

It does NOT support:

  • "the agent remembers across sessions" — the store is not wired, so nothing does yet;
  • "complete memory";
  • "everything is retrieval-backed" — two of nine agents are, on purpose.

The first of those is the one most likely to be said by accident, because the code and the table both exist. Built is not the same as switched on.