Add the in-cluster embedder, so production retrieval stops being keyword-only
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled

Production had no EMBED_PROVIDER, so every knowledge_chunk carried a null
embedding and a question only matched documents that shared its words. A person
asking about a family emergency got nothing from a document titled "shift cover
and cancellation".

Ollama rather than Voyage: internal/knowledge/embed.go calls it "the default
worth reaching for" — real semantics, no credential, no per-token cost, and no
tenant text leaving the cluster. Voyage needs an API key nobody has issued.

Bounded deliberately. The API pods share this node, so an unbounded model
server is a way to evict them; the memory limit means the kubelet kills the
embedder and nothing else. The 1Gi request is also what keeps it off the second
node, which has 1.2Gi allocatable and could not hold it.

Applied in three stages so nothing was pointed at an embedder that had not
been proven: deploy and pull the model, run reembed with the settings passed as
exec environment — 34 chunks in 11s, which proves connectivity without touching
live config — and only then patch krow-config and restart. Rolling back is
removing four keys and restarting.

Verified after: 55/55 on verify-deploy, and a question with no literal keyword
overlap with the corpus returned the relevant policy documents.

This file is the record of what was applied. It was applied by hand, which is
the same gap the README already admits for migrations — there is no deploy
pipeline, so a manifest in the repository is a description of the cluster
rather than the thing that produces it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
This commit is contained in:
2026-08-29 15:33:59 +05:30
parent 629d97181d
commit 109fc2f1c6

View File

@@ -0,0 +1,61 @@
# Embeddings for the knowledge layer.
#
# Production retrieval was keyword-only: no EMBED_PROVIDER, so every chunk had
# a null embedding and a question only matched documents sharing its words.
# Ollama is what internal/knowledge/embed.go calls "the default worth reaching
# for" — real semantics, no credential, no per-token cost, and no tenant text
# leaving the cluster.
#
# Bounded on purpose. The API pods share this node, so an unbounded model
# server is a way to evict them; the limit means the kubelet kills this and
# nothing else. The request is what keeps it off the 1.2Gi node, where it
# would not fit.
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: ollama-models, namespace: krow }
spec:
accessModes: [ReadWriteOnce]
resources: { requests: { storage: 4Gi } }
---
apiVersion: apps/v1
kind: Deployment
metadata: { name: ollama, namespace: krow }
spec:
replicas: 1
selector: { matchLabels: { app: ollama } }
strategy: { type: Recreate } # one volume, one writer
template:
metadata: { labels: { app: ollama } }
spec:
containers:
- name: ollama
image: ollama/ollama:0.33.1
ports: [{ containerPort: 11434, name: http }]
env:
- { name: OLLAMA_HOST, value: "0.0.0.0:11434" }
# One model, kept resident: reloading it per request would make
# every retrieval pay the load cost.
- { name: OLLAMA_KEEP_ALIVE, value: "24h" }
- { name: OLLAMA_MAX_LOADED_MODELS, value: "1" }
resources:
requests: { memory: "1Gi", cpu: "250m" }
limits: { memory: "3Gi", cpu: "2" }
volumeMounts: [{ name: models, mountPath: /root/.ollama }]
readinessProbe:
httpGet: { path: /api/version, port: http }
initialDelaySeconds: 5
periodSeconds: 10
livenessProbe:
httpGet: { path: /api/version, port: http }
initialDelaySeconds: 30
periodSeconds: 30
volumes:
- name: models
persistentVolumeClaim: { claimName: ollama-models }
---
apiVersion: v1
kind: Service
metadata: { name: ollama, namespace: krow }
spec:
selector: { app: ollama }
ports: [{ port: 11434, targetPort: http, name: http }]