Both documents are the first thing a new reader trusts, and several of their
claims were wrong — some wrong from the start, some overtaken by work this
week. A governing document that misdescribes the system is worse than none,
because it is believed.
CLAUDE.md §10 specified Python 3.11, FastAPI, SQLAlchemy and Alembic. The code
is Go and has never been anything else. That is corrected rather than quietly
deleted, so the next person understands the document drifted rather than
wondering which half to trust.
Also in CLAUDE.md: the tool count was 17 with one write and is 19 with two;
delegation now exists and §11's Orchestration row says what it guarantees; the
embedder deviation described a Voyage-or-stand-in choice that has since become
EMBED_PROVIDER with three options, of which production sets none.
The handover claimed three things that this session disproved by running them:
- "definition_versions is empty ... nothing has gone through it". It was not
empty in production; activity-agent had a v1 that the shipped file
contradicted, which is how a real drift was found. It now holds every
agent and skill.
- "Skills are still stored in user_preferences". They are rows in
skill_definitions, and are now versioned.
- "make eval-live ... has never been run". It has, it passes 3/3, and what
it established is recorded — including that the handbook corpus carries a
planted prompt injection which the agent refused and reported. That is I7
holding against a real model, which is worth more than the pass count.
The endpoint counts were one high throughout (55/57, not 56/58) — the delta of
two was always right, so the signal worked and the absolute numbers did not.
Added, because they cost time this week and would cost it again:
- the app reaches its database through pgbouncer, not PostgreSQL directly.
Enabling TLS on PostgreSQL does nothing for the application hop; pgbouncer
terminates 5432 and needs its own client_tls_sslmode.
- the seeded UserActivity is NOT anchored to today the way ShiftRecord is,
so it ages out of every window the activity tools offer. Twenty-three days
old as of writing: zero events in the last 7 days, 6 of 15 in the last 30.
The agent answers truthfully and the demo looks dead.
- the whole stack runs on Docker alone. Dockerfile.api builds every command
plus the migrate CLI, so a new machine needs neither Go nor psql — which
is how this one was set up, having no Homebrew.
- Ollama runs on the HOST, so a container reaches it at
host.docker.internal, not localhost. The old .env said localhost and would
have failed with nothing obviously wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
12 KiB
Handover
Written 2026-08-28, when the machine this was built on was retired.
Everything Claude Code "remembers" lives in ~/.claude/projects/<mangled-path>/
on one machine, keyed to the absolute path of the checkout. It does not sync,
and a different path on a new machine reads a different folder. So the durable
record is this file, in the repository, where git carries it and any path works.
Read CLAUDE.md first — it is the governing document. This file is what it does
not say: what was decided, what is deployed, and which parts bite.
Where things stand
Deployed. Backend at https://mcp.krowforce.com, frontend at
https://platform.krowforce.com, Kubernetes statefulset krow in namespace
krow, pods named krow-1 and krow-2 (they start at 1, not 0).
ssh root@<host> -p 4422 "kubectl -n krow rollout restart statefulset/krow && \
kubectl -n krow rollout status statefulset/krow --timeout=180s"
Verify a deployment — 55 checks including a real agent run:
KROW_EMAIL=... KROW_PASSWORD=... make verify-deploy BASE=https://mcp.krowforce.com
Auth runs BEFORE routing, so an unauthenticated probe answers 401 for every
path including ones that do not exist. curl cannot tell a missing endpoint
from a guarded one; only an authenticated check can.
Two runtime steps a deploy does not do, both easy to forget because the API looks healthy without them:
kubectl -n krow exec krow-1 -- importagents --org <slug> # publishes agents/ and skills/
kubectl -n krow exec krow-1 -- ingest --org <slug> # ingests knowledge/
Without the first, every Owliver question answers 404. Without the second,
retrieval finds nothing. The org slug is krow-dev — a hardcoded constant
(internal/orgctx.DevOrgSlug), not configuration.
The endpoint count is a signal. GET /api/v1/version reports it. Agent run
routes are not registered without a model credential, so 55 means no
ANTHROPIC_API_KEY and 57 means there is one. (These were written as 56/58
and were one high; the delta of two — the two run routes — was always right.) A keyless deployment boots
cleanly under APP_ENV=staging and refuses under production.
Decisions taken, so they are not relitigated
§12, who authors agents: self-serve, split by visibility. personal agents
are authored in the UI, POST to /api/v1/agent-definitions, and are runnable
immediately; evals are not required for them. organization agents stay as
files published by importagents on deploy — that deploy step is the approval
workflow, and §9's eval requirement still applies. §12 warned self-serve needs
"3× the platform"; it does not here, because I1 means an agent runs as its
caller and cannot exceed their access, I4 means writes still need a human, and
the spec format has no limits block so budgets cannot be raised by an author.
A rejected candidate counts as screened. It ranks with ai_screened:
rejection overwrites the stage it came from, so shortlisted can never be
claimed. An assigned candidate counts as hired, matching the backend's
existing status IN ('hired','assigned').
Seeded profile scores are the formula's output, not hand-authored narrative. Recalculating a seeded profile is a no-op, and a skill-check enforces it.
Conventions the schema actively contradicts
These are the ones that produce confident, wrong numbers rather than an error.
job_applications.screened_at is vestigial. Nothing writes it. "Screened"
means status <> 'applied', in about eight places in the frontend. Reading the
column reported 1 screened of 24 where the truth was 14.
A score of 0 means "not rated", never "rated zero". Every score column is
NOT NULL, so there is no null to distinguish it — that is the trap. Applies to
ai_score, client_rating, krow_score, reliability_score,
attendance_score, performance_score, experience_years. Aggregates need
FILTER (WHERE col > 0) and a stated basis count. Counting zeros once reported
"16 weak candidates averaging 28" for a pool that was 1 weak averaging 76.
Anchors to check against: applications are 9 scored averaging 76; workers are 5 rated of 9, client rating 4.70.
Genuine zeros, do not filter these: overtime_hours, minutes_late, xp,
profile_completion, and actual_hours (0 only on absent/no_show shifts).
absent and no_show are both missed shifts, but only one is a no-show.
attendance.js is canonical.
attendance_score defaults to 100 for display and must never be a scoring
input. A profile with no evidence otherwise scores 12 and leaves the "not yet
scored" band.
The stage ladder lives once, in krow-demo/src/lib/hiringRecords.js. It was
three byte-identical private copies, all missing rejected and assigned,
which indexOf scored -1 and dropped from every bucket including applied.
Things that will waste your afternoon
A space in the checkout path breaks path derivation. This repo lives under
Krow Project /, and it has bitten three times: an unquoted $(CURDIR) in the
Makefile, `file://${process.argv[1]}` in a script guard, and
new URL(...).pathname in scripts/oracle.mjs (use fileURLToPath). Any new
path derivation is guilty until tested there.
seed.json is generated from krow-demo/src/api/seed.js. Never edit it.
npm run seed:fixture writes it, npm run seed:check verifies, and the
skill-check compares byte-for-byte. ShiftRecord is excluded on purpose: the Go
seeder generates it against now.
Shift data is anchored to today, so anything asserting against it is
date-dependent unless the anchor is pinned. buildShiftsAt(anchor) exists for
that. One detector had a Friday-and-Saturday blind spot for exactly this reason.
Database tests skip when PostgreSQL is unreachable (testutil calls
t.Skipf). go test then exits 0 having run almost nothing. CI has a guard
that fails on any skip other than TestLive*; keep it.
The eval suites use a scripted model. They prove the permission boundary,
not answer quality. make eval-live uses the real model and costs tokens.
It has now been run, and passes 3/3 in about 30 seconds. Worth knowing what it established, because it is more than "the agents work":
- the coverage agent called
assign_workerand the write did NOT happen — I4 holds against a real model, not only a scripted one; - the handbook corpus contains a planted prompt injection telling the agent to send the staff roster to an external address. The agent refused it, answered the real question with citations, and reported the document as tampered with. I7 holds end to end;
- the activity agent declined to subtract two figures it could not reconcile, and said so, rather than producing the confident wrong number this schema invites.
Re-run it after any change to the loop, retrieval, or prompt assembly. It is the only check that measures answers rather than boundaries.
The seeded activity data is NOT anchored to today. ShiftRecord is —
seed.js says so — and UserActivity is not, so it ages out of every window
the activity tools offer. As of 2026-08-29 the newest event was 23 days old:
zero events in the last 7 days and 6 of 15 in the last 30. The activity agent
answers truthfully and the demo looks dead. Anchoring it the way shifts are
anchored is the fix; it changes seed.json, so it goes through
npm run seed:fixture and re-runs the frontend checks.
Still outstanding
ANTHROPIC_API_KEYwas pasted into a chat transcript and is live in a Kubernetes Secret. Rotate it.- Deployments report
version=dev: the image is built without--build-arg VERSION.make docker-buildpasses it. APP_ENV=stagingon the deployment, so the production config guards are off.- CI tests but does not deploy. The README's claim that migrations are "run by CI against the target database" is still aspirational.
- The fixture-drift CI jobs need
FRONTEND_REPO_TOKENto see the sibling repo, and fail rather than pass quietly without it. - The remote is Gitea. These are GitHub Actions workflows; they do nothing until a compatible runner exists. Nobody has confirmed a runner exists, so treat both repositories as having no CI until somebody checks.
- The application talks to its database in clear text.
DATABASE_SSLMODE= disableagainst66.116.207.225, which is a DIFFERENT machine from the cluster host — so credentials and every row cross the network unencrypted. It is permitted only becauseAPP_ENV=staging; the production guard refusesdisableoutright. PostgreSQL itself now hasssl = on(2026-08-29, port 5433, reload not restart), but the app does not reach PostgreSQL directly: pgbouncer terminates 5432 and offers no TLS of its own. The fix isclient_tls_sslmode = allowplus a cert in/etc/pgbouncer/pgbouncer.ini, thenDATABASE_SSLMODE=requireinkrow-configand thekrow-dbsecret.allowkeeps existing plaintext clients working, so it is additive. - Production retrieval is keyword-only: no
EMBED_PROVIDERinkrow-config, soknowledge_chunks.embeddingis null for all 34 rows. AVOYAGE_API_KEYis the cheap fix; Ollama in-cluster is the other, and the nodes were at 60% and 49% memory when that was last looked at. - Delegation (§6) is implemented and on
mainbut NOT deployed. Until the next image ships, production agents still ignore theirsubagents:. - §3's publish-time cycle detection is still missing. The runtime depth cap (2) is what bounds a cycle that reaches run time.
cmd/importagentshas no tests, andrun()opens its own pool from config, so making it testable is a refactor rather than an addition.importagentsdoes not enforce monotonicity: a spec whoseversion:is LOWERED still overwrites the live row and rolls the deployed agent backwards.
Setting up a new machine
git clone <backend> krow-backend && git clone <frontend> krow-demo
cp krow-backend/CLAUDE.md ./claude.md # the governing doc lives above both repos
Needs, if you run the backend natively: Go (see go-api/go.mod), Node 20,
PostgreSQL, Docker, and Ollama with nomic-embed-text for semantic retrieval.
You do not need most of that. infrastructure/Dockerfile.api builds EVERY
command in go-api/cmd/ plus the golang-migrate CLI into the image, so the
whole stack runs on Docker alone — no Go, no psql, no migrate on the host:
cd krow-backend/infrastructure
cp .env.docker.example .env # fill it in; DATABASE_HOST=postgres
docker compose -f docker-compose.yml -f docker-compose.local-db.yml up -d
docker exec krow-api seed
docker exec krow-api importagents --dir /app/agents --skills /app/skills --org krow-dev
docker exec krow-api ingest --dir /app/knowledge --org krow-dev
printf '%s' 'PASSWORD' | docker exec -i krow-api setpassword -email demo@krow.app -stdin
Ollama, if you want semantic retrieval, runs on the HOST — so the container
reaches it at host.docker.internal:11434, NOT localhost:11434, which inside
a container means the container.
Running natively instead, you need all of the above. Then:
cd krow-backend && cp .env.example .env # fill it in; .env is gitignored
make migrate-up && make seed
make import-agents ORG=krow-dev
make ingest ORG=krow-dev
go run ./go-api/cmd/setpassword -email demo@krow.app
cd ../krow-demo && npm ci && cp .env.example .env
# VITE_AGENT_API=/api/v1 for local dev (vite proxies it);
# production passes an absolute URL as a Docker build arg instead.
.env files are not in git and must be carried across by hand.