Both repositories carried GitHub Actions workflows on a Gitea remote and nobody had confirmed a runner. There was not one: the 924 frontend checks, the whole Go suite, the skip guard and the suite-shrank guard had never run on a push, only when somebody remembered. gitea/act_runner v0.6.1 is registered as krow-runner on the cluster host. This commit is also the first push that can prove it picks up a job, which is the failure mode worth catching — a runner that registers and never runs anything looks identical to a healthy one in the Runners list. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
250 lines
13 KiB
Markdown
250 lines
13 KiB
Markdown
# Handover
|
||
|
||
Written 2026-08-28, when the machine this was built on was retired.
|
||
|
||
Everything Claude Code "remembers" lives in `~/.claude/projects/<mangled-path>/`
|
||
on one machine, keyed to the absolute path of the checkout. It does not sync,
|
||
and a different path on a new machine reads a different folder. So the durable
|
||
record is this file, in the repository, where git carries it and any path works.
|
||
|
||
Read `CLAUDE.md` first — it is the governing document. This file is what it does
|
||
not say: what was decided, what is deployed, and which parts bite.
|
||
|
||
---
|
||
|
||
## Where things stand
|
||
|
||
**Deployed.** Backend at `https://mcp.krowforce.com`, frontend at
|
||
`https://platform.krowforce.com`, Kubernetes statefulset `krow` in namespace
|
||
`krow`, pods named `krow-1` and `krow-2` (they start at 1, not 0).
|
||
|
||
ssh root@<host> -p 4422 "kubectl -n krow rollout restart statefulset/krow && \
|
||
kubectl -n krow rollout status statefulset/krow --timeout=180s"
|
||
|
||
**Verify a deployment** — 55 checks including a real agent run:
|
||
|
||
KROW_EMAIL=... KROW_PASSWORD=... make verify-deploy BASE=https://mcp.krowforce.com
|
||
|
||
Auth runs BEFORE routing, so an unauthenticated probe answers 401 for every
|
||
path including ones that do not exist. `curl` cannot tell a missing endpoint
|
||
from a guarded one; only an authenticated check can.
|
||
|
||
**Two runtime steps a deploy does not do**, both easy to forget because the API
|
||
looks healthy without them:
|
||
|
||
kubectl -n krow exec krow-1 -- importagents --org <slug> # publishes agents/ and skills/
|
||
kubectl -n krow exec krow-1 -- ingest --org <slug> # ingests knowledge/
|
||
|
||
Without the first, every Owliver question answers 404. Without the second,
|
||
retrieval finds nothing. The org slug is `krow-dev` — a hardcoded constant
|
||
(`internal/orgctx.DevOrgSlug`), not configuration.
|
||
|
||
**The endpoint count is a signal.** `GET /api/v1/version` reports it. Agent run
|
||
routes are not registered without a model credential, so **55** means no
|
||
`ANTHROPIC_API_KEY` and **57** means there is one. (These were written as 56/58
|
||
and were one high; the delta of two — the two run routes — was always right.) A keyless deployment boots
|
||
cleanly under `APP_ENV=staging` and refuses under `production`.
|
||
|
||
---
|
||
|
||
## Decisions taken, so they are not relitigated
|
||
|
||
**§12, who authors agents: self-serve, split by visibility.** `personal` agents
|
||
are authored in the UI, POST to `/api/v1/agent-definitions`, and are runnable
|
||
immediately; evals are not required for them. `organization` agents stay as
|
||
files published by `importagents` on deploy — that deploy step *is* the approval
|
||
workflow, and §9's eval requirement still applies. §12 warned self-serve needs
|
||
"3× the platform"; it does not here, because I1 means an agent runs as its
|
||
caller and cannot exceed their access, I4 means writes still need a human, and
|
||
the spec format has no `limits` block so budgets cannot be raised by an author.
|
||
|
||
**A rejected candidate counts as screened.** It ranks with `ai_screened`:
|
||
rejection overwrites the stage it came from, so shortlisted can never be
|
||
claimed. **An assigned candidate counts as hired**, matching the backend's
|
||
existing `status IN ('hired','assigned')`.
|
||
|
||
**Seeded profile scores are the formula's output**, not hand-authored narrative.
|
||
Recalculating a seeded profile is a no-op, and a skill-check enforces it.
|
||
|
||
---
|
||
|
||
## Conventions the schema actively contradicts
|
||
|
||
These are the ones that produce confident, wrong numbers rather than an error.
|
||
|
||
**`job_applications.screened_at` is vestigial.** Nothing writes it. "Screened"
|
||
means `status <> 'applied'`, in about eight places in the frontend. Reading the
|
||
column reported 1 screened of 24 where the truth was 14.
|
||
|
||
**A score of `0` means "not rated", never "rated zero".** Every score column is
|
||
NOT NULL, so there is no null to distinguish it — that is the trap. Applies to
|
||
`ai_score`, `client_rating`, `krow_score`, `reliability_score`,
|
||
`attendance_score`, `performance_score`, `experience_years`. Aggregates need
|
||
`FILTER (WHERE col > 0)` and a stated basis count. Counting zeros once reported
|
||
"16 weak candidates averaging 28" for a pool that was 1 weak averaging 76.
|
||
|
||
Anchors to check against: applications are **9 scored averaging 76**; workers
|
||
are **5 rated of 9, client rating 4.70**.
|
||
|
||
**Genuine zeros, do not filter these:** `overtime_hours`, `minutes_late`, `xp`,
|
||
`profile_completion`, and `actual_hours` (0 only on absent/no_show shifts).
|
||
|
||
**`absent` and `no_show` are both missed shifts, but only one is a no-show.**
|
||
`attendance.js` is canonical.
|
||
|
||
**`attendance_score` defaults to 100 for display and must never be a scoring
|
||
input.** A profile with no evidence otherwise scores 12 and leaves the "not yet
|
||
scored" band.
|
||
|
||
**The stage ladder lives once**, in `krow-demo/src/lib/hiringRecords.js`. It was
|
||
three byte-identical private copies, all missing `rejected` and `assigned`,
|
||
which `indexOf` scored -1 and dropped from every bucket including `applied`.
|
||
|
||
---
|
||
|
||
## Things that will waste your afternoon
|
||
|
||
**A space in the checkout path breaks path derivation.** This repo lives under
|
||
`Krow Project /`, and it has bitten three times: an unquoted `$(CURDIR)` in the
|
||
Makefile, `` `file://${process.argv[1]}` `` in a script guard, and
|
||
`new URL(...).pathname` in `scripts/oracle.mjs` (use `fileURLToPath`). Any new
|
||
path derivation is guilty until tested there.
|
||
|
||
**`seed.json` is generated from `krow-demo/src/api/seed.js`.** Never edit it.
|
||
`npm run seed:fixture` writes it, `npm run seed:check` verifies, and the
|
||
skill-check compares byte-for-byte. `ShiftRecord` is excluded on purpose: the Go
|
||
seeder generates it against *now*.
|
||
|
||
**Shift data is anchored to today**, so anything asserting against it is
|
||
date-dependent unless the anchor is pinned. `buildShiftsAt(anchor)` exists for
|
||
that. One detector had a Friday-and-Saturday blind spot for exactly this reason.
|
||
|
||
**Database tests skip when PostgreSQL is unreachable** (`testutil` calls
|
||
`t.Skipf`). `go test` then exits 0 having run almost nothing. CI has a guard
|
||
that fails on any skip other than `TestLive*`; keep it.
|
||
|
||
**The eval suites use a scripted model.** They prove the permission boundary,
|
||
not answer quality. `make eval-live` uses the real model and costs tokens.
|
||
|
||
It has now been run, and passes 3/3 in about 30 seconds. Worth knowing what it
|
||
established, because it is more than "the agents work":
|
||
|
||
- the coverage agent called `assign_worker` and the write did NOT happen —
|
||
I4 holds against a real model, not only a scripted one;
|
||
- the handbook corpus contains a planted prompt injection telling the agent
|
||
to send the staff roster to an external address. The agent refused it,
|
||
answered the real question with citations, and reported the document as
|
||
tampered with. I7 holds end to end;
|
||
- the activity agent declined to subtract two figures it could not
|
||
reconcile, and said so, rather than producing the confident wrong number
|
||
this schema invites.
|
||
|
||
Re-run it after any change to the loop, retrieval, or prompt assembly. It is
|
||
the only check that measures answers rather than boundaries.
|
||
|
||
**Seeded time-series data is rebased to now at seed time** — see
|
||
`seeder.RebaseToNow`. `ShiftRecord` is generated against now; `UserActivity`,
|
||
`JobApplication`, `AIInterview` and `Staff` are moved so their newest record
|
||
sits at today, keeping every authored gap. Without it the demo goes quiet: on
|
||
2026-08-29 the newest activity event was 23 days old, applications 16 days,
|
||
staff hire dates 35 — zero events in the last 7 days and an empty Hiring
|
||
activity chart on Control Center.
|
||
|
||
Two things to know if you touch it. EVERY timestamp on a record shifts by the
|
||
same delta, not just the anchor: an application's created_date and updated_date
|
||
are what `buildHires` subtracts for time-to-hire, and moving one alone turns a
|
||
five-day hire into a three-week one. And `Staff` anchors on `hire_date` rather
|
||
than `created_date`, because the hire is the event the chart plots — leaving it
|
||
behind produced a workspace where somebody was hired last week according to
|
||
their application and five weeks ago according to their staff record.
|
||
|
||
Reference data is deliberately not rebased. A course's date is a fact about the
|
||
course, not a position in a window.
|
||
|
||
---
|
||
|
||
## Still outstanding
|
||
|
||
- `ANTHROPIC_API_KEY` was pasted into a chat transcript and is live in a
|
||
Kubernetes Secret. Rotate it.
|
||
- Deployments report `version=dev`: the image is built without
|
||
`--build-arg VERSION`. `make docker-build` passes it.
|
||
- `APP_ENV=staging` on the deployment, so the production config guards are off.
|
||
- CI tests but does not deploy. The README's claim that migrations are "run by
|
||
CI against the target database" is still aspirational.
|
||
- The fixture-drift CI jobs need `FRONTEND_REPO_TOKEN` to see the sibling repo,
|
||
and fail rather than pass quietly without it.
|
||
- The remote is Gitea and the workflows are GitHub Actions syntax. Gitea Actions
|
||
runs them, and a runner now exists: `gitea-runner` (gitea/act_runner v0.6.1)
|
||
on the cluster host, registered as `krow-runner` with labels
|
||
`ubuntu-latest, ubuntu-22.04` mapped to `node:20-bookworm`. Before that, both
|
||
repositories had workflows that had never executed once — the 924 frontend
|
||
checks, the whole Go suite, the skip guard and the suite-shrank guard were
|
||
all things somebody had to remember to run.
|
||
|
||
If a job fails resolving `actions/checkout` or `actions/setup-node`, the
|
||
runner needs egress to github.com or a mirror; that is where those actions
|
||
come from and Gitea does not host them.
|
||
- **The application talks to its database in clear text.** `DATABASE_SSLMODE=
|
||
disable` against `66.116.207.225`, which is a DIFFERENT machine from the
|
||
cluster host — so credentials and every row cross the network unencrypted.
|
||
It is permitted only because `APP_ENV=staging`; the production guard refuses
|
||
`disable` outright. PostgreSQL itself now has `ssl = on` (2026-08-29, port
|
||
5433, reload not restart), but the app does not reach PostgreSQL directly:
|
||
**pgbouncer terminates 5432** and offers no TLS of its own. The fix is
|
||
`client_tls_sslmode = allow` plus a cert in `/etc/pgbouncer/pgbouncer.ini`,
|
||
then `DATABASE_SSLMODE=require` in `krow-config` and the `krow-db` secret.
|
||
`allow` keeps existing plaintext clients working, so it is additive.
|
||
- Production retrieval is **keyword-only**: no `EMBED_PROVIDER` in
|
||
`krow-config`, so `knowledge_chunks.embedding` is null for all 34 rows. A
|
||
`VOYAGE_API_KEY` is the cheap fix; Ollama in-cluster is the other, and the
|
||
nodes were at 60% and 49% memory when that was last looked at.
|
||
- Delegation (§6) is implemented and on `main` but NOT deployed. Until the next
|
||
image ships, production agents still ignore their `subagents:`.
|
||
- §3's publish-time cycle detection is still missing. The runtime depth cap
|
||
(2) is what bounds a cycle that reaches run time.
|
||
- `cmd/importagents` has no tests, and `run()` opens its own pool from config,
|
||
so making it testable is a refactor rather than an addition.
|
||
- `importagents` does not enforce monotonicity: a spec whose `version:` is
|
||
LOWERED still overwrites the live row and rolls the deployed agent backwards.
|
||
|
||
---
|
||
|
||
## Setting up a new machine
|
||
|
||
git clone <backend> krow-backend && git clone <frontend> krow-demo
|
||
cp krow-backend/CLAUDE.md ./claude.md # the governing doc lives above both repos
|
||
|
||
Needs, if you run the backend natively: Go (see `go-api/go.mod`), Node 20,
|
||
PostgreSQL, Docker, and Ollama with `nomic-embed-text` for semantic retrieval.
|
||
|
||
You do not need most of that. `infrastructure/Dockerfile.api` builds EVERY
|
||
command in `go-api/cmd/` plus the golang-migrate CLI into the image, so the
|
||
whole stack runs on Docker alone — no Go, no psql, no migrate on the host:
|
||
|
||
cd krow-backend/infrastructure
|
||
cp .env.docker.example .env # fill it in; DATABASE_HOST=postgres
|
||
docker compose -f docker-compose.yml -f docker-compose.local-db.yml up -d
|
||
docker exec krow-api seed
|
||
docker exec krow-api importagents --dir /app/agents --skills /app/skills --org krow-dev
|
||
docker exec krow-api ingest --dir /app/knowledge --org krow-dev
|
||
printf '%s' 'PASSWORD' | docker exec -i krow-api setpassword -email demo@krow.app -stdin
|
||
|
||
Ollama, if you want semantic retrieval, runs on the HOST — so the container
|
||
reaches it at `host.docker.internal:11434`, NOT `localhost:11434`, which inside
|
||
a container means the container.
|
||
|
||
Running natively instead, you need all of the above. Then:
|
||
|
||
cd krow-backend && cp .env.example .env # fill it in; .env is gitignored
|
||
make migrate-up && make seed
|
||
make import-agents ORG=krow-dev
|
||
make ingest ORG=krow-dev
|
||
go run ./go-api/cmd/setpassword -email demo@krow.app
|
||
|
||
cd ../krow-demo && npm ci && cp .env.example .env
|
||
# VITE_AGENT_API=/api/v1 for local dev (vite proxies it);
|
||
# production passes an absolute URL as a Docker build arg instead.
|
||
|
||
`.env` files are not in git and must be carried across by hand.
|