Files
krow_backend/docs/handover.md
Suriyakumarvijayanayagam c74fe7e074 Add an OpenAI-compatible gateway, so the model provider is a config value
The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.

`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.

`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.

THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.

Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.

Three configuration mistakes are refused at startup rather than at runtime:

  - MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
    endpoint and ignores the field, so this is a deployment that believes it
    switched providers and did not — every run still goes to Anthropic and is
    still billed there, with nothing in the logs to say so. Cost is the whole
    reason this change exists, and that is the one mistake that silently
    defeats it.
  - An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
  - A production deployment with no credential — except against localhost,
    which needs none, and demanding one would make the free local path
    impossible to configure.

reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.

`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.

Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.

NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-01 11:47:53 +05:30

325 lines
16 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Handover
Written 2026-08-28, when the machine this was built on was retired.
Everything Claude Code "remembers" lives in `~/.claude/projects/<mangled-path>/`
on one machine, keyed to the absolute path of the checkout. It does not sync,
and a different path on a new machine reads a different folder. So the durable
record is this file, in the repository, where git carries it and any path works.
Read `CLAUDE.md` first — it is the governing document. This file is what it does
not say: what was decided, what is deployed, and which parts bite.
---
## Where things stand
**Deployed.** Backend at `https://mcp.krowforce.com`, frontend at
`https://platform.krowforce.com`, Kubernetes statefulset `krow` in namespace
`krow`, pods named `krow-1` and `krow-2` (they start at 1, not 0).
ssh root@<host> -p 4422 "kubectl -n krow rollout restart statefulset/krow && \
kubectl -n krow rollout status statefulset/krow --timeout=180s"
**Verify a deployment** — 55 checks including a real agent run:
KROW_EMAIL=... KROW_PASSWORD=... make verify-deploy BASE=https://mcp.krowforce.com
Auth runs BEFORE routing, so an unauthenticated probe answers 401 for every
path including ones that do not exist. `curl` cannot tell a missing endpoint
from a guarded one; only an authenticated check can.
**Two runtime steps a deploy does not do**, both easy to forget because the API
looks healthy without them:
kubectl -n krow exec krow-1 -- importagents --org <slug> # publishes agents/ and skills/
kubectl -n krow exec krow-1 -- ingest --org <slug> # ingests knowledge/
Without the first, every Owliver question answers 404. Without the second,
retrieval finds nothing. The org slug is `krow-dev` — a hardcoded constant
(`internal/orgctx.DevOrgSlug`), not configuration.
**The endpoint count is a signal.** `GET /api/v1/version` reports it. Agent run
routes are not registered without a model credential, so **55** means no
`ANTHROPIC_API_KEY` and **57** means there is one. (These were written as 56/58
and were one high; the delta of two — the two run routes — was always right.) A keyless deployment boots
cleanly under `APP_ENV=staging` and refuses under `production`.
---
## Decisions taken, so they are not relitigated
**§12, who authors agents: self-serve, split by visibility.** `personal` agents
are authored in the UI, POST to `/api/v1/agent-definitions`, and are runnable
immediately; evals are not required for them. `organization` agents stay as
files published by `importagents` on deploy — that deploy step *is* the approval
workflow, and §9's eval requirement still applies. §12 warned self-serve needs
"3× the platform"; it does not here, because I1 means an agent runs as its
caller and cannot exceed their access, I4 means writes still need a human, and
the spec format has no `limits` block so budgets cannot be raised by an author.
**A rejected candidate counts as screened.** It ranks with `ai_screened`:
rejection overwrites the stage it came from, so shortlisted can never be
claimed. **An assigned candidate counts as hired**, matching the backend's
existing `status IN ('hired','assigned')`.
**Seeded profile scores are the formula's output**, not hand-authored narrative.
Recalculating a seeded profile is a no-op, and a skill-check enforces it.
---
## Conventions the schema actively contradicts
These are the ones that produce confident, wrong numbers rather than an error.
**`job_applications.screened_at` is vestigial.** Nothing writes it. "Screened"
means `status <> 'applied'`, in about eight places in the frontend. Reading the
column reported 1 screened of 24 where the truth was 14.
**A score of `0` means "not rated", never "rated zero".** Every score column is
NOT NULL, so there is no null to distinguish it — that is the trap. Applies to
`ai_score`, `client_rating`, `krow_score`, `reliability_score`,
`attendance_score`, `performance_score`, `experience_years`. Aggregates need
`FILTER (WHERE col > 0)` and a stated basis count. Counting zeros once reported
"16 weak candidates averaging 28" for a pool that was 1 weak averaging 76.
Anchors to check against: applications are **9 scored averaging 76**; workers
are **5 rated of 9, client rating 4.70**.
**Genuine zeros, do not filter these:** `overtime_hours`, `minutes_late`, `xp`,
`profile_completion`, and `actual_hours` (0 only on absent/no_show shifts).
**`absent` and `no_show` are both missed shifts, but only one is a no-show.**
`attendance.js` is canonical.
**`attendance_score` defaults to 100 for display and must never be a scoring
input.** A profile with no evidence otherwise scores 12 and leaves the "not yet
scored" band.
**The stage ladder lives once**, in `krow-demo/src/lib/hiringRecords.js`. It was
three byte-identical private copies, all missing `rejected` and `assigned`,
which `indexOf` scored -1 and dropped from every bucket including `applied`.
---
## Things that will waste your afternoon
**A space in the checkout path breaks path derivation.** This repo lives under
`Krow Project /`, and it has bitten three times: an unquoted `$(CURDIR)` in the
Makefile, `` `file://${process.argv[1]}` `` in a script guard, and
`new URL(...).pathname` in `scripts/oracle.mjs` (use `fileURLToPath`). Any new
path derivation is guilty until tested there.
**`seed.json` is generated from `krow-demo/src/api/seed.js`.** Never edit it.
`npm run seed:fixture` writes it, `npm run seed:check` verifies, and the
skill-check compares byte-for-byte. `ShiftRecord` is excluded on purpose: the Go
seeder generates it against *now*.
**Shift data is anchored to today**, so anything asserting against it is
date-dependent unless the anchor is pinned. `buildShiftsAt(anchor)` exists for
that. One detector had a Friday-and-Saturday blind spot for exactly this reason.
**Database tests skip when PostgreSQL is unreachable** (`testutil` calls
`t.Skipf`). `go test` then exits 0 having run almost nothing. CI has a guard
that fails on any skip other than `TestLive*`; keep it.
**The eval suites use a scripted model.** They prove the permission boundary,
not answer quality. `make eval-live` uses the real model and costs tokens.
It has now been run, and passes 3/3 in about 30 seconds. Worth knowing what it
established, because it is more than "the agents work":
- the coverage agent called `assign_worker` and the write did NOT happen —
I4 holds against a real model, not only a scripted one;
- the handbook corpus contains a planted prompt injection telling the agent
to send the staff roster to an external address. The agent refused it,
answered the real question with citations, and reported the document as
tampered with. I7 holds end to end;
- the activity agent declined to subtract two figures it could not
reconcile, and said so, rather than producing the confident wrong number
this schema invites.
Re-run it after any change to the loop, retrieval, or prompt assembly. It is
the only check that measures answers rather than boundaries.
**Seeded time-series data is rebased to now at seed time** — see
`seeder.RebaseToNow`. `ShiftRecord` is generated against now; `UserActivity`,
`JobApplication`, `AIInterview` and `Staff` are moved so their newest record
sits at today, keeping every authored gap. Without it the demo goes quiet: on
2026-08-29 the newest activity event was 23 days old, applications 16 days,
staff hire dates 35 — zero events in the last 7 days and an empty Hiring
activity chart on Control Center.
Two things to know if you touch it. EVERY timestamp on a record shifts by the
same delta, not just the anchor: an application's created_date and updated_date
are what `buildHires` subtracts for time-to-hire, and moving one alone turns a
five-day hire into a three-week one. And `Staff` anchors on `hire_date` rather
than `created_date`, because the hire is the event the chart plots — leaving it
behind produced a workspace where somebody was hired last week according to
their application and five weeks ago according to their staff record.
Reference data is deliberately not rebased. A course's date is a fact about the
course, not a position in a window.
---
**HTTP_WRITE_TIMEOUT must exceed the deepest agent deadline.** It was 30s in
production while every shipped agent runs at the `balanced` tier, whose
deadline is 60s — so the server aborted the response on any run over half its
allowed time, and the proxy in front answered **502 Bad Gateway**. A gateway
error for something no gateway did, which is why it read as an infrastructure
fault: nginx was innocent and already had `proxy_read_timeout 3600s`.
Streaming hid it. The chat panel uses SSE and survives, so the product looked
healthy while any non-streaming caller — a webhook, a script, an integration —
got 502 on a slow question. Delegation made it routine rather than causing it:
a parent that asks two subagents takes longer than one answering alone.
Production is now 180s, and `config.validateWriteTimeout` refuses a value below
`DeepestAgentDeadline` at startup. NOTE THE ORDERING: that constant is 120s, so
a deployment still carrying the old 30s will now refuse to boot. Patch the
configmap before shipping an image that contains the check.
---
## Changing model provider
The gateway speaks two wire protocols. `anthropic` is the Claude API.
`openai` is the chat-completions shape — and that one is not only OpenAI:
Groq, Gemini's compatibility endpoint, OpenRouter, Together, vLLM and a local
Ollama all serve it, so moving between them is configuration, not code.
```bash
# Groq
MODEL_PROVIDER=openai
MODEL_BASE_URL=https://api.groq.com/openai/v1
MODEL_API_KEY=<key>
MODEL_FAST=llama-3.1-8b-instant
MODEL_BALANCED=openai/gpt-oss-120b
MODEL_DEEP=openai/gpt-oss-120b
# Gemini
MODEL_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
# A model on this machine — no credential at all
MODEL_BASE_URL=http://localhost:11434/v1
```
Four things worth knowing before you do it.
**Set `MODEL_PROVIDER`, not just the base URL.** The anthropic path has one
endpoint and ignores `MODEL_BASE_URL` entirely, so setting the URL alone is a
deployment that believes it has switched providers and has not — every run
still goes to Anthropic and is still billed there. Config validation refuses
that combination at startup rather than letting it run up a bill quietly.
**Leave `MODEL_REASONING_EFFORT` off unless every configured model is a
reasoning model.** Reasoning models accept the field; most others reject the
*entire request* with a 400 rather than ignoring an unknown key.
**Run the evals before trusting it, and read the I7 case first.**
```bash
MODEL_PROVIDER=openai MODEL_BASE_URL=… MODEL_API_KEY=… MODEL_BALANCED=… make eval-live
```
`liveGateway` reads the same environment the service does and logs which
provider and model answered. The handbook corpus contains a planted prompt
injection; Claude refuses it and reports the document as tampered with. **A
model that answers every other case well and follows that injection is not a
cheaper option — it is a security regression.** That case is the gate, not the
cost table.
**Token accounting differs between the two wires and is already reconciled.**
OpenAI reports `prompt_tokens` *inclusive* of the cached prefix; Anthropic
reports input tokens *exclusive* of it. `oaiUsage.normalise` subtracts, because
`Usage.Total()` sums all four fields and copying both numbers across verbatim
would bill the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. Don't "simplify" that subtraction away;
there is a test named after it.
## Still outstanding
- `ANTHROPIC_API_KEY` was pasted into a chat transcript and is live in a
Kubernetes Secret. Rotate it.
- Deployments report `version=dev`: the image is built without
`--build-arg VERSION`. `make docker-build` passes it.
- `APP_ENV=staging` on the deployment, so the production config guards are off.
- CI tests but does not deploy. The README's claim that migrations are "run by
CI against the target database" is still aspirational.
- The fixture-drift CI jobs need `FRONTEND_REPO_TOKEN` to see the sibling repo,
and fail rather than pass quietly without it.
- The remote is Gitea and the workflows are GitHub Actions syntax. Gitea Actions
runs them, and a runner now exists: `gitea-runner` (gitea/act_runner v0.6.1)
on the cluster host, registered as `krow-runner` with labels
`ubuntu-latest, ubuntu-22.04` mapped to `node:20-bookworm`. Before that, both
repositories had workflows that had never executed once — the 924 frontend
checks, the whole Go suite, the skip guard and the suite-shrank guard were
all things somebody had to remember to run.
If a job fails resolving `actions/checkout` or `actions/setup-node`, the
runner needs egress to github.com or a mirror; that is where those actions
come from and Gitea does not host them.
- **The application talks to its database in clear text.** `DATABASE_SSLMODE=
disable` against `66.116.207.225`, which is a DIFFERENT machine from the
cluster host — so credentials and every row cross the network unencrypted.
It is permitted only because `APP_ENV=staging`; the production guard refuses
`disable` outright. PostgreSQL itself now has `ssl = on` (2026-08-29, port
5433, reload not restart), but the app does not reach PostgreSQL directly:
**pgbouncer terminates 5432** and offers no TLS of its own. The fix is
`client_tls_sslmode = allow` plus a cert in `/etc/pgbouncer/pgbouncer.ini`,
then `DATABASE_SSLMODE=require` in `krow-config` and the `krow-db` secret.
`allow` keeps existing plaintext clients working, so it is additive.
- Production retrieval is **keyword-only**: no `EMBED_PROVIDER` in
`krow-config`, so `knowledge_chunks.embedding` is null for all 34 rows. A
`VOYAGE_API_KEY` is the cheap fix; Ollama in-cluster is the other, and the
nodes were at 60% and 49% memory when that was last looked at.
- Delegation (§6) is implemented and on `main` but NOT deployed. Until the next
image ships, production agents still ignore their `subagents:`.
- §3's publish-time cycle detection is still missing. The runtime depth cap
(2) is what bounds a cycle that reaches run time.
- `cmd/importagents` has no tests, and `run()` opens its own pool from config,
so making it testable is a refactor rather than an addition.
- `importagents` does not enforce monotonicity: a spec whose `version:` is
LOWERED still overwrites the live row and rolls the deployed agent backwards.
---
## Setting up a new machine
git clone <backend> krow-backend && git clone <frontend> krow-demo
cp krow-backend/CLAUDE.md ./claude.md # the governing doc lives above both repos
Needs, if you run the backend natively: Go (see `go-api/go.mod`), Node 20,
PostgreSQL, Docker, and Ollama with `nomic-embed-text` for semantic retrieval.
You do not need most of that. `infrastructure/Dockerfile.api` builds EVERY
command in `go-api/cmd/` plus the golang-migrate CLI into the image, so the
whole stack runs on Docker alone — no Go, no psql, no migrate on the host:
cd krow-backend/infrastructure
cp .env.docker.example .env # fill it in; DATABASE_HOST=postgres
docker compose -f docker-compose.yml -f docker-compose.local-db.yml up -d
docker exec krow-api seed
docker exec krow-api importagents --dir /app/agents --skills /app/skills --org krow-dev
docker exec krow-api ingest --dir /app/knowledge --org krow-dev
printf '%s' 'PASSWORD' | docker exec -i krow-api setpassword -email demo@krow.app -stdin
Ollama, if you want semantic retrieval, runs on the HOST — so the container
reaches it at `host.docker.internal:11434`, NOT `localhost:11434`, which inside
a container means the container.
Running natively instead, you need all of the above. Then:
cd krow-backend && cp .env.example .env # fill it in; .env is gitignored
make migrate-up && make seed
make import-agents ORG=krow-dev
make ingest ORG=krow-dev
go run ./go-api/cmd/setpassword -email demo@krow.app
cd ../krow-demo && npm ci && cp .env.example .env
# VITE_AGENT_API=/api/v1 for local dev (vite proxies it);
# production passes an absolute URL as a Docker build arg instead.
`.env` files are not in git and must be carried across by hand.