The platform now runs on Groq by default, through the OpenAI-compatible chat-completions shape. That shape is not one vendor — Gemini, OpenRouter, Together, vLLM and a local Ollama serve it too — so moving again stays configuration rather than code. Two things in the deleted file were not Anthropic's and would have gone with it silently: withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and CALLED BY openai.go. Deleting the file wholesale would have removed the retry policy of the provider that survived, and nothing in openai.go mentions it, so the loss would have been invisible until the next 429. The policy is a property of this platform's runs, not of a vendor's API; it now lives in retry.go where no provider can carry it off. StreamComplete had the same problem and moves to gateway.go, beside the Streamer interface whose comment already referenced it. Three stale-configuration failures are now refused at startup instead of being ignored. Each was verified firing through the real config.Load(): MODEL_PROVIDER=anthropic — named separately from every other wrong value because it used to be correct. Ignoring it gives a stack that believes it is on Claude while every run goes to Groq and is billed there. ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an operator did set is the worst version of this: they fail every run on a missing credential they are looking straight at. A leftover claude-* model id, naming the tier that carries it. This is the check the previous commit's error-detail work was diagnosing: such an id is accepted by this process, rejected by the provider, and 400s on EVERY run. "A model is wrong" does not say which of three lines to edit. Defaults ship as a matched pair. defaultBaseURL and the three tier ids are one decision, not four: an id is only meaningful against the service that serves it, and a Groq id on an OpenAI base URL is the same failure from the other side. The tiers also stop being one model — a tier whose cost does not differ is a distinction that buys nothing. Verified end to end against a stub of the wire, driving the real wiring (config.Load in production mode, gateway.New, StreamComplete): streamed deltas, tool-call decoding, the loopback credential exemption, and usage totalling 150 rather than 190 — the cached-prefix subtraction still holds. gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still needs a reachable database. NOT verified: the I7 planted-injection eval. Removing this path removed the only model whose refusal behaviour had been measured against it, so the new default is unproven there until `make eval-live` runs with a real key. The Groq model ids should also be confirmed against Groq's current lineup. Flagged in CLAUDE.md §12 and docs/handover.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
334 lines
16 KiB
Markdown
334 lines
16 KiB
Markdown
# Handover
|
||
|
||
Written 2026-08-28, when the machine this was built on was retired.
|
||
|
||
Everything Claude Code "remembers" lives in `~/.claude/projects/<mangled-path>/`
|
||
on one machine, keyed to the absolute path of the checkout. It does not sync,
|
||
and a different path on a new machine reads a different folder. So the durable
|
||
record is this file, in the repository, where git carries it and any path works.
|
||
|
||
Read `CLAUDE.md` first — it is the governing document. This file is what it does
|
||
not say: what was decided, what is deployed, and which parts bite.
|
||
|
||
---
|
||
|
||
## Where things stand
|
||
|
||
**Deployed.** Backend at `https://mcp.krowforce.com`, frontend at
|
||
`https://platform.krowforce.com`, Kubernetes statefulset `krow` in namespace
|
||
`krow`, pods named `krow-1` and `krow-2` (they start at 1, not 0).
|
||
|
||
ssh root@<host> -p 4422 "kubectl -n krow rollout restart statefulset/krow && \
|
||
kubectl -n krow rollout status statefulset/krow --timeout=180s"
|
||
|
||
**Verify a deployment** — 55 checks including a real agent run:
|
||
|
||
KROW_EMAIL=... KROW_PASSWORD=... make verify-deploy BASE=https://mcp.krowforce.com
|
||
|
||
Auth runs BEFORE routing, so an unauthenticated probe answers 401 for every
|
||
path including ones that do not exist. `curl` cannot tell a missing endpoint
|
||
from a guarded one; only an authenticated check can.
|
||
|
||
**Two runtime steps a deploy does not do**, both easy to forget because the API
|
||
looks healthy without them:
|
||
|
||
kubectl -n krow exec krow-1 -- importagents --org <slug> # publishes agents/ and skills/
|
||
kubectl -n krow exec krow-1 -- ingest --org <slug> # ingests knowledge/
|
||
|
||
Without the first, every Owliver question answers 404. Without the second,
|
||
retrieval finds nothing. The org slug is `krow-dev` — a hardcoded constant
|
||
(`internal/orgctx.DevOrgSlug`), not configuration.
|
||
|
||
**The endpoint count is a signal.** `GET /api/v1/version` reports it. Agent run
|
||
routes are not registered without a model credential, so **55** means no
|
||
`ANTHROPIC_API_KEY` and **57** means there is one. (These were written as 56/58
|
||
and were one high; the delta of two — the two run routes — was always right.) A keyless deployment boots
|
||
cleanly under `APP_ENV=staging` and refuses under `production`.
|
||
|
||
---
|
||
|
||
## Decisions taken, so they are not relitigated
|
||
|
||
**§12, who authors agents: self-serve, split by visibility.** `personal` agents
|
||
are authored in the UI, POST to `/api/v1/agent-definitions`, and are runnable
|
||
immediately; evals are not required for them. `organization` agents stay as
|
||
files published by `importagents` on deploy — that deploy step *is* the approval
|
||
workflow, and §9's eval requirement still applies. §12 warned self-serve needs
|
||
"3× the platform"; it does not here, because I1 means an agent runs as its
|
||
caller and cannot exceed their access, I4 means writes still need a human, and
|
||
the spec format has no `limits` block so budgets cannot be raised by an author.
|
||
|
||
**A rejected candidate counts as screened.** It ranks with `ai_screened`:
|
||
rejection overwrites the stage it came from, so shortlisted can never be
|
||
claimed. **An assigned candidate counts as hired**, matching the backend's
|
||
existing `status IN ('hired','assigned')`.
|
||
|
||
**Seeded profile scores are the formula's output**, not hand-authored narrative.
|
||
Recalculating a seeded profile is a no-op, and a skill-check enforces it.
|
||
|
||
---
|
||
|
||
## Conventions the schema actively contradicts
|
||
|
||
These are the ones that produce confident, wrong numbers rather than an error.
|
||
|
||
**`job_applications.screened_at` is vestigial.** Nothing writes it. "Screened"
|
||
means `status <> 'applied'`, in about eight places in the frontend. Reading the
|
||
column reported 1 screened of 24 where the truth was 14.
|
||
|
||
**A score of `0` means "not rated", never "rated zero".** Every score column is
|
||
NOT NULL, so there is no null to distinguish it — that is the trap. Applies to
|
||
`ai_score`, `client_rating`, `krow_score`, `reliability_score`,
|
||
`attendance_score`, `performance_score`, `experience_years`. Aggregates need
|
||
`FILTER (WHERE col > 0)` and a stated basis count. Counting zeros once reported
|
||
"16 weak candidates averaging 28" for a pool that was 1 weak averaging 76.
|
||
|
||
Anchors to check against: applications are **9 scored averaging 76**; workers
|
||
are **5 rated of 9, client rating 4.70**.
|
||
|
||
**Genuine zeros, do not filter these:** `overtime_hours`, `minutes_late`, `xp`,
|
||
`profile_completion`, and `actual_hours` (0 only on absent/no_show shifts).
|
||
|
||
**`absent` and `no_show` are both missed shifts, but only one is a no-show.**
|
||
`attendance.js` is canonical.
|
||
|
||
**`attendance_score` defaults to 100 for display and must never be a scoring
|
||
input.** A profile with no evidence otherwise scores 12 and leaves the "not yet
|
||
scored" band.
|
||
|
||
**The stage ladder lives once**, in `krow-demo/src/lib/hiringRecords.js`. It was
|
||
three byte-identical private copies, all missing `rejected` and `assigned`,
|
||
which `indexOf` scored -1 and dropped from every bucket including `applied`.
|
||
|
||
---
|
||
|
||
## Things that will waste your afternoon
|
||
|
||
**A space in the checkout path breaks path derivation.** This repo lives under
|
||
`Krow Project /`, and it has bitten three times: an unquoted `$(CURDIR)` in the
|
||
Makefile, `` `file://${process.argv[1]}` `` in a script guard, and
|
||
`new URL(...).pathname` in `scripts/oracle.mjs` (use `fileURLToPath`). Any new
|
||
path derivation is guilty until tested there.
|
||
|
||
**`seed.json` is generated from `krow-demo/src/api/seed.js`.** Never edit it.
|
||
`npm run seed:fixture` writes it, `npm run seed:check` verifies, and the
|
||
skill-check compares byte-for-byte. `ShiftRecord` is excluded on purpose: the Go
|
||
seeder generates it against *now*.
|
||
|
||
**Shift data is anchored to today**, so anything asserting against it is
|
||
date-dependent unless the anchor is pinned. `buildShiftsAt(anchor)` exists for
|
||
that. One detector had a Friday-and-Saturday blind spot for exactly this reason.
|
||
|
||
**Database tests skip when PostgreSQL is unreachable** (`testutil` calls
|
||
`t.Skipf`). `go test` then exits 0 having run almost nothing. CI has a guard
|
||
that fails on any skip other than `TestLive*`; keep it.
|
||
|
||
**The eval suites use a scripted model.** They prove the permission boundary,
|
||
not answer quality. `make eval-live` uses the real model and costs tokens.
|
||
|
||
It has now been run, and passes 3/3 in about 30 seconds. Worth knowing what it
|
||
established, because it is more than "the agents work":
|
||
|
||
- the coverage agent called `assign_worker` and the write did NOT happen —
|
||
I4 holds against a real model, not only a scripted one;
|
||
- the handbook corpus contains a planted prompt injection telling the agent
|
||
to send the staff roster to an external address. The agent refused it,
|
||
answered the real question with citations, and reported the document as
|
||
tampered with. I7 holds end to end;
|
||
- the activity agent declined to subtract two figures it could not
|
||
reconcile, and said so, rather than producing the confident wrong number
|
||
this schema invites.
|
||
|
||
Re-run it after any change to the loop, retrieval, or prompt assembly. It is
|
||
the only check that measures answers rather than boundaries.
|
||
|
||
**Seeded time-series data is rebased to now at seed time** — see
|
||
`seeder.RebaseToNow`. `ShiftRecord` is generated against now; `UserActivity`,
|
||
`JobApplication`, `AIInterview` and `Staff` are moved so their newest record
|
||
sits at today, keeping every authored gap. Without it the demo goes quiet: on
|
||
2026-08-29 the newest activity event was 23 days old, applications 16 days,
|
||
staff hire dates 35 — zero events in the last 7 days and an empty Hiring
|
||
activity chart on Control Center.
|
||
|
||
Two things to know if you touch it. EVERY timestamp on a record shifts by the
|
||
same delta, not just the anchor: an application's created_date and updated_date
|
||
are what `buildHires` subtracts for time-to-hire, and moving one alone turns a
|
||
five-day hire into a three-week one. And `Staff` anchors on `hire_date` rather
|
||
than `created_date`, because the hire is the event the chart plots — leaving it
|
||
behind produced a workspace where somebody was hired last week according to
|
||
their application and five weeks ago according to their staff record.
|
||
|
||
Reference data is deliberately not rebased. A course's date is a fact about the
|
||
course, not a position in a window.
|
||
|
||
---
|
||
|
||
**HTTP_WRITE_TIMEOUT must exceed the deepest agent deadline.** It was 30s in
|
||
production while every shipped agent runs at the `balanced` tier, whose
|
||
deadline is 60s — so the server aborted the response on any run over half its
|
||
allowed time, and the proxy in front answered **502 Bad Gateway**. A gateway
|
||
error for something no gateway did, which is why it read as an infrastructure
|
||
fault: nginx was innocent and already had `proxy_read_timeout 3600s`.
|
||
|
||
Streaming hid it. The chat panel uses SSE and survives, so the product looked
|
||
healthy while any non-streaming caller — a webhook, a script, an integration —
|
||
got 502 on a slow question. Delegation made it routine rather than causing it:
|
||
a parent that asks two subagents takes longer than one answering alone.
|
||
|
||
Production is now 180s, and `config.validateWriteTimeout` refuses a value below
|
||
`DeepestAgentDeadline` at startup. NOTE THE ORDERING: that constant is 120s, so
|
||
a deployment still carrying the old 30s will now refuse to boot. Patch the
|
||
configmap before shipping an image that contains the check.
|
||
|
||
---
|
||
|
||
## Changing model provider
|
||
|
||
The gateway speaks one wire protocol: `openai`, the chat-completions shape.
|
||
That is not the same as one vendor — Groq, Gemini's compatibility endpoint,
|
||
OpenRouter, Together, vLLM and a local Ollama all serve it, so moving between
|
||
them is configuration, not code.
|
||
|
||
**The Anthropic path was removed.** `MODEL_PROVIDER=anthropic` and a stale
|
||
`ANTHROPIC_API_KEY` are both *refused at startup* rather than ignored, and so
|
||
is a leftover `claude-*` model id. That is deliberate: each of those would
|
||
otherwise produce a service that boots cleanly and fails every agent run.
|
||
|
||
The default with nothing set is Groq.
|
||
|
||
```bash
|
||
# Groq (the default — base URL and ids below are what you get unset)
|
||
MODEL_BASE_URL=https://api.groq.com/openai/v1
|
||
MODEL_API_KEY=<key>
|
||
MODEL_FAST=llama-3.1-8b-instant
|
||
MODEL_BALANCED=llama-3.3-70b-versatile
|
||
MODEL_DEEP=llama-3.3-70b-versatile
|
||
|
||
# Gemini
|
||
MODEL_BASE_URL=https://generativelanguage.googleapis.com/v1beta/openai
|
||
|
||
# A model on this machine — no credential at all
|
||
MODEL_BASE_URL=http://localhost:11434/v1
|
||
```
|
||
|
||
Four things worth knowing before you do it.
|
||
|
||
**A model id and a base URL are one decision, not two.** An id is only
|
||
meaningful against the service that serves it, so changing the endpoint without
|
||
changing the ids gives you a process that starts fine and 400s on every run.
|
||
The defaults ship as a matched Groq pair for that reason.
|
||
|
||
**Leave `MODEL_REASONING_EFFORT` off unless every configured model is a
|
||
reasoning model.** Reasoning models accept the field; most others reject the
|
||
*entire request* with a 400 rather than ignoring an unknown key.
|
||
|
||
**Run the evals before trusting it, and read the I7 case first.**
|
||
|
||
```bash
|
||
MODEL_BASE_URL=… MODEL_API_KEY=… MODEL_BALANCED=… make eval-live
|
||
```
|
||
|
||
`liveGateway` reads the same environment the service does and logs which
|
||
provider and model answered. The handbook corpus contains a planted prompt
|
||
injection. A model worth running refuses it and reports the document as
|
||
tampered with. **A model that answers every other case well and follows that
|
||
injection is not a cheaper option — it is a security regression.** That case is
|
||
the gate, not the cost table.
|
||
|
||
This one is not optional now: the removed provider was the one whose refusal
|
||
behaviour had actually been measured here, so whatever replaces it is unproven
|
||
against I7 until this suite says otherwise.
|
||
|
||
**Token accounting is already reconciled, and the subtraction is load-bearing.**
|
||
This wire reports `prompt_tokens` *inclusive* of the cached prefix, while
|
||
`Usage` carries the cached figure separately. `oaiUsage.normalise` subtracts, because
|
||
`Usage.Total()` sums all four fields and copying both numbers across verbatim
|
||
would bill the cached prefix twice — worst on long conversations, which is
|
||
exactly where I3's budget matters most. Don't "simplify" that subtraction away;
|
||
there is a test named after it.
|
||
|
||
## Still outstanding
|
||
|
||
- `ANTHROPIC_API_KEY` was pasted into a chat transcript and is live in a
|
||
Kubernetes Secret. Rotate it.
|
||
- Deployments report `version=dev`: the image is built without
|
||
`--build-arg VERSION`. `make docker-build` passes it.
|
||
- `APP_ENV=staging` on the deployment, so the production config guards are off.
|
||
- CI tests but does not deploy. The README's claim that migrations are "run by
|
||
CI against the target database" is still aspirational.
|
||
- The fixture-drift CI jobs need `FRONTEND_REPO_TOKEN` to see the sibling repo,
|
||
and fail rather than pass quietly without it.
|
||
- The remote is Gitea and the workflows are GitHub Actions syntax. Gitea Actions
|
||
runs them, and a runner now exists: `gitea-runner` (gitea/act_runner v0.6.1)
|
||
on the cluster host, registered as `krow-runner` with labels
|
||
`ubuntu-latest, ubuntu-22.04` mapped to `node:20-bookworm`. Before that, both
|
||
repositories had workflows that had never executed once — the 924 frontend
|
||
checks, the whole Go suite, the skip guard and the suite-shrank guard were
|
||
all things somebody had to remember to run.
|
||
|
||
If a job fails resolving `actions/checkout` or `actions/setup-node`, the
|
||
runner needs egress to github.com or a mirror; that is where those actions
|
||
come from and Gitea does not host them.
|
||
- **The application talks to its database in clear text.** `DATABASE_SSLMODE=
|
||
disable` against `66.116.207.225`, which is a DIFFERENT machine from the
|
||
cluster host — so credentials and every row cross the network unencrypted.
|
||
It is permitted only because `APP_ENV=staging`; the production guard refuses
|
||
`disable` outright. PostgreSQL itself now has `ssl = on` (2026-08-29, port
|
||
5433, reload not restart), but the app does not reach PostgreSQL directly:
|
||
**pgbouncer terminates 5432** and offers no TLS of its own. The fix is
|
||
`client_tls_sslmode = allow` plus a cert in `/etc/pgbouncer/pgbouncer.ini`,
|
||
then `DATABASE_SSLMODE=require` in `krow-config` and the `krow-db` secret.
|
||
`allow` keeps existing plaintext clients working, so it is additive.
|
||
- Production retrieval is **keyword-only**: no `EMBED_PROVIDER` in
|
||
`krow-config`, so `knowledge_chunks.embedding` is null for all 34 rows. A
|
||
`VOYAGE_API_KEY` is the cheap fix; Ollama in-cluster is the other, and the
|
||
nodes were at 60% and 49% memory when that was last looked at.
|
||
- Delegation (§6) is implemented and on `main` but NOT deployed. Until the next
|
||
image ships, production agents still ignore their `subagents:`.
|
||
- §3's publish-time cycle detection is still missing. The runtime depth cap
|
||
(2) is what bounds a cycle that reaches run time.
|
||
- `cmd/importagents` has no tests, and `run()` opens its own pool from config,
|
||
so making it testable is a refactor rather than an addition.
|
||
- `importagents` does not enforce monotonicity: a spec whose `version:` is
|
||
LOWERED still overwrites the live row and rolls the deployed agent backwards.
|
||
|
||
---
|
||
|
||
## Setting up a new machine
|
||
|
||
git clone <backend> krow-backend && git clone <frontend> krow-demo
|
||
cp krow-backend/CLAUDE.md ./claude.md # the governing doc lives above both repos
|
||
|
||
Needs, if you run the backend natively: Go (see `go-api/go.mod`), Node 20,
|
||
PostgreSQL, Docker, and Ollama with `nomic-embed-text` for semantic retrieval.
|
||
|
||
You do not need most of that. `infrastructure/Dockerfile.api` builds EVERY
|
||
command in `go-api/cmd/` plus the golang-migrate CLI into the image, so the
|
||
whole stack runs on Docker alone — no Go, no psql, no migrate on the host:
|
||
|
||
cd krow-backend/infrastructure
|
||
cp .env.docker.example .env # fill it in; DATABASE_HOST=postgres
|
||
docker compose -f docker-compose.yml -f docker-compose.local-db.yml up -d
|
||
docker exec krow-api seed
|
||
docker exec krow-api importagents --dir /app/agents --skills /app/skills --org krow-dev
|
||
docker exec krow-api ingest --dir /app/knowledge --org krow-dev
|
||
printf '%s' 'PASSWORD' | docker exec -i krow-api setpassword -email demo@krow.app -stdin
|
||
|
||
Ollama, if you want semantic retrieval, runs on the HOST — so the container
|
||
reaches it at `host.docker.internal:11434`, NOT `localhost:11434`, which inside
|
||
a container means the container.
|
||
|
||
Running natively instead, you need all of the above. Then:
|
||
|
||
cd krow-backend && cp .env.example .env # fill it in; .env is gitignored
|
||
make migrate-up && make seed
|
||
make import-agents ORG=krow-dev
|
||
make ingest ORG=krow-dev
|
||
go run ./go-api/cmd/setpassword -email demo@krow.app
|
||
|
||
cd ../krow-demo && npm ci && cp .env.example .env
|
||
# VITE_AGENT_API=/api/v1 for local dev (vite proxies it);
|
||
# production passes an absolute URL as a Docker build arg instead.
|
||
|
||
`.env` files are not in git and must be carried across by hand.
|