§3 says a published version is immutable and editing publishes a new one.
The machinery for that was all present — an append-only definition_versions
table, a trigger, and repo.VersionsRepo.Snapshot, which already refuses to
store a version number whose content differs from what is stored.
Nothing acted on that refusal. snapshotIfPublished's error was discarded at
both call sites (`_ = s.snapshotIfPublished(...)`), and deliberately so: the
comment there explains that losing an author's work to protect a record of it
is the wrong trade. That is right for a recording failure and wrong for
exactly one case. A conflict is not the history failing to record; it is the
invariant firing.
The effect was silent. Editing a published agent without raising the
frontmatter version answered 200: the live row took the new text, the history
kept the old, and two different definitions were both called v1. Because
runtime.LoadAgentVersion resolves a pin by returning the CURRENT definition
whenever the pinned number equals the current one, a conversation pinned to v1
then ran the rewritten instructions while the audit trail showed the
originals. Verified against a live stack before the fix: PATCH answered 200,
agent_definitions held "SILENTLY CHANGED" and definition_versions still held
the published text, both labelled v2.
So the conflict is now detected before anything is written, where refusing
costs the author nothing but a version bump. The post-write snapshot keeps its
original contract for every other kind of failure, and republishing a version
unchanged stays the no-op it was. Drafts are untouched: they carry no promise,
and are still rewritten in place.
Not addressed here, and each its own change:
- cmd/importagents never creates versions at all (documented at main.go:10),
so the nine file-published organization agents are outside this entirely
and every deploy still mutates v1 in place.
- skill definitions never snapshot, so KindSkill exists with nothing writing
it. Fixing that changes skill authoring behaviour and wants its own pass.
- a concurrent publish of one version number with differing content can still
pass this check and be caught by the unique index afterwards, where it is
swallowed as before. That is the pre-existing behaviour, narrowed rather
than removed.
Tests: the new case fails without the fix — the live row takes the rewritten
text at version 1 — and passes with it. Full suite green against PostgreSQL,
with only TestLive* skipped, which is what CI allows.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
The overlay had never been run against a fresh volume. Two faults, the first
hiding the second:
- postgres:16-alpine ships libssl but not the openssl CLI, so the first-boot
certificate generation exited 127 in a restart loop. It failed invisibly:
the 2>/dev/null on the openssl line swallowed sh's "not found" as well, so
`docker logs` was completely empty. openssl is now installed on the boot
that generates the certificate, inside the same guard, so a restart still
needs no network.
- the certificate was written into /var/lib/postgresql/data BEFORE initdb
ran, and initdb refuses to initialise a directory that is not empty. That
made a fresh volume unstartable regardless of the first fault. The
certificate now lives in its own volume, which keeps it persistent — the
reason it was put in the data directory — without touching the cluster's.
Separately, docker-compose.yml did not pass ANTHROPIC_API_KEY to the api
container, so a compose deployment could never register the agent run routes:
POST /agents/{id}/runs answered 404 and /version reported two endpoints fewer.
The model and embedder variables are now passed through, all defaulting to
empty so a deployment without them behaves exactly as it did.
Verified on a fresh volume: 56/56 verify-deploy checks against the resulting
stack, including a live agent run and 34 chunks embedded through Ollama.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
Neither survived a machine change. CLAUDE.md sat in the directory ABOVE both
repositories, which is not a git repository at all, so the governing document
for the project existed on exactly one laptop. It is now in this repository;
place a copy at the parent level on a new machine, where it covers both.
docs/handover.md records what CLAUDE.md does not: what was decided and why,
what is deployed and how to verify it, and the conventions that produce
confident wrong numbers rather than errors — a score of 0 meaning "not rated",
screened_at being vestigial, shift data anchored to today.
Written because Claude Code's own memory is per-machine and keyed to the
absolute path of the checkout: it does not sync, and a different path on a new
machine reads a different folder. A file in the repository travels with the
code and is useful to a person besides.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
§9 says no agent ships without evals. Eight of the nine had none: the two
other suites in evals/ are harness fixtures rather than agents in the
registry, so the rule was being met by one agent in nine.
Evals — 40 new cases, five per agent, every one carrying mustNotLeak:
- the agent is loaded from its real spec in agents/*.md rather than
written out again in Go. A hand-copied agent tests the copy: it keeps
passing after somebody edits the spec, which is the moment it most
needed to fail.
- callNamed calls the tool a case names. toolThenAnswer always called
tools[0], so seven of positions-agent's eight tools were unreachable,
and a boundary nothing calls is a boundary nothing tests.
- seedWorkspace fills BOTH tenants. A leak test against an empty second
tenant cannot fail.
Verified by breaking workersByScore's org predicate: six cases across four
agents fail with LEAKED "RIVAL".
Knowledge — six policy documents, taking the corpus from 2 to 8 (34
chunks). Three restricted to admin and employer, five tenant-wide. They
cover what the tools cannot: a tool reports how many shifts went unworked,
a policy says what cover costs inside 24 hours.
corpus_test.go treats those documents as product rather than fixtures. The
first version was tautological — it read audience: from a file and checked
that file's audience was enforced, so opening a restricted document passed.
mustNotBeTenantWide now holds that judgement apart from the files, with the
reason recorded for each.
CI — the checks this repository already had, made unskippable. testutil
calls t.Skipf on an unreachable database, so a dead service container would
produce a green build over a suite that ran almost nothing. Simulated: go
test exits 0 with 74 tests skipped, including every tenant-isolation test.
The guard exits 1 and names them, while still allowing TestLive* to skip
without a model key.
This CI tests; it does not deploy. The README's claim that migrations are
run by CI against the target database remains aspirational.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
CORS
cors.go never set Access-Control-Allow-Credentials, so the
cookie-authenticated API was unreadable from any cross-origin frontend:
the server answered correctly and the browser blocked the page from
reading it. Set for allowlisted origins on both the preflight and the
actual response. Three tests added.
HTTP_COOKIE_SAMESITE (lax|none|strict, default lax) is new. CORS is only
half of what a cross-origin browser call needs; SameSite is judged on
registrable domain, so a frontend on an unrelated domain gets perfect CORS
headers and still no cookie. "none" is the only value that survives that,
and validate() refuses it without the Secure flag.
The "*" rejection now explains itself: browsers refuse Allow-Origin "*"
together with credentials, so it would break every authenticated call
rather than loosen anything.
Transactional endpoints (api-contract.md 12.1)
POST /api/v1/job-applications/{id}/hire
POST /api/v1/job-postings/{id}/assignments
Replaces two client-side loops that wrote several records with no
transaction and no rollback. Each is now one endpoint and one transaction,
built over repo.Repo so org scoping, derived columns, type casts and error
translation are not re-derived. Authorization reuses the existing policy
table rather than adding a parallel one: a workflow is exactly as
privileged as the writes it performs. 13 tests, including both rollback
paths.
Bug fix in the repository layer
repo.bindValue handled int64/int/float64/string but not int32, which is
what pgx returns for a PostgreSQL `int` column. Nothing previously read a
record and wrote one of its fields elsewhere, so it never surfaced; the
hire flow does exactly that and failed with "ai_score must be a number".
Both KindInt and KindFloat now accept the widths pgx actually produces.
Deployment
infrastructure/Dockerfile.api multi-stage, cross-compiling (BUILDPLATFORM
+ GOARCH) so linux/amd64 builds from arm64 are compiled rather than
emulated. Alpine runtime, non-root uid 10001, 22.1 MB. Ships api, seed,
setpassword and migrate, plus the migrations, so a Kubernetes
initContainer can apply the schema from the same image and tag as the
API. HEALTHCHECK keys on status code, not body, so a "degraded" instance
is not pulled from rotation during a migration window.
infrastructure/docker-compose.yml migrations run to completion before the
API starts. Assumes a managed PostgreSQL; the local-db overlay adds one
with TLS enabled so APP_ENV=production is met rather than dodged.
scripts/drop_public_tables.go the one-off used to clear an unrelated
schema from krowdb on 2026-08-24, kept for the record. Build-tagged
ignore and gated on CONFIRM_DROP=yes.
Verified against PostgreSQL: 16/16 new tests pass, and the image was built,
run and exercised end to end (login, CORS preflight, authenticated reads,
transaction rollback).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CmQiGq73Uyfq7J4yR8Vxxw