The krow-2 deploy failed on the ANTHROPIC_API_KEY guard, which is the guard doing its job. Checking what an operator hits *after* fixing it turned up two older faults in the files they are told to copy — both predating the Groq switch, both fatal at boot. HTTP_WRITE_TIMEOUT shipped as 30s in .env.example, .env.docker.example and the compose default, while validateWriteTimeout refuses anything at or under the deep tier's 2m deadline. `cp .env.docker.example .env && docker compose up` could not start. Now 180s. krow-2 never saw this because someone had already overridden it in that environment. .env.docker.example carried no model block at all, so a production stack built from it is refused for a missing MODEL_API_KEY. Added, with the Groq defaults and the reasoning-effort note (most non-reasoning models reject the request rather than ignoring the key). Neither was subtle. Both survived because the examples were prose to every test in this package: the validator and the file documenting it had no mechanical connection, so tightening one silently invalidated the other. That connection is now TestShippedExampleEnvActuallyBoots, which parses each example and runs Load() on it under the APP_ENV the file itself declares — production for the docker one, development for the root one, each internally consistent. Verified by mutation: reverting the timeout, removing the key line, and restoring a claude-* id each fail it with the message an operator would see. Go does not treat these files as test inputs, so an example-only edit can be served a stale pass from the test cache. Noted in the test; use -count=1. Also documented the upgrade path in handover.md, including the one thing startup validation cannot catch: renaming ANTHROPIC_API_KEY to MODEL_API_KEY without replacing the value boots fine and 401s on every run. gofmt clean, go vet clean, 15/15 packages pass. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
krow-backend
The backend for Krow — a Go API over PostgreSQL, with a Python Owliver service
to follow. The Krow frontend lives in a separate repository (krow-demo)
and is not vendored, copied or modified here.
Status: authenticated, authorized, multi-tenant. The API serves the
endpoints in docs/api-contract.md against PostgreSQL, loaded with the
frontend's own demo dataset. Every endpoint except GET /health and the two
/api/v1/auth/* routes requires a session: sign in with a password, hold an
HttpOnly cookie, and the server resolves it to a real user on every request.
Roles are enforced — a deny-by-default policy table answers 403 before any query
runs, and organization scope and talent ownership are SQL predicates, so a row
you may not see answers 404. On top of that sits the agent and skill definition
surface: Markdown with YAML frontmatter, parsed and stored, with full CRUD.
Not built: no Owliver, no RAG, no Redis, no object storage. The agent/skill runtime is an internal boundary with no executor behind it and no HTTP route into it.
docs/KROW_BACKEND_COMPLETE_SUMMARY.md is the long-form technical account —
architecture, history, and a frank list of the current limitations.
Layout
krow-backend/
├── go-api/
│ ├── cmd/
│ │ ├── api/ the HTTP service entrypoint
│ │ ├── seed/ loads the demo dataset
│ │ └── setpassword/ the only way a password enters the database
│ ├── internal/
│ │ ├── config/ environment loading + validation
│ │ ├── db/ pgx pool and the database health check
│ │ ├── domain/ resource descriptors (generated) + the policy table
│ │ ├── repo/ pgx access layer — every statement built here
│ │ ├── service/ validation, org scoping, contract semantics
│ │ ├── auth/ argon2id passwords, sessions, the user store
│ │ ├── authctx/ the authenticated identity a request carries
│ │ ├── orgctx/ the organization a request runs as
│ │ ├── definition/ the agent/skill Markdown + YAML frontmatter parser
│ │ ├── runtime/ agent/skill load, validate, resolve, dispatch
│ │ ├── httpserver/ router, handlers, /health, graceful shutdown
│ │ ├── seeder/ fixture loading + the attendanceSeed.js port
│ │ └── testutil/ disposable migrated + seeded test database
│ └── go.mod
├── migrations/ SQL migrations — the source of truth for the schema
├── seed/fixtures/ seed.json, generated from the frontend repository
├── docs/
│ ├── api-contract.md the frozen client contract
│ └── KROW_BACKEND_COMPLETE_SUMMARY.md the long-form technical account
├── infrastructure/ deployment definitions (empty until later)
├── scripts/ verify_schema.sql, gen_resources.py, oracle.mjs
├── Makefile developer, migration and seed tasks
└── .env.example
migrations/, seed/, scripts/ and infrastructure/ sit beside go-api/
rather than inside it because none of them is Go: the migrations are run by the
golang-migrate CLI, the generator is Python, the parser oracle is Node, and a
second service (Owliver, Python) is expected to consume the same schema.
Architecture
HTTP handler → service → repository → pgx → PostgreSQL
No ORM. Every statement is built in internal/repo from a *domain.Resource:
column lists are explicit, every value is a bind parameter cast to its declared
type, and no identifier ever comes from user input — a filter or sort name is
resolved to a real column before any SQL is assembled.
The resource descriptors in internal/domain/resources_gen.go are generated
from the live schema (make gen-resources), so column names, types, enum
values and nullability cannot drift from the migrations. The hand-maintained
part is the per-resource metadata — path, default sort, default limit, supported
operations, required fields — which comes from docs/api-contract.md.
One shared query builder rather than fourteen repositories is deliberate: the
contract's awkward semantics (NULLS LAST in both directions, the , id
tiebreaker, per-endpoint limits, shallow PATCH, idempotent DELETE) are then
implemented once and apply identically everywhere.
Requirements
| Tool | Version used | Install |
|---|---|---|
| Go | 1.27 | brew install go |
| golang-migrate | 4.19.1 | brew install golang-migrate |
| PostgreSQL | 18.6 | brew install postgresql@18 |
Getting started
cp .env.example .env # then fill in DATABASE_USER / DATABASE_PASSWORD
make db-create # creates the database if it does not exist
make migrate-up # applies every migration in migrations/
make seed # loads the frontend's demo dataset
make run # http://127.0.0.1:8080/health
The API
54 registered routes: 34 entity routes across 14 resources, 10 agent and skill
definition routes, 4 /me routes, 2 /auth routes, 2 workflow routes, the
Owliver suggestion route and GET /health. The entity routes and the suggestion
route are specified in docs/api-contract.md (§2 and §2A); the definition and
workflow routes are not yet in the contract and are documented in the source and
its tests. Every
entity route exists because a frontend call site exists — a table in the
database is never a reason for an endpoint. DELETE /job-postings/{id} and
POST /shift-records return 405 because nothing in the frontend deletes a
posting or creates a shift record.
curl localhost:8080/api/v1/job-postings
curl "localhost:8080/api/v1/job-applications?job_posting_id=<id>&limit=50"
curl "localhost:8080/api/v1/job-applications?status=hired&status=interview"
curl localhost:8080/api/v1/me
Authentication is required. POST /api/v1/auth/login verifies a password
with argon2id and issues a server-side session; the raw token goes out in an
HttpOnly, SameSite=Lax cookie and only its SHA-256 is stored. Middleware resolves
that cookie to a user on every request and puts the user's organization on the
context — the same seam devOrgMiddleware used to occupy, so nothing in the
service or repository layers changed.
Set a password before signing in: the seeded user's password_hash is NULL until
go run ./cmd/setpassword -email demo@krow.app is run.
curl -c jar -X POST localhost:8080/api/v1/auth/login \
-H 'Content-Type: application/json' \
-d '{"email":"demo@krow.app","password":"…","remember_me":false}'
curl -b jar localhost:8080/api/v1/me
curl -b jar -X POST localhost:8080/api/v1/auth/logout
Authorization is enforced. users.role — admin, employer or talent,
never the self-editable account_type — is the only authority. Entity routes
are gated by the deny-by-default policy table in internal/domain/policy.go:
a resource with no policy permits nothing to anyone, and Server.authorize
answers 403 before any query runs. Row visibility is a SQL predicate rather than
a filter — the organization scope always, plus an ownership clause for talent
callers — so a row outside it is never fetched and answers 404, which does not
distinguish "exists but not yours" from "does not exist".
The ten definition routes are the exception: they are not domain.Resource
values, so the descriptor machinery does not reach them and their role checks
are written inline in internal/service/definitions.go. Tenancy and ownership
are enforced for them, in the repository predicates.
Seeding
make seed loads seed/fixtures/seed.json, which is generated by executing the
frontend's src/api/seed.js through Vite — so ids, dates, numbers and enum
values arrive exactly as the demo has them, with no transcription step.
Shift records are the exception: they are generated by a Go port of
attendanceSeed.js rather than snapshotted, because their dates are anchored to
now. dataResolver.inPeriod windows every collection on created_date, so a
frozen snapshot would read as permanently empty a fortnight later.
Idempotency: explicit upsert inside one transaction. Every record's key is derived deterministically from its source id (uuid v5), so re-running targets the same rows and restores them to their seeded values. Nothing is deleted, so records created through the API survive a re-seed.
Tests
make test # go test ./...
175 test functions. The database-backed ones build a disposable database —
dropped, recreated, migrated and seeded per run, named
krow_backend_autotest_<pid> so concurrent test packages cannot collide. They
skip rather than fail when PostgreSQL is unreachable.
Seed assertions compare the database against the fixture field-by-field rather
than against numbers typed into a test. The two ordering guarantees
(NULLS LAST in both directions, and the , id tiebreaker) are covered by tests
verified to fail when the guarantee is removed.
Parser conformance. The Go definition parser must agree with the frontend's
JavaScript one. internal/definition/conformance_test.go asserts against
internal/definition/testdata/oracle.json, which is not written by hand: it is
captured by running the real frontend module graph through Vite, so the fixture
records what the JS parser actually does rather than what anyone believes it
does. Regenerating it needs a krow-demo checkout, and is not part of
make test:
node scripts/oracle.mjs go-api/internal/definition/testdata/oracle.json
scripts/cases.mjs holds the adversarial corpus that file is captured over.
GET /health is public and says only whether traffic should be sent here:
{"status": "ok"} with 200, {"status": "degraded"} with 200 when the
database is up but unmigrated or left dirty, and {"status": "unavailable"}
with 503 when it is unreachable. Nothing about the server, the database or
the schema appears in the body — an unauthenticated caller gets the verdict,
not the reasoning.
The check itself is unchanged: db.Check still gathers the PostgreSQL version,
the database and schema names, the applied migration version, the table count
and the connection error, and the handler logs all of it. Read it from the
server log (debug when healthy, warn otherwise), or from psql.
Migrations
Migration files are the source of truth for the schema. Nothing in go-api/
issues DDL, no ORM generates it, and no table is ever created by hand — if the
database and migrations/ disagree, migrations/ is right and the database is
wrong.
make migrate-up # apply everything pending
make migrate-status # current version
make migrate-down # roll back exactly one step
make migrate-new NAME=add_sessions # scaffold the next pair
make verify-schema # print tables, enums, FKs, indexes
Ordering. -seq numbering: 000001, 000002, … golang-migrate applies
them in ascending order and records the highest applied version in
schema_migrations. Two developers who both branch off 000001 and both write
000002 get a collision at merge — rename the later one rather than resolving
it in the database.
Never edit an applied migration. Once a migration has run anywhere other
than your own machine, it is immutable. Change it and every database that
already applied it silently diverges from every one that has not. Write
000002 instead.
Guards against destructive change:
- There is no
make migrate-drop. golang-migrate'sdropcommand deletes every table in the database; it is deliberately not one keystroke fromdown. make migrate-downrefuses to run unlessAPP_ENV=development.- Down migrations name every object they drop. No
DROP SCHEMA, noDROP DATABASE, noCASCADEon a schema. config.validate()rejectsDATABASE_SCHEMA=pg_catalog,pg_toast,information_schemaor anything starting withpg_, so the application can never be pointed at a system schema.DATABASE_SSLMODE=disableis rejected whenAPP_ENV=production.
Dirty state. If a migration fails partway, golang-migrate marks the version
dirty and refuses to continue. Fix the SQL, restore from backup if the failure
left data behind, then make migrate-force VERSION=<last-good> and re-run.
/health answers "degraded" while the version is dirty, and the handler logs
migration_dirty alongside the applied version, so the detail is in the server
log rather than in the public response.
Staging and production. The same files, run by CI against the target database as a discrete deploy step before the new binary rolls out — which means every migration must be backwards-compatible with the currently-running version. Expand, migrate, contract: add a nullable column, backfill, switch reads, drop the old column in a later release. Never in one migration.
Schema
migrations/000001_initial_schema.up.sql creates 17 tables in public. Fifteen
correspond to entities the frontend actually reads or writes through
src/api/base44Client.js; organizations and user_preferences are the two
additions, and both are documented in the migration itself.
Conventions:
uuidprimary keys, with a nullable uniquelegacy_id textso the existing seeded record ids survive a data migration.timestamptzthroughout.created_date/updated_datekeep their frontend names deliberately. The frontend windows every collection oncreated_dateand its default sort string is-created_date; renaming these tocreated_atwould mean touching the resolver, the store's sort parser and every hook.- Native enums for closed vocabularies,
text+CHECKwhere the set is still moving. - Email columns are
citextand stay populated even where a uuid FK also exists — the frontend joins workers by email today, and those joins must keep working.