Files
krow_backend/docs/KROW_BACKEND_COMPLETE_SUMMARY.md
2026-08-25 16:37:05 +05:30

2417 lines
133 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Krow Backend — Complete Technical Summary
**Audience:** backend, frontend, AI/agent and DevOps engineers, and engineering leads.
**Scope:** what exists in the `krow-backend` repository today, how it got here, and where it is heading.
**How this document was produced.** Every factual claim below was read out of the
repository at `/Users/apple/Krow/krow-backend` as it currently stands — Go source,
`go.mod`, migrations, `Makefile`, `README.md`, `docs/api-contract.md`, tests and
fixtures — plus read-only queries against the local development database. Where an
earlier report or an in-repository comment disagrees with the code, **the code is
treated as authoritative** and the disagreement is recorded in §17.
Anything that could not be checked against this repository is marked
**"Not verified from current repository."** Nothing here is inferred from a system
that is not present in this checkout.
This is a technical history and current-state document. It does not certify phases.
---
## 1. Executive Summary
### What the Krow backend is
Krow is a hiring and workforce platform: job postings, applications, AI screening
interviews, staff records, worker profiles, assignments, attendance, a training
library, evidence of work, and an audit log. On top of that sits an authoring
system for **Agents** and **Skills** — Markdown documents with YAML frontmatter that
describe what the product's assistant can be offered on each page.
`krow-backend` is the Go HTTP API and PostgreSQL database behind that product. The
Krow frontend lives in a **separate repository** (`krow-demo`) and is not vendored,
copied or modified here.
### Why the backend was introduced
The frontend shipped first, as a demo. Its data layer was a client-side store:
`src/api/store.js` held arrays in memory, `src/api/seed.js` filled them at boot, and
`src/api/base44Client.js` was the seam every hook called through. That arrangement
has three properties that stop being acceptable the moment more than one person uses
the product:
- **Data does not survive.** It lives in the browser tab.
- **There is no tenancy and no identity.** Every viewer sees the same array.
- **Nothing is enforceable.** Any rule the client does not apply does not exist.
The backend exists to move persistence, identity, tenancy and enforcement to a
server, **without rewriting the frontend above the transport seam**. That constraint
is written into `docs/api-contract.md` as its acceptance criterion:
> Replacing the transport inside `src/api/base44Client.js` — and changing no other
> frontend file — must leave the application behaving identically.
Every design decision downstream follows from that sentence.
### The transformation
```
OLD CURRENT
React React
| |
base44Client base44Client
| |
localStorage / mock data HTTP
|
Go API
|
PostgreSQL
```
The shape above the seam is unchanged. What moved is everything below it: the
records now live in PostgreSQL, the organization a request runs as comes from a
server-side session rather than from nowhere, and the rules about who may read or
write what are applied in SQL and in the handler rather than in the browser.
### The Agent / Skill path
Agents and Skills are authored as Markdown. The backend parses them, validates them
against the same rules the frontend editor applies, projects a few queryable columns
out of them, and stores the Markdown itself verbatim:
```
Markdown
|
Parser (internal/definition — YAML subset + frontmatter, ported from JS)
|
Validator (ValidateAgent / ValidateSkill — frontend messages, plus DB bounds)
|
PostgreSQL (agent_definitions / skill_definitions; markdown stored byte-for-byte)
|
CRUD API (/api/v1/agent-definitions, /api/v1/skill-definitions)
|
Runtime Loader (internal/runtime — loads by id or definition_id, re-parses)
|
Dependency Resolution (agent's `skills:` list resolved within tenant + ownership scope)
|
Execution Boundary (AgentExecutor / SkillExecutor interfaces)
```
### What the runtime can and cannot do today
**It can:** load an authored agent or skill from the database by UUID or by
`definition_id`; re-parse and re-validate its Markdown; refuse it if its status makes
it ineligible (a draft or archived agent, an inactive skill); resolve an agent's
declared skill dependencies within the caller's organization and ownership scope,
preferring a personal definition over an organization one with the same
`definition_id`; deduplicate the resolved list; and hand the resolved agent to an
executor through a typed interface.
**It cannot:** execute anything. The only executor implementation in the repository
is `UnavailableExecutor`, which returns `ErrExecutorUnavailable` for both agents and
skills. There is no LLM client, no prompt assembly, no tool dispatch and no external
AI call anywhere in `go-api/`.
**It is also not reachable over HTTP.** No route in `internal/httpserver` references
the `runtime` package; a repository-wide search for `runtime.` outside the package
itself and its tests returns nothing. The runtime is an internal boundary exercised
only by its own test suite.
### Current architectural direction
The backend is a layered, no-ORM Go service where the shape of every resource is
**generated from the live database schema** and the rules about every resource are
**hand-written next to it**. Authorization is deny-by-default. Tenancy and ownership
are SQL predicates rather than post-fetch filters. The definition system treats the
authored Markdown as the authoritative artefact and every column derived from it as
a cache. The runtime establishes where AI execution will attach, without attaching
it.
---
## 2. Development History
The repository documents its own work in phases, in migration headers, package
documentation and `docs/api-contract.md`. What follows describes what each phase
introduced and what of it is in the code today.
### Phase 1 — Backend Foundation
**What was built.** The Go module, the PostgreSQL connection, the migration
discipline and the initial schema.
- **Go module** — `github.com/krow/krow-backend/go-api`, `go 1.27`. Three direct
dependencies and nothing else: `github.com/jackc/pgx/v5`, `golang.org/x/crypto`,
`golang.org/x/term`. No web framework, no ORM, no YAML library, no test framework.
- **Database access** — `internal/db` owns a `pgxpool.Pool`. `db.Open` deliberately
acquires and pings once at startup, because `pgxpool.New` is lazy and would
otherwise defer a misconfiguration to the first request.
- **Configuration** — `internal/config` loads from the environment with typed
defaults and a `validate()` pass that refuses: an unknown `APP_ENV`; a port out of
range; a `DATABASE_SCHEMA` that is a PostgreSQL system schema or starts with
`pg_`; `MIN_IDLE_CONNS` above `MAX_OPEN_CONNS`; `DATABASE_SSLMODE=disable` when
`APP_ENV=production`; and a CORS entry of `*` or one without a scheme.
- **Health** — `GET /health`, and `db.Check` behind it.
- **Migrations** — `migrations/` is the source of truth for the schema. Nothing in
`go-api/` issues DDL, and no ORM generates any. The `Makefile` wraps the
`golang-migrate` CLI (README records 4.19.1).
**Implementation decisions worth carrying forward.**
- `uuid` primary keys plus a nullable unique `legacy_id text`, so the frontend's
existing string ids (`jobposting_m1a2b3c001`) survive as data without becoming the
addressable id.
- `timestamptz` throughout; `created_date` / `updated_date` keep the frontend's own
names, because the frontend windows every collection on `created_date` and its
default sort string is `-created_date`.
- Native enums for closed vocabularies, `text` + `CHECK` where the set is still
moving.
- Email columns are `citext`, and stay populated even where a UUID foreign key also
exists, because the frontend joins workers by email today.
- Guards against destructive change: there is deliberately **no** `make migrate-drop`
target; `make migrate-down` refuses unless `APP_ENV=development`; down migrations
name every object they drop, with no `DROP SCHEMA` and no `CASCADE` on a schema.
### Phase 2A — Architecture Decisions
Four questions were raised, investigated against the frontend, and recorded in
`docs/api-contract.md` §11 as deferred with reasons. All four are still recorded
there as deferred.
| Ref | Question | Recorded outcome |
| --- | --- | --- |
| **D2** | Is `company` a real entity? | It is free text on `job_postings`, displayed and read through a fallback chain, but never grouped by id, joined, or given a route. Kept as `company text NOT NULL DEFAULT ''`. No `clients` table, no endpoint. |
| **D3** | Are `assigned` / `rejected` pipeline stages? | The frontend's `STAGE_ORDER` excludes both from every funnel count via `indexOf(...) >= from`. The enum carries all seven values; the API stores and returns `status` verbatim and never filters, reinterprets or normalises it. The funnel exclusion is recorded as a live frontend behaviour that Phase 2 does not touch. |
| **D6** | Do the unmounted pages come back? | Fourteen page files plus a layout are imported by `App.jsx` and mounted on no route. Three operations are reachable only through them. All three endpoints were included anyway and marked *unreachable-today*, because the calling code exists and excluding them would make the transport shim need conditional methods. |
| **U1** | Where do shift records come from? | `ShiftRecord` has exactly one frontend seam operation: `.list('-created_date', 500)`. No create, no update, no delete. The contract specifies `GET /api/v1/shift-records` only — no `POST` — because there is no call site to derive a write contract from. |
**Organization / tenant modelling.** `organizations` and `user_preferences` are the
two tables in migration 000001 that do not correspond to an entity the frontend
reads through `base44Client.js`; both are documented in the migration itself. Every
tenant-scoped table carries `org_id`. Two tables — `courses` and `learning_paths` —
are *organization-nullable*: a `NULL org_id` means the shared platform library,
visible to every tenant.
### Phase 2B — API Contract
`docs/api-contract.md` (1,060 lines) was written before implementation, derived
entirely from the frontend repository and the Phase 1 schema. Its stated rule: every
endpoint exists because a call site exists, and every semantic was read out of
`src/api/store.js` rather than designed.
- **Operation matrix.** An operation with no call site gets no endpoint.
`DELETE /job-postings/{id}` and `POST /shift-records` are absent for that reason,
and the router answers 405 for them because the path pattern exists under another
method.
- **REST shape.** Base path `/api/v1`; kebab-case plural resource names, with mass
nouns staying singular (`/staff`, `/evidence`, `/user-activity`); `PATCH` with
shallow merge and no `PUT`.
- **Field naming.** snake_case, identical to the frontend's own field names — no
renaming, no camelCase conversion. The single documented exception is
`/me/preferences`.
- **Envelope.** Responses are wrapped in `{"data": …}`, with `meta` on collections.
The frontend's methods return bare values; the shim unwraps with one `.data`. What
the envelope buys is `meta.total`, which the bare shape has nowhere to put.
- **Filtering (§6).** Two operators only: equality for a single value, membership
for a repeated parameter. Every non-reserved query parameter is a field filter.
- **Sorting (§7).** `NULLS LAST` in **both** directions, because the frontend's
comparator returns before it can place a null; and `, id` appended to every
`ORDER BY`, because JavaScript's sort is not stable in the way the UI relies on.
- **Pagination (§8).** `limit` / `offset`, with per-endpoint defaults taken from the
literal arguments at each frontend call site.
- **Errors (§5).** A single envelope shape with a code, a message and a details map.
### Phase 2C — Backend Implementation
**Layers.** `internal/domain` (descriptors and errors) → `internal/service`
(parsing, validation, contract semantics) → `internal/repo` (all SQL) → pgx.
**Generated descriptors.** `internal/domain/resources_gen.go` is produced by
`scripts/gen_resources.py` (`make gen-resources`), which reads column names, types,
enum values and nullability out of `information_schema`. The hand-maintained part is
the per-resource metadata — path, default sort, default limit, supported operations,
required fields — which comes from the contract. Column definitions therefore cannot
drift from the migrations.
**One query builder, not fourteen repositories.** `internal/repo/repo.go` builds
every statement from a `*domain.Resource`: column lists are explicit, every value is
a bind parameter cast to its declared type, and no identifier ever comes from user
input — a filter or sort name is resolved to a real `*domain.Column` before any SQL
is assembled. The contract's awkward semantics are then implemented once and apply
identically everywhere.
**Seed system.** `seed/fixtures/seed.json` is generated by executing the frontend's
`src/api/seed.js` through Vite and serialising what it exports, so ids, dates,
numbers and enum values arrive exactly as the demo has them with no transcription
step. `internal/seeder` loads it inside a single transaction with explicit upserts.
Shift records are the exception and are generated by a Go port of
`attendanceSeed.js` — see §6.
**Bugs and gaps this phase exposed**, all recorded in `docs/api-contract.md` §13:
- **Three columns had no home** (§13.1). `job_applications.interview_id`,
`courses.training_outline` and `worker_profiles.score_breakdown` are written or
read by the frontend and were absent from migration 000001. Migration 000001 was
not edited; the columns were added in **000002**. `interview_id` is deliberately
*not* a foreign key, because a seeded application references an interview id for
which no record exists, and nulling it would have transformed source data.
- **A constraint that rejected real data** (§13.2). `job_applications_screened_consistent`
from 000001 rejected 9 of 24 seeded applications — precisely the 9 AI-scored ones.
`screened_at` is never read or written by the frontend; the constraint came from a
blueprint rather than from evidence. Dropped in **000003**; the column is retained,
nullable and unused. All 53 CHECK constraints were then evaluated against all 245
seeded records; the other 52 hold.
- **Unknown fields are rejected, not ignored** (§13.5). A body field that maps to no
column answers **422** with the field named in `details`. The reasoning is
explicit: silently ignoring them is exactly how the three columns above would have
been lost. Server-owned fields (`id`, `org_id`, `created_date`, `updated_date`,
`legacy_id`) remain ignored rather than rejected.
- **Three deliberate divergences from `store.js`** (§13.4): email matching is
case-insensitive because the columns are `citext`; `updated_date` is always
populated because the column is `NOT NULL` and the seeder sets it to
`created_date`; and `_order`, a positional index the frontend used while building
its course list, is dropped.
### Phase 2D — Frontend Transport Integration
This phase belongs to the frontend repository, which is not present in this
checkout. What **is** verifiable here is the backend side of the seam:
- **CORS** (`internal/httpserver/cors.go`). Origins are matched exactly against an
allowlist and echoed back one at a time; `*` is rejected at config validation. A
request whose `Origin` is not on the list is served normally with no CORS headers —
the API does not refuse it, the browser simply will not hand the response to the
page. That distinction is deliberate so that curl, health checkers and
server-to-server callers, which send no `Origin`, are unaffected. With an empty
allowlist the middleware is not installed at all.
- **Development defaults.** With `APP_ENV=development` and `HTTP_CORS_ORIGINS`
unset, the allowlist defaults to `http://localhost:5173`, `http://127.0.0.1:5173`
and the `vite preview` port — both hostnames, because a browser treats them as
different origins. Unset anywhere else it defaults to empty, meaning same-origin
only.
- **JSON everywhere.** `net/http`'s own plain-text 404 and 405 replies are
intercepted and rewritten into the error envelope, so every response from the API
is JSON.
The frontend files named in the contract — `base44Client.js`, `krowHooks.js`,
`store.js` — appear only as references inside this repository's documentation and
comments. `httpClient.js` and any `VITE_API_*` environment variable are **not
referenced anywhere in this repository**: *Not verified from current repository.*
**Known frontend/backend mismatches** recorded in the contract §12: multi-record
writes are not transactional; collection caps truncate silently; all aggregation is
client-side; search is browser-side substring matching; `DELETE` on a missing record
returns 200. See §17.
### Phase 3 — Authentication & Authorization
Split across several steps in the repository's own narrative; what exists today is
described in §8. Chronologically:
**3B — schema and primitives.** Migration **000004** adds two facts and nothing
else: a globally unique index on `users.email` (because the login form supplies no
organization, so an email must resolve to exactly one user), and a `sessions` table.
The migration explicitly declines to add roles, permissions, organization_members,
credentials or refresh_tokens tables. `internal/auth` was written in the same step
and deliberately knows nothing about HTTP: password hashing (`password.go`), token
generation and hashing (`token.go`), the session lifecycle (`session.go`), the
PostgreSQL stores (`store.go`, `users.go`) and credential verification
(`credentials.go`).
**3C — the HTTP surface.** `POST /api/v1/auth/login`, `POST /api/v1/auth/logout`,
and the `authenticate` middleware. This step removed `devOrgMiddleware`, which had
put a fixed organization on every request with no credential behind it. Because the
organization had always been threaded explicitly through the service and repository
boundaries rather than defaulted inside the SQL, that swap changed one line and
nothing below it.
**3D — authorization.** `internal/domain/policy.go` — a hand-written table, kept
separate from the generated descriptors precisely so that regenerating descriptors
can never silently drop an access rule, and a new column can never grant anyone
anything by accident. The handler gate is `Server.authorize` in
`internal/httpserver/api.go`; the row predicates are in `builder.ownership` in
`internal/repo/repo.go`. `docs/api-contract.md` §9A documents the resulting contract.
**Health information exposure reduction.** `GET /health` is public, so its body was
reduced to a single field: `{"status": "ok" | "degraded" | "unavailable"}`. The
PostgreSQL version, database name, schema name, applied migration version, table
count, connection error text, environment and uptime were removed from the response
and moved to the server log — `debug` when healthy, `warn` otherwise. `db.Check`
itself is unchanged and still gathers all of it.
**Future authentication direction** is discussed in §18 and is separated there from
what the code does.
### Phase 4A — Agent & Skill Architecture
Agents and Skills are configuration written as Markdown with YAML frontmatter. Three
tiers exist, and the reasoning for keeping them apart is recorded in migration
000005:
| Tier | Where it lives | Rows in this database |
| --- | --- | --- |
| **shipped** | `src/agents/**/*.md`, `src/skills/**/*.md` in the frontend repo, versioned in Git, bundled at build time | none — they are code, and putting them in a table would trade `git log`, code review and atomic deploy for nothing |
| **organization** | authored in the app, shared across one tenant | `agent_definitions` / `skill_definitions`, `visibility = 'organization'` |
| **personal** | authored in the app, private to one user | `agent_definitions` / `skill_definitions`, `visibility = 'personal'` |
Before this work, the last two lived in `user_preferences.extra` — a jsonb blob with
no owner, no tenancy, no size bound, no server-side validation and no query surface,
returned in full by `GET /api/v1/me` on every page load.
### Phase 4B — Definition Contract
A definition is a Markdown document whose leading fenced block is YAML frontmatter
and whose remainder is the body. The parser contract is stated in
`internal/definition/definition.go`:
- **The document layer** — byte-order mark, line endings, leading blank lines, fence
recognition, body extraction (`frontmatter.go`).
- **The YAML subset** — block maps and sequences, scalars, quoting, comments
(`yaml.go`).
- **Definition-level normalization and validation** — id, name, description, status,
version, pages, icons, reasoning, permissions, starters, knowledge, subagents.
Those cover every column migration 000005 projects out of a definition:
`definition_id`, `status`, `version`, `name`, `description`, `pages`.
Field-by-field detail for both kinds is in §9.
### Phase 4C — Definition Database
Migration **000005** creates `agent_definitions` and `skill_definitions` — two
tables, deliberately, not one. The reasoning is recorded in the migration: agents
have `draft | published | archived` plus a monotonic integer version; skills have
`active | inactive` and **no version at all**, because the frontend has no notion of
a skill version. One table would need a union CHECK permitting "version 5, status
inactive", and a version column forever equal to 1 for half the rows.
The migration also records what it deliberately does not create:
`definition_versions` (nothing retains prior Markdown and no rollback feature exists
to serve), `definition_permissions` (the `permissions:` block stays inside the
Markdown, parsed and unenforced), `agent_skills` / `agent_subagents` (the `skills:`
list names ids in a namespace that includes shipped definitions, which have no rows
here, so a join table would need foreign keys to rows that do not exist),
`agent_knowledge`, and `conversations`.
Schema detail is in §5.
### Phase 4D — Parser Compatibility
Two parsers read the same definition: the JavaScript in the frontend's
`src/lib/skills` and `src/lib/agents`, and the Go in `internal/definition`. If they
disagree, one of two silent failures occurs: a definition the editor accepts and the
API rejects looks valid while it is being written and fails when it is saved; or a
definition the API accepts and the editor rejects is stored and then cannot be
rendered by the product that owns it.
The Go side is therefore a **port**, not an independent implementation, and
compatibility is enforced by replay rather than by description. Detail is in §10.
### Phase 4E — Agent & Skill CRUD
Ten HTTP routes — list, create, get, update, delete for each of agents and skills —
plus `internal/service/definitions.go` and `internal/repo/definitions.go`. Detail is
in §11.
### Phase 4F — Runtime Boundary
`internal/runtime` introduces the runtime representations (`runtime.Agent`,
`runtime.Skill`), the execution value types (`ExecutionInput`, `ExecutionResult`,
`RuntimeError`), a `Loader` that reads a stored definition and re-parses it, status
eligibility rules, dependency resolution with personal-over-organization shadowing,
the `AgentExecutor` / `SkillExecutor` interfaces, an `Engine` that ties them
together, and `UnavailableExecutor`.
**The current runtime establishes the execution boundary.** `UnavailableExecutor`
returns `ErrExecutorUnavailable` from both `ExecuteAgent` and `ExecuteSkill`, and
performs no AI execution of any kind. Detail is in §12.
---
## 3. Current Architecture
### Request path
```
React (krow-demo — separate repository)
|
base44Client.js the one frontend file allowed to change
|
HTTP (JSON, /api/v1, cookie-bearing)
|
===================== krow-backend =====================
|
Go HTTP server net/http.Server + ServeMux
|
requestLogger method, path, status, duration
|
cors installed only when an allowlist is configured
|
recoverer a panic becomes a logged 500, not a dropped connection
|
authenticate cookie -> session row -> user row -> Identity + OrgID
|
jsonErrors rewrites the mux's plain-text 404/405 into the envelope
|
HTTP handlers role gate (403), body decode, path values
|
Service layer query parsing, validation, contract semantics
|
Repository layer all SQL; org + ownership predicates; bind parameters only
|
pgx / pgxpool
|
PostgreSQL
```
### Definition and runtime path
```
Agent/Skill Markdown authored in the app, submitted as a JSON string
|
Parser internal/definition — frontmatter + YAML subset
|
Validator ValidateAgent / ValidateSkill
|
Definition Projection id, status, version, name, description, pages
|
PostgreSQL agent_definitions / skill_definitions
markdown stored verbatim; columns derived from it
|
Definition CRUD /api/v1/agent-definitions, /api/v1/skill-definitions
|
Runtime Loader internal/runtime — load by uuid or definition_id,
re-parse, re-validate
|
Dependency Resolution agent.Skills -> LoadSkill each, tenant + ownership
scoped, personal shadows organization, deduplicated
|
Executor Boundary AgentExecutor / SkillExecutor interfaces
(only UnavailableExecutor exists)
```
### Layer responsibilities
| Layer | Package | Owns | Deliberately does not own |
| --- | --- | --- | --- |
| Entrypoint | `cmd/api` | config load, pool, logger, signal handling, graceful shutdown, session sweeper goroutine | routing, business rules |
| Config | `internal/config` | environment loading, typed defaults, validation | secrets (nothing is baked in) |
| Database | `internal/db` | pool lifecycle, startup ping, `Check` for `/health` | any DDL |
| HTTP | `internal/httpserver` | routing, middleware chain, role gate, response envelopes, cookies, rate limiting | SQL |
| Auth primitives | `internal/auth` | argon2id, token generation/hashing, session lifecycle, credential verification | HTTP, cookies, logging (nothing in the package logs) |
| Identity context | `internal/authctx` | the authenticated `Identity` on the request context | any setter that takes a user id from a client |
| Org context | `internal/orgctx` | the organization a request runs as | deciding which organization that is |
| Service | `internal/service` | query-string parsing, body validation, contract semantics | statement construction |
| Domain | `internal/domain` | resource descriptors (generated), the authorization policy table (hand-written), error constructors, record types | I/O |
| Repository | `internal/repo` | every SQL statement, organization scope, ownership predicates, derived columns | deciding *whether* a role may act |
| Definitions | `internal/definition` | frontmatter, the YAML subset, agent/skill normalization and validation | storage, HTTP, and the `ui:`/`owliver:` blocks |
| Runtime | `internal/runtime` | loading, eligibility, dependency resolution, the executor interfaces | executing anything |
| Seeder | `internal/seeder` | fixture loading, deterministic ids, upsert, shift generation and pruning | schema changes |
| Test harness | `internal/testutil` | a disposable migrated + seeded database per test process | touching the development database |
**Where authorization happens, and why it is split.** The role check runs in the
handler *before any query*, so it always answers 403 and can never reveal whether a
row exists. Organization scope and ownership are SQL predicates in the `WHERE`
clause, so a row outside them is simply absent and answers 404. A caller cannot
distinguish "exists and is not yours" from "does not exist". Because the predicate
is in SQL, `count(*)` runs over the same clause — so `meta.total` for a talent
caller is their own count, not the organization's.
---
## 4. Technology Stack
### Current implementation — verified from this repository
| Concern | What is used | Evidence |
| --- | --- | --- |
| Language | Go, module directive `go 1.27`; local toolchain `go1.27.0 darwin/arm64` | `go-api/go.mod`, `go version` |
| PostgreSQL driver | `github.com/jackc/pgx/v5 v5.10.0` (with `pgxpool`) | `go-api/go.mod` |
| Password hashing | `golang.org/x/crypto v0.42.0` (`argon2`) | `go-api/go.mod`, `internal/auth/password.go` |
| Terminal input | `golang.org/x/term v0.35.0` (hidden password prompt) | `go-api/go.mod`, `cmd/setpassword` |
| Indirect deps | `pgpassfile v1.0.0`, `pgservicefile`, `puddle/v2 v2.2.2`, `x/sync v0.17.0`, `x/sys v0.37.0`, `x/text v0.29.0` | `go-api/go.mod` |
| HTTP server | Go standard library only — `net/http.Server` and `http.ServeMux` with method+pattern routes (`"GET /api/v1/job-postings"`, `"{id}"` path values). No third-party router or framework. | `internal/httpserver/server.go`, `api.go` |
| Logging | `log/slog`, JSON handler to stdout, level from `LOG_LEVEL` | `cmd/api/main.go` |
| Migrations | `golang-migrate` CLI, wrapped by the `Makefile`; `-seq` numbering | `Makefile`, `README.md` (records version 4.19.1) |
| PostgreSQL | **18.6 (Homebrew)** on the local development machine | read-only `SHOW server_version` |
| Applied migration version | **5**, not dirty, in the local development database | read-only `SELECT version, dirty FROM schema_migrations` |
| Extensions installed | `citext`, `plpgsql`. **`pgvector` is not installed.** | read-only `SELECT extname FROM pg_extension` |
| Testing | Go's own `testing` package. No assertion library, no mocking framework, no test containers. Database-backed tests build a disposable database per test process. | `internal/testutil/db.go` |
| Parser | Hand-written, no YAML dependency — a port of the frontend's `yaml.js`. Conformance is asserted against a captured JS oracle fixture. | `internal/definition/yaml.go`, `conformance_test.go` |
| Runtime | Go, in-process, no external calls. Executor is an interface with one stub implementation. | `internal/runtime` |
| Code generation | `scripts/gen_resources.py` (Python 3, shells out to `psql`) | `Makefile` target `gen-resources` |
| Oracle capture | `scripts/oracle.mjs` + `scripts/cases.mjs` (Node, runs the frontend module graph through Vite) | `scripts/` |
**Frontend transport relationship.** The frontend repository is not part of this
checkout. This backend serves the contract in `docs/api-contract.md`; the frontend
consumes it through `base44Client.js`. No frontend source is vendored, copied or
modified here.
### Future / planned direction — not implemented
Nothing in this list is present in the repository. See §18 for detail and §17 for
what the repository does and does not say about each.
- Owliver (the planned Python service), LangGraph, or any orchestration layer
- Any LLM provider client
- RAG, embeddings, `pgvector`
- FastMCP or any tool-execution layer
- Redis, or any shared cache/rate-limit store
- Docker images or compose files
- Object storage
- OpenTelemetry or any tracing/metrics exporter
---
## 5. PostgreSQL & Database Architecture
### Database and schema
- Local development database name and credentials come from environment variables;
`.env` is gitignored and `.env.example` contains only placeholders and safe local
defaults. **No password, key or token appears in any tracked file.**
- The application owns exactly one schema, defaulting to `public`, and
`config.validate()` refuses to point it at a PostgreSQL system schema.
- `migrations/` holds five up/down pairs. The local development database reports
applied version **5**, not dirty.
### Migrations
| File | Introduces |
| --- | --- |
| `000001_initial_schema` | 17 tables, 16 enum types, the `citext` extension, and the index and constraint set |
| `000002_application_interview_id` | `job_applications.interview_id`, `courses.training_outline`, `worker_profiles.score_breakdown` |
| `000003_drop_screened_consistent_check` | drops `job_applications_screened_consistent`; keeps the unused `screened_at` column |
| `000004_auth_sessions` | global unique index on `users.email`; the `sessions` table |
| `000005_agent_skill_definitions` | `agent_definitions`, `skill_definitions` |
Every migration has a matching down file, and a test asserts that (`TestEveryMigrationHasADownFile`). Two further tests assert that 000004 and 000005 are reversible, and one asserts that 000005 adds exactly two tables and no others.
**Migration rules recorded in `README.md`:** files are the source of truth; an
applied migration is immutable (write the next one instead); staging and production
run the same files as a discrete deploy step *before* the new binary rolls out,
which means every migration must be backwards-compatible with the currently-running
version — expand, migrate, contract, never in one migration.
### Tables present in the local `public` schema
Twenty application tables plus `schema_migrations` (21 base tables, verified
read-only):
```
organizations users user_preferences sessions
job_postings job_applications ai_interviews staff
worker_profiles assignments shift_records evidence
user_activity courses learning_paths role_categories
certifications badges agent_definitions skill_definitions
```
### Important tables in plain terms
| Table | What it holds | Notable |
| --- | --- | --- |
| `organizations` | the tenant | every tenant-scoped table references it |
| `users` | people who can sign in | `role` (`admin`/`employer`/`talent`) is the authorization field; `account_type` is a display attribute; `password_hash` is nullable and NULL until set; `email` is `citext` and globally unique since 000004 |
| `sessions` | server-side sessions | stores only SHA-256 of the token; two expiries; `ON DELETE CASCADE` from `users` |
| `user_preferences` | per-user settings | three real columns plus an `extra` jsonb blob |
| `job_postings` | the shop window | `status` enum `draft/active/paused/closed`; `created_by` is server-owned; GIN index on `skill_requirements` |
| `job_applications` | applications | `status` enum carries seven values; `email` is `citext`; `interview_id` is a deliberate soft reference with no FK |
| `ai_interviews` | screening interview results | references an application; ownership is by reference, not by column |
| `worker_profiles` | a worker's own profile | `user_id` is server-owned for talent callers; GIN index on `completed_courses` |
| `staff` | the employment record of the workforce | operator-facing: endorsements, review dates, reviewer names |
| `assignments` | who is on which position | a GiST index over the assignment period |
| `shift_records` | attendance | read-only over the API; the seeder owns these rows |
| `evidence` | proof of work | submitted by a worker, verified by the organization |
| `user_activity` | the audit log | append-only by schema; four server-derived identity columns |
| `courses`, `learning_paths` | the training library | organization-nullable: `NULL org_id` is the shared platform library |
| `badges` | badge definitions | the table exists; **no endpoint reads it** (see §17) |
| `agent_definitions`, `skill_definitions` | authored definitions | see below |
### Conventions
- `uuid` primary keys (`gen_random_uuid()`), plus a nullable unique `legacy_id text`.
- `timestamptz` throughout; `created_date` / `updated_date` keep their frontend names.
- Native enums for closed vocabularies (16 types in 000001), `text` + `CHECK` where
the set is still moving. The definition tables use `text` + `CHECK` for both
`visibility` and `status`.
- Email columns are `citext`.
- 53 CHECK constraints were introduced in 000001; one was dropped in 000003.
### Ownership and visibility model on the definition tables
```
agent_definitions / skill_definitions
org_id NOT NULL, even for a personal definition
(a user belongs to exactly one organization, so the org
predicate applies to every read whether ownership does or not)
visibility 'personal' | 'organization' CHECK
owner_user_id set if and only if visibility='personal' CHECK
((visibility = 'personal') = (owner_user_id IS NOT NULL))
ON DELETE CASCADE — a personal definition dies with its owner
created_by always the author, for attribution
ON DELETE SET NULL — a shared definition survives its author
```
**Uniqueness is per tier, by two partial unique indexes** — not one global unique
constraint, because shadowing by id is the point:
```
..._personal_key UNIQUE (owner_user_id, definition_id) WHERE visibility='personal'
..._org_key UNIQUE (org_id, definition_id) WHERE visibility='organization'
```
Other constraints on both tables: `definition_id ~ '^[a-z0-9][a-z0-9-]*$'`;
`length(markdown) BETWEEN 1 AND 65536`; agent `status IN ('draft','published','archived')`
and `version >= 1`; skill `status IN ('active','inactive')` and no version column.
Indexes: `(org_id, visibility)` for tier listing (which also covers the `org_id` FK),
`(owner_user_id)` for a user's own definitions and the cascade check, and a partial
index for the live rows — `(org_id, visibility) WHERE status='published'` for agents,
and the equivalent active-only index for skills. There is deliberately **no** index
on `created_by`, because it is attribution only.
### Verified seed counts
Counted directly from `seed/fixtures/seed.json` in this repository:
| Entity | Records in fixture |
| --- | --- |
| Course | 40 |
| JobApplication | 24 |
| UserActivity | 15 |
| RoleCategory | 9 |
| WorkerProfile | 9 |
| JobPosting | 8 |
| Certification | 8 |
| AIInterview | 4 |
| Badge | 4 |
| Staff | 3 |
| Evidence | 3 |
| LearningPath | 2 |
| User | 1 |
| Assignment | 0 (empty in the source by design) |
`ShiftRecord` is not in the fixture — it is generated. `docs/api-contract.md` §13.7
records the generated figure as 115 records (2 absent, 1 no-show, 5 late, 107
present) with the qualifier that the set is anchored to the day the seeder runs.
`docs/api-contract.md` §13.7 also records that "6 job postings" in an earlier brief
was the *active* count; the fixture carries 8 (6 active, 1 paused, 1 closed).
---
## 6. Seed & Data Strategy
### Source
`seed/fixtures/seed.json` is **generated**, not written: the frontend's
`src/api/seed.js` is executed through Vite and what it exports is serialised. The
reason is stated in `internal/seeder/seeder.go`: ids, dates, numbers and enum values
then arrive exactly as the demo has them, with no transcription step and nothing to
drift.
### Shift records — generated, not snapshotted
`internal/seeder/shifts.go` is a Go port of the frontend's `attendanceSeed.js`.
Shift records are the one collection whose dates are anchored to *now* rather than to
a fixed calendar: the frontend's `dataResolver.inPeriod` windows every collection on
`created_date`, so a frozen snapshot would read as permanently empty a fortnight
later, and the attendance and overtime features would have nothing to show.
Nothing in the generator is random. The distribution is deterministic given the date
the seeder runs, over a 56-day window, across a three-person roster designed so the
demo has a control, an anomaly and a trend:
| Person | Pattern |
| --- | --- |
| Marco Rivera | the control — reliable, weekend event overtime |
| Marcus Williams | attendance degrading over the last fortnight (the anomaly) |
| Chef Antoine Dubois | present throughout, overtime climbing week on week (the trend) |
### UUID strategy
`seeder.DeterministicUUID` derives a UUID v5 from the source id over a fixed
namespace (`krow-seed-v1...`), implemented with SHA-1 and the RFC 4122 version and
variant bits. Changing the namespace re-keys the entire dataset, so it is a constant
rather than configuration. Foreign key values in the fixture are run through the same
derivation, so a reference resolves to the row it named.
### Idempotency
Explicit upsert inside **one transaction**. Every record's primary key is derived
from its source id, so re-running targets exactly the same rows and
`ON CONFLICT (id) DO UPDATE` restores each one to its seeded values. Consequences,
stated in the package documentation:
- Records created through the API **survive** a re-seed — nothing is deleted.
- A column the fixture does not carry is left as it is.
- Entities are inserted in a fixed order chosen so every foreign key is satisfied by
the time it is referenced.
### The one exception: pruning
Shift records are also **pruned** (`pruneShiftRecords`). Upsert alone cannot converge
a rolling window: yesterday's generated set and today's overlap but are not the same
set, so without a prune, re-seeding on a later day would leave stale rows behind. The
prune is scoped to the seeded organization, and the number pruned is reported in the
seeder's result rather than being silent — a delete during a seed should not be
something you have to read the source to discover.
Four tests cover exactly this behaviour: `TestReseedOnALaterDayLeavesNoStaleShifts`,
`TestReseedIsConvergentAcrossAWeek`, `TestReseedSameInstantPrunesNothing`, and
`TestPruneIsScopedToTheSeededOrganization`.
### Limitations
- The demo user's `password_hash` is **NULL** after seeding. Nothing in the fixture,
the migrations or this repository contains, generates or defaults a password. A
password enters the system only through `cmd/setpassword`, typed by a person or
piped on stdin.
- `Assignment` is empty in the source and stays empty.
- Attendance and overtime are among the richest features in the product and will be
empty in any real deployment until a rostering source exists — an import, an
integration, or a scheduling UI, none of which exist. This is recorded as U1 in the
contract.
---
## 7. API Architecture
### Conventions
| Aspect | Behaviour |
| --- | --- |
| Base path | `/api/v1` |
| Resource naming | kebab-case plural; mass nouns singular (`/staff`, `/evidence`, `/user-activity`) |
| Identifiers | `uuid` in the path; the frontend's original string ids live in `legacy_id` |
| Content type | `application/json; charset=utf-8` both ways |
| Field naming | snake_case, identical to the frontend's names. The single exception is `/me/preferences`, which is camelCase |
| Dates | ISO 8601 with offset, rendered by the database as `YYYY-MM-DDTHH:MM:SS.mmmZ` |
| Partial updates | `PATCH`, shallow merge. There is no `PUT` |
| Reserved query params | `sort`, `limit`, `offset`. Every other parameter is a field filter |
| Request body cap | 4 MiB (`maxBodyBytes`) |
| Cache headers | every JSON response sets `Cache-Control: no-store` |
### Response envelopes
Success — a record:
```json
{ "data": { "id": "…", "title": "…", "created_date": "…" } }
```
Success — a collection:
```json
{
"data": [ … ],
"meta": { "total": 24, "limit": 100, "offset": 0, "returned": 24, "truncated": false }
}
```
`meta.truncated` is `total > offset + returned`. It exists so that the silent
truncation recorded in contract §12.2 is fixable without another contract change.
Failure:
```json
{ "error": { "code": "validation_failed", "message": "…", "details": { "field": "…" } } }
```
`details` is always present, as an object, even when empty.
### Filtering, sorting, pagination
- **Filtering** — a single value is equality (`?status=hired`); a repeated parameter
is membership (`?status=hired&status=interview`). A filter name is resolved to a
real column before any SQL is assembled; an unknown name answers 400. Array and
JSON columns are not filterable and say so.
- **Sorting** — `?sort=field` ascending, `?sort=-field` descending, `?sort=` (empty)
means unordered. Every ordered query renders `ORDER BY <col> <dir> NULLS LAST,
<table>.id` — nulls last in **both** directions, and the id tiebreaker always.
- **Pagination** — `limit` and `offset`, both non-negative integers. `limit` is
clamped to `MaxLimit = 1000`. Each resource's default limit is the literal argument
at its frontend call site.
### Error codes and HTTP status
| Code | HTTP | When |
| --- | --- | --- |
| `invalid_query` | 400 | a query parameter is malformed or names an unknown/unfilterable/unsortable field |
| `unauthorized` | 401 | no valid session |
| `forbidden` | 403 | the caller's role does not permit the operation |
| `not_found` | 404 | no such row within the caller's organization and ownership scope |
| `method_not_allowed` | 405 | the path exists but not under this method |
| `conflict` | 409 | a uniqueness constraint was violated |
| `validation_failed` | 422 | a body field is unknown, blank, null on a NOT NULL column, an invalid enum value, or a required field is absent |
| `rate_limited` | 429 | too many failed sign-in attempts; accompanied by `Retry-After` |
| `internal` | 500 | anything else; detail goes to the log, never to the client |
A resource with no item route at all — `assignments`, which is only listed and
created — answers **404** rather than 405, because 405 requires the path pattern to
exist under some other method.
### Endpoint groups
**Routes registered by the server: 51.**
`routeResources` 34 + `routeMe` 4 + `routeAuth` 2 + `routeDefinitions` 10 +
`GET /health` 1.
`docs/api-contract.md` §2 enumerates 38 endpoints; those are exactly the 34 resource
routes plus the 4 `/me` routes. The 2 auth routes are specified in §9 and the 10
definition routes are not in the contract document at all (see §17).
#### Health
| Method | Path | Notes |
| --- | --- | --- |
| `GET` | `/health` | Public. Body is one field: `{"status": …}`. `"ok"` with 200; `"degraded"` with 200 when the database is up but unmigrated or left dirty; `"unavailable"` with 503 when it is unreachable, so a load balancer can act on the status code alone. Nothing about the server, database or schema appears in the body. |
#### Authentication
| Method | Path | Notes |
| --- | --- | --- |
| `POST` | `/api/v1/auth/login` | Public. Body `{"email", "password", "remember_me"}`. 200 returns the user in the same shape as `GET /me` and sets the session cookie. 422 when a field is absent. 429 when rate limited. Every credential failure is the same 401. |
| `POST` | `/api/v1/auth/logout` | Public and idempotent. Revokes the session behind the cookie if there is one, expires the cookie, and answers 200 with `{"data":{"status":"signed_out"}}` — including when the cookie is absent, stale, or was never valid. |
#### Current user
| Method | Path | Notes |
| --- | --- | --- |
| `GET` | `/api/v1/me` | The user behind the session cookie, with `preferences` embedded |
| `PATCH` | `/api/v1/me` | Shallow merge. Only `full_name` and `account_type` are writable |
| `GET` | `/api/v1/me/preferences` | camelCase keys |
| `PATCH` | `/api/v1/me/preferences` | Shallow merge, returns the merged object |
#### Core application entities
Registered only where the resource declares the operation, so an unsupported one is
answered by the mux rather than by a handler that has to know to refuse.
| Resource | Registered operations |
| --- | --- |
| `job-postings` | list, get, create, update |
| `job-applications` | list, create, update, delete |
| `ai-interviews` | list, create |
| `staff` | list, create, update |
| `worker-profiles` | list, create, update |
| `courses` | list, get, create, update |
| `learning-paths` | list |
| `role-categories` | list, create |
| `certifications` | list, create, delete |
| `user-activity` | list, create |
| `evidence` | list, create, update |
| `assignments` | list, create |
| `shift-records` | list |
| `badges` | none — `Ops: 0` |
#### Agent Definitions
| Method | Path |
| --- | --- |
| `GET` | `/api/v1/agent-definitions` |
| `POST` | `/api/v1/agent-definitions` |
| `GET` | `/api/v1/agent-definitions/{id}` |
| `PATCH` | `/api/v1/agent-definitions/{id}` |
| `DELETE` | `/api/v1/agent-definitions/{id}` |
#### Skill Definitions
| Method | Path |
| --- | --- |
| `GET` | `/api/v1/skill-definitions` |
| `POST` | `/api/v1/skill-definitions` |
| `GET` | `/api/v1/skill-definitions/{id}` |
| `PATCH` | `/api/v1/skill-definitions/{id}` |
| `DELETE` | `/api/v1/skill-definitions/{id}` |
Accepted query parameters on both collections: `visibility`, `status`,
`definition_id`, `sort`, `limit`, `offset`. Any other parameter answers 400. Default
limit 100; default sort `-created_date`. Sortable fields: `created_date`,
`updated_date`, `name`, `definition_id`, `status`, `version`.
#### Runtime
**Runtime execution is currently an internal backend boundary; no public runtime
execution endpoint is implemented.** No route in `internal/httpserver` references the
`runtime` package.
---
## 8. Authentication, RBAC & Multi-Tenancy
### Users
`users` carries `email` (`citext`, globally unique since migration 000004),
`password_hash` (nullable, and NULL until a password is set), `role`, `account_type`,
`status`, `org_id` and `last_login_at`. `User.IsActive()` is `status == "active"`;
`User.CanAuthenticate()` additionally requires a non-empty password hash.
### Password handling
- **argon2id** via `golang.org/x/crypto/argon2`, with the OWASP-recommended
parameters: 64 MiB memory, 3 iterations, 4 lanes, 16-byte salt, 32-byte key.
- Hashes are stored as a complete self-describing PHC record, so verification reads
the parameters out of the stored string rather than assuming today's defaults — a
future cost increase does not invalidate existing hashes. `NeedsRehash` exists and
is tested.
- Minimum length **12 bytes** (a byte floor, because that is what the KDF consumes),
maximum **1024 bytes** (so an unbounded body cannot be hashed at 64 MiB per
attempt).
- Nothing in `internal/auth` logs, and neither a password, a token nor a hash is ever
returned in an error or formatted into a string.
- Passwords enter the system only through `cmd/setpassword`. There is deliberately no
`-password` flag — a password in argv is visible through `ps` and lands in shell
history — so the routes are an interactive hidden prompt (confirmed twice) or
stdin.
### Sessions
- A session is a **row**, not a token payload. The token is 32 random bytes from
`crypto/rand`, encoded as 43 characters of unpadded base64url. The database stores
only its lowercase hex **SHA-256**, pinned by a CHECK constraint — a raw token
fails that pattern, so the catastrophic-and-silent mistake of storing the secret
cannot pass.
- Plain SHA-256 rather than argon2 is deliberate and the reasoning is recorded: a
session token is 256 uniformly random bits, so there is no dictionary to run
against it and no work factor worth paying on every request.
- **Two expiries.** `expires_at` slides forward as the session is used;
`absolute_expires_at` is fixed at creation and never moves. Without the second, the
first could be slid indefinitely.
| Session kind | Idle lifetime | Absolute cap |
| --- | --- | --- |
| normal | 12 hours | 24 hours |
| `remember_me` | 30 days | 90 days |
- Sliding only happens once a session is past a threshold of its idle window, so a
page firing ten requests does not fire ten `UPDATE`s. The idle window is recovered
from the row rather than from the policy, so a session keeps the lifetime it was
issued under.
- A background sweeper deletes expired rows every **15 minutes**, running once
immediately at startup. Sweeping is housekeeping, not correctness: `Authenticate`
already refuses an expired session and deletes the row as it finds it.
- `ON DELETE CASCADE` from `users`, so a deleted user cannot leave a live session
behind.
### Cookie
`krow_session`, `Path=/`, `HttpOnly`, `SameSite=Lax`, `Max-Age` matching the
session's lifetime, and `Secure` whenever `APP_ENV != development`. The `__Host-`
prefix was considered and declined because it *requires* `Secure`, which cannot be
set over plain HTTP on localhost; the hardening is done by attributes instead, where
it can be conditional. `SameSite=Lax` rather than `Strict` (which would drop the
cookie on any cross-site navigation) or `None` (which would require `Secure` and send
the cookie on cross-site POSTs).
**The token is never in a response body.** It goes out in a `Set-Cookie` header and
nowhere else.
### Login flow
1. Parse and validate the *shape* of the request. A missing field is a malformed
request (422), not a failed login.
2. Check both rate limiters, **before** any expensive work — an attacker must not be
able to make the server hash on their behalf.
3. Verify the credentials. `auth.Credentials.Verify` takes the same measurable time
whether the email exists or not: an unknown email, a user with no password set,
and a wrong password all pay a full argon2id derivation against a lazily-built
decoy hash. Account status is checked **after** the password, so a suspended
account is indistinguishable from a wrong password in both answer and timing.
4. Issue the session and set the cookie. A correct password clears the email's
failure counter; the address counter is left alone.
5. `last_login_at` is stamped best-effort, after the session exists — a failure there
is a lost diagnostic, not a reason to refuse a sign-in that already succeeded.
Every credential failure produces one identical 401. The reason (`no_such_user`,
`no_password`, `bad_password`, `not_active`) goes to the log at `warn`, with the
email — which was already in the request — and never the password.
**Rate limiting** is a fixed-window counter of *failed* attempts held in this
process's memory: 5 per email and 20 per client address per 15-minute window. Two
limiters rather than one, because the budgets are deliberately different sizes — an
email is one account, while an address may be a whole office behind NAT. Its
limitations are documented in the source itself and repeated in §17.
### Authentication middleware
```
publicPaths = { /health, /api/v1/auth/login, /api/v1/auth/logout }
```
An **allowlist**, not a list of protected prefixes, so the failure mode of forgetting
to update it is a route that refuses everyone rather than one that serves everyone.
It sits below the router, so even the mux's own 404 is behind it.
For every other request:
```
cookie -> Manager.Authenticate(token)
hash the token, look up the row, refuse if expired (and delete it),
slide the expiry if due
-> users.FindByID(session.UserID)
re-read on EVERY request, not cached in the session row, so suspending
an account takes effect on its next request
-> refuse and revoke if the user is not active
-> authctx.Identity{UserID, OrgID, Email, FullName, Role, AccountType,
Status, SessionID, ExpiresAt} on the context
-> orgctx.With(ctx, user.OrgID)
```
**Nothing in the request influences any of it.** The middleware reads no body, no
query string and no header other than `Cookie`. A `user_id` or `org_id` in a body or
query string is ignored.
### Roles and the authority
`users.role`, and only `users.role`. Three values, fixed by the `users_role_check`
constraint. An unrecognised role authorizes nothing.
| Role | Who |
| --- | --- |
| `admin` | runs the platform for the organization |
| `employer` | runs the organization's hiring and workforce |
| `talent` | a worker, acting for themselves |
`account_type` is **not** an authorization field. It is a display attribute the user
may change on themselves through `PATCH /me`, and nothing in the API reads it to make
a decision.
### Access resolution
```
User (session -> users row)
|
+--> Organization org_id, from the user's row, never from the request
| |
| +--> org predicate on every read and write
| (courses / learning_paths also match org_id IS NULL —
| the shared platform library)
|
+--> Role (admin | employer | talent)
|
+--> handler gate: may this role perform this operation?
| no -> 403, before any query runs
|
+--> row predicate: which rows may this role see?
admin, employer -> the whole organization
talent -> an extra WHERE clause
-> outside it: 404
```
### Permission matrix
R = read (list/get), C = create, U = update, D = delete. `own` means restricted by
a SQL predicate.
| Resource | Admin | Employer | Talent |
| --- | --- | --- | --- |
| `/me`, `/me/preferences` | R U | R U | R U |
| `job-postings` | R C U | R C U | R *(active only)* |
| `job-applications` | R C U D | R C U D | R C *(own)* |
| `ai-interviews` | R C | R C | R C *(own)* |
| `staff` | R C U | R C U | — |
| `worker-profiles` | R C U | R C U | R C U *(own)* |
| `assignments` | R C | R C | R *(own)* |
| `shift-records` | R | R | R *(own)* |
| `courses` | R C U | R | R |
| `learning-paths` | R | R | R |
| `role-categories` | R C | R C | R |
| `certifications` | R C D | R C | R |
| `user-activity` | R C | R C | R *(own)* C |
| `evidence` | R C U | R C U | R C *(own)* |
| `badges` | — | — | — |
Two admin-only operations, each with a reason beyond seniority: `POST`/`PATCH
/courses`, because a course with a `NULL org_id` is the shared library and a write
there can reach beyond the writer's own tenant; and `DELETE /certifications/{id}`,
because deleting one changes what every existing posting that required it means.
### Ownership predicates (talent callers)
| Resource | Predicate |
| --- | --- |
| `worker-profiles` | `user_id = <session user id>` |
| `job-applications` | `email = <session email>` |
| `assignments` | `worker_email = <session email>` |
| `shift-records` | `worker_email = <session email>` |
| `evidence` | `worker_email = <session email>` |
| `user-activity` | `user_email = <session email>` |
| `ai-interviews` | `application_id IN (SELECT id FROM job_applications WHERE org_id = … AND email = <session email>)` |
| `job-postings` | `status = 'active'` — visibility rather than ownership |
The `ai-interviews` case also has a **write** guard: an insert whose ownership is
expressed by reference has the reference checked against the caller's own
applications, and a reference that is not theirs answers the same 404 a nonexistent
application would.
### Server-owned identity columns
Six columns are filled from the session and ignored if present in a request body,
because they are the columns every ownership rule rests on:
| Column | Filled with | When |
| --- | --- | --- |
| `job_postings.created_by` | session user id | always |
| `user_activity.user_id` | session user id | always |
| `user_activity.user_email` | session email | always |
| `user_activity.user_name` | session full name | always |
| `user_activity.account_type` | session account type | always |
| `worker_profiles.user_id` | session user id | talent callers only |
Two more are overridden for talent callers and left writable for operators:
`job_applications.email` and `evidence.worker_email`. The distinction is between *who
acted* and *who the row is about* — when an admin creates a candidate's profile or
files an application on their behalf, the subject is the candidate, not the operator.
`PATCH /me` has its own allowlist: only `full_name` and `account_type` are writable.
`role` used to be in that map, which would have been a one-line privilege-escalation
path the instant sessions existed. Attempts to write a server-owned user field are
ignored and **logged**, so a client asking to change its own role is visible even
though the answer is no.
### Cross-user and cross-tenant protection
Both are predicates, not filters, so a row outside them is never fetched. Tests
covering this: `TestCrossOrganizationIsolation`, `TestTalentSeesOnlyTheirOwnRecords`,
`TestTenancyFollowsTheSession`, `TestIdentityCannotBeSuppliedByTheRequest`,
`TestForbiddenVersusNotFound`, `TestOrganizationScoping`,
`TestTalentCannotInterviewForAnotherApplication`.
### Deny by default
`internal/domain/policy.go` maps a URL path to a `*Policy`. A resource with **no
entry keeps a nil policy and permits nothing, to anyone** — a table added to the
schema tomorrow is unreachable until someone writes down who may reach it.
`TestEveryResourceHasAPolicy` makes the omission loud. `badges` has an empty policy
written out explicitly, so the resource is deliberately closed rather than merely
forgotten.
### What is not implemented
No permissions table, no policy engine, no per-record ACL, no role hierarchy, no
delegation, no refresh tokens, no multi-factor, no SSO, no OAuth, no password reset
flow, no self-service registration, and no email delivery of any kind.
**Firebase Auth is an architectural direction, not the current implementation.** The
string "Firebase" does not appear anywhere in this repository — *Not verified from
current repository* as a recorded plan; it is recorded here because it was named as a
direction for this document.
---
## 9. Agent & Skill System
### The three representations
| Representation | Where | Purpose |
| --- | --- | --- |
| **Source-controlled (shipped)** | `src/agents/**/*.md`, `src/skills/**/*.md` in the frontend repository | product source, versioned in Git, bundled at build time. **No rows in this database.** |
| **PostgreSQL authored** | `agent_definitions`, `skill_definitions` | authored in the app; the Markdown is the authoritative artefact, the columns are projections |
| **Runtime** | `runtime.Agent`, `runtime.Skill` | the parsed definition plus its database identity and ownership, prepared for execution |
The three tiers share one id namespace, and shadowing across them is the point: a
personal definition may carry the same `definition_id` as an organization one, which
may carry the same id as a shipped one.
### Agent — fields parsed by `internal/definition`
`definition.Agent`:
| Field | Type | Notes |
| --- | --- | --- |
| `ID` | string | from `id:`; falls back through name and then the origin path |
| `Name` | string | falls back to `Untitled agent`, which the validator then refuses |
| `Description` | string | |
| `Status` | string | `draft` / `published` / `archived` |
| `Version` | int | |
| `Pages` | []string | **canonical** surface ids — mapped through the frontend's `canonicalPage` |
| `Icon` | string | checked against a closed vocabulary (10 values in the captured oracle) |
| `Reasoning` | string | `fast` / `balanced` / `deep` |
| `Trigger` | string | |
| `WebSearch` | bool | |
| `Skills` | []string | ids of skills the agent depends on |
| `Subagents` | []string | |
| `Starters` | []Starter | `{Label, Prompt}` — conversation starters |
| `Knowledge` | []Knowledge | `{ID, Label, Kind, Body, URL}`; kinds `note` / `link` / `skill-reference` |
| `Permissions` | Permissions | `{Owner, Access, People[]}`; access `all` / `specific`; roles `manager` / `editor` / `viewer`. **Parsed and not enforced** |
| `Instructions` | string | the body's `## Instructions` section — prose belongs under a heading |
| `Errors` | []string | what this definition lost on the way in, carried on the record rather than thrown |
| `Body` | string | the Markdown after the frontmatter |
`runtime.Agent` additionally carries `DatabaseID`, `Visibility`, `OwnerUserID`,
`ResolvedSkills` and `RawMarkdown`.
### Skill — fields parsed by `internal/definition`
`definition.Skill`:
| Field | Type | Notes |
| --- | --- | --- |
| `ID` | string | |
| `Name` | string | falls back to `Untitled skill` |
| `Description` | string | |
| `Status` | string | `active` / `inactive` |
| `Pages` | []string | **as the author wrote them**, not canonicalised |
| `Kind` | string | |
| `Category` | string | |
| `Actions` | []string | |
| `Triggers` | []string | |
| `DeclaredTriggers` | bool | whether the author declared any |
| `Prompt` | *string | nullable |
| `SkillID` | *string | nullable |
| `Levels` | []Level | `{Level, Label, Summary}`, read from the body's own headings |
| `Body` | string | Markdown after the frontmatter, trimmed |
| `Deferred` | []string | frontmatter blocks this package does not check — see §10 |
`runtime.Skill` additionally carries `DatabaseID`, `Visibility`, `OwnerUserID` and
`RawMarkdown`.
### One asymmetry that must not be "fixed"
`Agent.Pages` holds **canonical** surface ids; `Skill.Pages` holds the strings the
author wrote. That is not an oversight: the frontend's `normalizeAgent` maps every
page through `canonicalPage` and its `parseSkill` does not. Making the two agree in
Go would make each one disagree with its own editor. The comment in the source says
so explicitly.
### Vocabularies (captured from the frontend)
18 pages with aliases; agent statuses `draft`/`published`/`archived`; reasoning
`fast`/`balanced`/`deep`; 10 icons; knowledge kinds `note`/`link`/`skill-reference`;
access `all`/`specific`; permission roles `manager`/`editor`/`viewer`.
`TestVocabularyMatchesFrontend` asserts the Go tables against the captured ones.
### Backend-only bounds
Two rules the frontend editor does not have, both flagged `BackendOnly` so they are
distinguishable from shared rules:
- `MaxMarkdownLength = 65536` **characters** (PostgreSQL `length()` counts
characters), matching the `markdown_size` CHECK.
- `MaxVersion = 2147483647`, the range of `agent_definitions.version` as a PostgreSQL
`integer`.
Both turn a constraint violation into a message an author can act on.
---
## 10. Parser & Definition Compatibility
### Why the two parsers must agree
```
Frontend JS parser (src/lib/skills, src/lib/agents)
|
compatibility contract (internal/definition + testdata/oracle.json)
|
Backend Go parser (internal/definition)
```
A definition is authored in the browser and stored by the server, so both parsers see
it. Disagreement is silent in both directions: a definition the editor accepts and
the API rejects fails at save time after looking valid; a definition the API accepts
and the editor rejects is stored and then cannot be rendered.
### The supported YAML subset
`internal/definition/yaml.go` is a line-for-line port of the frontend's `yaml.js`.
**Supported, and nothing else:** block maps and block sequences nested to any depth;
scalars (strings, integers, floats, booleans, null); quoted strings for values
containing `:` or `#`; `- key: value` (a mapping whose first key sits on the dash);
`#` comments and blank lines.
**Not supported:** anchors, aliases, merge keys, multi-document files, flow mappings,
flow sequences, block scalars, tags. These are not silently half-read — an
unparseable line is an error carrying its 1-based line number *within the
frontmatter*, worded exactly as the frontend words it, because an author who sees one
message in the editor and another from the API is being told about two different
problems.
**No YAML dependency, deliberately.** A general library would accept a much larger
language than the frontend does, and every construct it accepted and the frontend did
not would be a definition the backend stores and the editor cannot read.
**Nothing here evaluates anything.** There is no reflection, no template, and no code
path from a definition to execution of any kind. A definition is configuration, and
the parser is the boundary that keeps it configuration.
### Document layer
`frontmatter.go` handles byte-order mark stripping, CRLF normalisation, leading blank
lines, fence recognition (including a trailing tab after the closing fence), and body
extraction. A second `---` in the document is body, not a new frontmatter block.
### Validation and normalization
- Everything is optional. A definition declaring only an id and a name normalizes to
a working agent with documented defaults.
- Nothing unknown survives — statuses, reasoning modes, pages, icons, knowledge kinds
and permission roles are checked against closed tables, and an unrecognised value
is a *named error* rather than a dropped key.
- What validates is kept. One bad entry costs its author that entry and a message,
never the rest of the file (`Agent.Errors`).
- An agent with **no skills** is deliberately not refused, because several product
pages have no assistant skills and answer from their own responder.
- The rejection messages reuse the frontend's own wording wherever the rule is
shared, including its quirks — `ValidateAgent` performs a literal
`strings.Contains(raw, "name:")` substring test because the JavaScript does.
### Raw Markdown is never rewritten
Normalization reads; it does not rewrite what is stored. The Markdown handed in is
the Markdown that goes to the database, byte for byte.
`TestParsingDoesNotMutateSource`, `TestNormalizationIsNotStorage` and
`TestMarkdownIsStoredVerbatim` assert it from three directions.
### Deferred blocks
A skill may carry a `ui:` block (declarative page sections) or an `owliver:` block
(assistant capabilities). Validating those means reproducing roughly 1,500 lines of
closed vocabulary describing what the **frontend** can render — placements, data
sources, section types, periods — none of which the backend stores, projects or acts
on.
The package therefore does not check them. It records their presence on
`Skill.Deferred`, so the gap is a value a caller can see rather than an assumption.
The one consequence is stated exactly in the source: `skill-examples/board-invalid-context.md`
is rejected by the frontend on a rule about which placement may supply which data
source, and accepted here. It is the only definition in the corpus where the two
disagree, and the conformance suite asserts that it stays the only one.
Visibility is deliberately absent from the parser: the frontend ignores a
`visibility:` key in a definition entirely. It is a storage tier chosen by the
request and checked by a database constraint. **A definition cannot name its own
tenancy.**
### The conformance corpus
`internal/definition/testdata/oracle.json` (9,633 lines) is **not hand-written**. It
is captured by `scripts/oracle.mjs`, which loads the real frontend module graph
through Vite — `import.meta.glob`, the `@/` alias and raw Markdown loading behave
exactly as they do in the app — and records what the JavaScript parser did with every
definition. Counted from the fixture in this repository:
| Set | Entries | Breakdown | Accepted by JS |
| --- | --- | --- | --- |
| `corpus` (shipped definitions) | **37** | 28 skills, 9 agents | 36 accepted, 1 rejected |
| `cases` (adversarial) | **132** | 87 skills, 45 agents | 86 accepted, 46 rejected |
The adversarial cases come from `scripts/cases.mjs`, which stores raw bytes as an
author could actually produce them — nothing is normalized on the way in, because the
point is what the two parsers do with the awkward form.
The assertion is therefore not "Go agrees with a description of the frontend" but
"Go agrees with the frontend", replayed. A frontend change that alters parsing fails
these tests, which is the intent: the contract cannot drift silently in either
direction.
### Mutation checks
`TestMutationsWouldBeCaught` exists because tests that pass against a broken parser
are not tests. Each entry is a plausible mistake in the Go package, and each must be
caught by a real definition changing its meaning rather than by an assertion written
to notice it. The mutations covered include: dropping the BOM strip; dropping CR
normalisation (so `candidates\r` stops being a page); trimming the closing fence too
eagerly or not at all; treating a second `---` as frontmatter; treating a `#` inside
a word as a comment; failing to treat a spaced `#` as a comment; letting a colon
inside quotes split the value; taking the first rather than the last of a duplicate
key; accepting ragged indentation instead of reporting its line; and canonicalising a
skill's page name.
### Parser behaviours worth knowing
- `pages: candidates` (a bare string, not a sequence) is **accepted for an agent and
refused for a skill**, because `normalizeAgent` coerces and `parseSkill` requires a
real sequence.
- A definition with no `id:`, no `name:` and a valid `pages:` list gets the id
`custom` on the frontend and is accepted; the Go parser derives the same id from
the same default parameter (`AuthoredPath = "custom"`) so that it does not refuse a
definition the editor accepts.
---
## 11. Agent & Skill CRUD
### Create flow
```
POST /api/v1/agent-definitions { "markdown": "...", "visibility": "personal" }
|
authctx.MustFrom(ctx) identity from the session
|
body["markdown"] present? absent -> 422 "Paste or upload a Markdown definition."
|
ValidateAgent(markdown) frontend rules + backend bounds -> 422 with the message
|
visibility resolved defaults to "personal"; must be personal|organization
|
organization? role gate talent (or an unrecognised role) -> 403
|
ParseAgent(markdown) -> id, status, version, name, description, pages
|
AgentInsertInput org_id = session org
created_by = session user
owner_user_id = session user (personal only)
markdown = the caller's bytes, unchanged
|
INSERT ... RETURNING partial unique index -> 409 on a duplicate id per tier
|
{ "data": { … the stored row … } }
```
Skills follow the same path through `ValidateSkill` / `ParseSkill`, minus `version`.
### List, get, update, delete
| Operation | Behaviour |
| --- | --- |
| `GET` collection | Scoped to the caller's organization **and** ownership: `visibility='organization' OR (visibility='personal' AND owner_user_id = <session user>)`. Optional `visibility` filter narrows to one tier. Returns the standard page envelope with `meta`. |
| `GET /{id}` | Same predicate. A non-UUID id, or a row outside the predicate, answers 404 — the caller cannot distinguish the two. |
| `PATCH /{id}` | Reads the row first (within scope). If the stored row is `organization`, a talent caller is refused 403. `visibility` cannot be changed after creation — an attempt answers 422 `immutable`. If `markdown` is supplied it is re-validated, re-parsed, and **every projection is recomputed from it**. Otherwise a bare `status` change is accepted against the closed status list. |
| `DELETE /{id}` | Idempotent, matching the contract's §12.7 semantics: a non-UUID id or an already-absent row answers 200 with `{"data":{"id":"…"}}`. An `organization` row still enforces the role gate before deleting. |
### Scope, ownership and tenancy
- **Personal scope** — `visibility='personal'`, `owner_user_id` set to the session
user by the server. A caller never supplies it.
- **Organization scope** — `visibility='organization'`, `owner_user_id` NULL,
writable by admin and employer only.
- **Tenant isolation** — `org_id = <session org>` is on every read, update and
delete, including for personal definitions, whose `org_id` is `NOT NULL` for
exactly this reason.
- **Duplicate handling** — the two partial unique indexes make an id unique per owner
or per tenant, never globally. A collision surfaces as `conflict` / 409 through the
repository's error translation.
- **Server-owned fields** — `id`, `org_id`, `owner_user_id`, `created_by`,
`created_date`, `updated_date`, and every projection (`definition_id`, `status`,
`version`, `name`, `description`, `pages`). The only client-supplied values are
`markdown`, `visibility` (at creation), and a bare `status` on `PATCH`.
- **Projection synchronization** — a `PATCH` that changes `markdown` recomputes every
projected column in the same statement, so a column can never describe a different
document than the one stored.
- **Verbatim Markdown** — the bytes the caller sent are the bytes stored.
### Tests covering this surface
`TestAgentCreate`, `TestSkillCreate`, `TestDefinitionsList`, `TestDefinitionsGet`,
`TestDefinitionsPatch`, `TestDefinitionsDelete`, `TestSecurityAndSQLInjection`,
`TestFullCRUDFlowAndProjections` (`internal/httpserver/definitions_api_test.go`, 848
lines).
---
## 12. Runtime Architecture
### What the package contains
| Type | Role |
| --- | --- |
| `Loader` | loads a stored definition by UUID or `definition_id`, re-parses and re-validates it, and returns a runtime representation |
| `runtime.Agent` | parsed agent + `DatabaseID`, `Visibility`, `OwnerUserID`, `ResolvedSkills`, `RawMarkdown` |
| `runtime.Skill` | parsed skill + the same identity and ownership fields |
| `ExecutionInput` | `{Identity, TargetID, Input, Parameters, Context}` |
| `ExecutionResult` | `{Success, Output, AgentID, AgentVersion, ResolvedSkills, Error}` |
| `RuntimeError` | structured `{Code, Message, Target, Cause}` with `Unwrap` |
| `AgentExecutor` / `SkillExecutor` | the execution boundary interfaces |
| `UnavailableExecutor` | the only implementation; refuses execution |
| `Engine` | ties loader + executors together; `RunAgent`, `RunSkill` |
### Typed errors
`ErrNotFound`, `ErrUnauthorized`, `ErrInvalidDefinition`, `ErrDraftAgent`,
`ErrArchivedAgent`, `ErrInactiveSkill`, `ErrNotExecutable`, `ErrDependencyMissing`,
`ErrDependencyInactive`, `ErrCircularDependency`, `ErrExecutorUnavailable`.
### Loading
`LoadAgent` / `LoadSkill` accept either a UUID (matched by regex) or a
`definition_id`, and dispatch to the repository accordingly. The stored `markdown` is
then re-validated and re-parsed — the runtime does not trust the projected columns
for anything but identity, visibility and ownership. A row whose `markdown` is absent
or empty answers `ErrInvalidDefinition`.
Both lookups go through `DefinitionsRepo`, so **tenant and ownership scoping applies
to the runtime exactly as it does to the CRUD API**: `org_id = <session org>` and
`visibility='organization' OR (visibility='personal' AND owner_user_id = <session user>)`.
### Status eligibility
```
Agent Skill
published -> eligible active -> eligible
draft -> ErrDraftAgent inactive -> ErrInactiveSkill
archived -> ErrArchivedAgent other -> ErrNotExecutable
other -> ErrNotExecutable
```
`LoadExecutableAgent` applies the agent rule and then resolves dependencies;
`LoadExecutableSkill` applies the skill rule.
### Dependency resolution
```
Agent
|
agent.Skills (ids from the `skills:` frontmatter)
|
for each id, in declared order:
|
+-- already resolved? -> skip (deterministic deduplication)
|
+-- LoadSkill(ident, id) tenant + ownership scoped
| |
| +-- by definition_id: ORDER BY
| CASE WHEN visibility='personal' THEN 1 ELSE 2 END
| LIMIT 1 personal shadows organization
| |
| +-- not found -> ErrDependencyMissing
|
+-- status != 'active' -> ErrDependencyInactive
|
v
agent.ResolvedSkills ([]*Skill, in declared order)
```
An agent with no `skills:` entries resolves to an empty list rather than an error.
**Cross-tenant dependency protection** falls out of the repository predicate: a skill
belonging to another organization is simply not found, so the dependency resolves to
`ErrDependencyMissing` rather than to another tenant's definition.
`TestRuntime_DependencyResolution` covers the cross-org case explicitly, and
`TestRuntime_PersonalSkillShadowing` covers the shadowing precedence.
**Cycle detection** is present in the code (`inProgress` map, `ErrCircularDependency`)
but is not reachable in the current design — see §17.
### The executor boundary
```go
type AgentExecutor interface {
ExecuteAgent(ctx, agent *Agent, input ExecutionInput) (*ExecutionResult, error)
}
type SkillExecutor interface {
ExecuteSkill(ctx, skill *Skill, input ExecutionInput) (*ExecutionResult, error)
}
```
`NewEngine` defaults both to `UnavailableExecutor` and accepts `WithAgentExecutor` /
`WithSkillExecutor` overrides. `UnavailableExecutor.ExecuteAgent` returns a result
carrying the agent id, version and resolved skill ids, with `Success: false` and
`Error: ErrExecutorUnavailable`, and returns that error.
**Real LLM execution is not part of the current executor boundary.** There is no LLM
client, no prompt assembly, no tool dispatch and no external call anywhere in
`go-api/`. The boundary exists so that adding one is an implementation of an
interface rather than a change to the loading, scoping and eligibility rules.
### Reachability
Nine test functions in `internal/runtime/runtime_test.go` (890 lines) exercise the
package: loader, status eligibility, version semantics, dependency resolution,
executor boundary, personal skill shadowing, malformed Markdown, and skill execution.
Outside those tests, **nothing in the repository calls into the package** — see §17.
---
## 13. Security Architecture
### Controls that are present and verifiable
| Control | How it is implemented |
| --- | --- |
| **SQL parameterization** | Every value reaches PostgreSQL as a bind parameter cast to its declared type. No identifier ever comes from user input: a filter or sort name is resolved to a `*domain.Column` before any SQL is assembled, and an unresolved name is rejected. `TestSecurityAndSQLInjection` covers the definitions surface. |
| **Deny-by-default authorization** | A resource with no policy permits nothing to anyone. `TestEveryResourceHasAPolicy`, `TestNilPolicyDeniesEverything`, `TestUnknownRoleIsDenied`. |
| **Tenant isolation** | `org_id = <session org>` is a predicate on every read and write, taken from the user's row and never from the request. |
| **Ownership checks** | Predicates in the same `WHERE` clause, so `count(*)` matches what the caller may see and no unauthorized row is ever fetched. |
| **Server-owned fields** | Eight identity columns are filled from the session and override anything in the body; `PATCH /me` writes only two fields, and attempts on the rest are ignored **and logged**. |
| **Password handling** | argon2id at OWASP parameters, self-describing PHC records, a 12-byte floor and a 1024-byte ceiling, no password ever logged or echoed, and no `-password` CLI flag. |
| **Credential-failure uniformity** | One 401 for every failure mode, and a decoy argon2id derivation so an unknown email costs the same time as a known one. Account status is checked after the password. |
| **Sessions** | 256-bit random opaque tokens; only SHA-256 stored, pinned by a CHECK; sliding plus absolute expiry; per-request user re-read so suspension takes effect immediately; expired sessions deleted on contact and swept every 15 minutes. |
| **Cookies** | `HttpOnly`, `SameSite=Lax`, `Path=/`, `Secure` outside development, `Max-Age` matching the session. The token never appears in a response body. |
| **Rate limiting** | Failed sign-in attempts counted per email (5) and per client address (20) in a 15-minute window, checked **before** any hashing. |
| **CORS** | Exact-match allowlist, echoed one origin at a time, `*` rejected at config load, `Vary: Origin` always set, and the middleware not installed at all when the allowlist is empty. |
| **Input validation** | Unknown body fields answer 422 with the field named; enum values are checked against the column's declared set; `NOT NULL` columns refuse an explicit null; required fields are enforced on create; request bodies are capped at 4 MiB. |
| **Health endpoint exposure** | One status word and nothing else. Version, database name, schema, migration version, table count, error text, environment and uptime all moved to the server log. `TestHealthLeaksNoInfrastructure`, `TestHealthUnavailableSaysNothingAboutWhy`. |
| **Error exposure** | An unrecognised error is flattened to `{"code":"internal","message":"internal error"}`; the detail goes to the log. A 403 names neither the caller's role nor the roles that would have worked. |
| **Existence non-disclosure** | 403 for role, 404 for row. A caller cannot tell "exists and is not yours" from "does not exist". `TestForbiddenVersusNotFound`. |
| **Panic containment** | `recoverer` turns a panic into a logged 500 rather than a dropped connection. |
| **Config guards** | System schemas rejected; `sslmode=disable` rejected in production; CORS `*` rejected. |
| **Operational guards** | No `make migrate-drop`; `migrate-down` gated on `APP_ENV=development`; test databases are named `krow_backend_autotest_<pid>` and never touch the development database. |
| **Secret hygiene** | `.env` is gitignored; `.env.example` carries no real values; the connection string is redacted in logs (`DBConfig.Redacted`). |
### Honest limitations
These are the repository's own words where it states them, and direct observation
where it does not:
- **Rate limiting is per-process and in-memory.** Two API instances behind a load
balancer each allow the full budget, so the effective limit is the limit times the
instance count, and a restart clears every counter. Multi-instance deployment needs
shared state.
- **The client address is `RemoteAddr`.** Behind a reverse proxy every request appears
to come from the proxy, so the per-address budget becomes global. Reading
`X-Forwarded-For` instead would be *worse* until a trusted-proxy list exists,
because a client can send that header itself and mint a fresh budget per request.
- **The limiter map is bounded by pruning, not by a hard cap**, so a flood from many
distinct addresses grows it until the next prune.
- **There is no CSRF token.** The defence is `SameSite=Lax`, which does not send the
cookie on cross-site POST/PATCH/DELETE. That is adequate for the current
same-origin and localhost-development posture and would need re-examination
alongside any change to the cookie's `SameSite` value.
- **CORS never sets `Access-Control-Allow-Credentials`.** See §17.
- **The `permissions:` block in a definition is parsed and not enforced.** No
`definition_permissions` table exists.
- **No rate limiting exists on any endpoint other than login.**
- **No audit trail of authorization refusals beyond the application log.**
- **No secret management, key rotation, or encryption at rest** is configured in this
repository; those are deployment concerns and `infrastructure/` is empty.
Nothing here should be read as a claim that the service is production-hardened. It is
a claim about which controls exist and which do not.
---
## 14. Testing & Verification
### Toolchain checks — run against this repository
| Check | Command | Result observed |
| --- | --- | --- |
| Formatting | `gofmt -l ./cmd ./internal` | no files listed |
| Static analysis | `go vet ./...` | no diagnostics |
| Build | `go build ./...` | builds |
| Tests | `go test ./... -count=1` | every package's tests ran without failure |
### Test counts — counted from this repository
| Measure | Count |
| --- | --- |
| Top-level `Test…` functions | **175** |
| Total test entries executed, including subtests | **920** |
| Failures | **0** |
| Skipped | **0** (PostgreSQL was reachable on the machine used) |
Per package:
| Package | Top-level test functions | Lines of test code (largest files) |
| --- | --- | --- |
| `internal/httpserver` | 78 | `api_test.go` 1,114; `auth_test.go` 926; `definitions_api_test.go` 848; `rbac_test.go` 730 |
| `internal/auth` | 26 | `session_test.go` 488; `schema_test.go` 376; `password_test.go` 254 |
| `internal/domain` | 22 | `definitions_schema_test.go` 748; `policy_test.go` 191 |
| `internal/seeder` | 18 | `seeder_test.go` 300; `shifts_convergence_test.go` 178 |
| `internal/definition` | 13 | `conformance_test.go` 928 |
| `internal/runtime` | 9 | `runtime_test.go` 890 |
| `internal/service` | 9 | `service_test.go` 217 |
`cmd/api`, `cmd/seed`, `cmd/setpassword`, `internal/authctx`, `internal/config`,
`internal/db`, `internal/orgctx`, `internal/repo` and `internal/testutil` have no test
files of their own; `internal/repo` is exercised through the service and httpserver
suites.
### Database-backed tests
`internal/testutil` builds a disposable database per test **process**: dropped,
recreated, migrated and seeded, named `krow_backend_autotest_<pid>` so that
concurrently-running test packages cannot collide, and named distinctly enough that
it cannot be confused with a real database. The package documentation states that
nothing in it ever connects to, reads or drops the development database. Tests skip
rather than fail when PostgreSQL is unreachable.
### What the important tests prove
| Area | Representative tests | What they establish |
| --- | --- | --- |
| Contract semantics | `TestNullsSortLastInBothDirections`, `TestStableSortWithIDTiebreaker`, `TestEndpointSpecificDefaults`, `TestLimitAndTruncationMeta`, `TestOffsetPaging` | The two ordering guarantees hold, and each endpoint carries its own defaults. The README records that the ordering tests were verified to fail when the guarantee is removed. |
| Validation | `TestCreateRejectsUnknownFields`, `TestCreateRejectsMissingRequiredAndBlank`, `TestCreateRejectsInvalidEnum`, `TestCreateIgnoresServerOwnedFields` | Unknown fields are a 422 rather than silent data loss; server-owned fields are ignored. |
| Idempotent delete | `TestDeleteIsIdempotent` | 200 whether or not a row matched, as §12.7 requires. |
| Seed fidelity | `TestSeedMatchesFixtureCounts`, `TestSeedPreservesSourceValues`, `TestSeedRegressionAnchors` | The database is compared against the fixture field-by-field, not against numbers typed into a test. |
| Seed convergence | `TestSeedIsIdempotent`, `TestReseedOnALaterDayLeavesNoStaleShifts`, `TestReseedIsConvergentAcrossAWeek`, `TestPruneIsScopedToTheSeededOrganization` | Re-seeding converges the rolling shift window without touching other organizations or API-created rows. |
| Password and session | `TestDefaultPasswordParamsMeetOWASP`, `TestSessionCannotOutliveItsAbsoluteDeadline`, `TestDatabaseRefusesARawToken`, `TestUserDeletionCascadesToSessions` | The cost parameters, the absolute cap, the CHECK that refuses a raw token, and the cascade all hold at the database level. |
| Authentication behaviour | `TestLoginFailuresAreIndistinguishable`, `TestLoginSetsHardenedCookieAndNeverReturnsTheToken`, `TestCookieIsSecureOutsideDevelopment`, `TestSessionSlidesButNotForever`, `TestSuspendedUserIsRejectedMidSession`, `TestLoginRateLimitIsPerEmailAndPerAddress` | Enumeration resistance, cookie hardening, sliding-with-a-ceiling, mid-session suspension, and both limiter dimensions. |
| Authorization | `TestRoleMatrix`, `TestTalentSeesOnlyTheirOwnRecords`, `TestCrossOrganizationIsolation`, `TestForbiddenVersusNotFound`, `TestServerOwnedIdentityCannotBeSupplied`, `TestTalentCannotInterviewForAnotherApplication`, `TestTalentSeesOnlyActivePostings`, `TestExpiredSessionIsRefusedBeforeRoleCheck` | The full matrix, the row predicates, the 403/404 distinction, and that identity cannot be supplied by the request. |
| Policy invariants | `TestEveryResourceHasAPolicy`, `TestDerivedColumnsAreReadOnlyOrTalentScoped`, `TestOnlyTalentIsRowScoped`, `TestTalentScopesNameRealColumns` | Regenerating the descriptors cannot silently drop a rule, and a scope cannot name a column that does not exist. |
| Schema invariants | `TestDefinitionTablesShape`, `TestUniquenessPerTier`, `TestForeignKeysAndDeleteBehaviour`, `TestMigrationAddsExactlyTwoTables`, `TestMigration000005IsReversible`, `TestEveryMigrationHasADownFile` | The definition schema is asserted against the live database, including the deferred tables that must **not** exist. |
| Parser conformance | `TestFrontmatterTreeParity`, `TestAcceptanceParity`, `TestRejectionMessageParity`, `TestProjectionParity`, `TestAdversarialCoverage`, `TestDeferredBlocksAreReported` | Go matches the captured JS output over 37 shipped definitions and 132 adversarial cases, including the exact rejection wording. |
| Parser anti-regression | `TestMutationsWouldBeCaught`, `TestParsingDoesNotMutateSource`, `TestNormalizationIsNotStorage` | Plausible parser mistakes are caught by real definitions changing meaning, and normalization never rewrites what is stored. |
| Runtime | `TestRuntime_StatusEligibility`, `TestRuntime_DependencyResolution`, `TestRuntime_PersonalSkillShadowing`, `TestRuntime_ExecutorBoundary`, `TestRuntime_MalformedMarkdown` | Eligibility, resolution across tenants and tiers, shadowing precedence, and that the boundary refuses rather than pretends. |
| Health | `TestHealthEndpoint`, `TestHealthLeaksNoInfrastructure`, `TestHealthUnavailableSaysNothingAboutWhy` | The public body carries a verdict and no reconnaissance. |
| CORS | `TestCORSAllowsConfiguredOrigin`, `TestCORSPreflight`, `TestCORSRefusesUnknownOrigin`, `TestCORSIgnoresRequestsWithoutOrigin`, `TestCORSOffByDefault` | Exact matching, and that an origin-less caller is unaffected. |
---
## 15. Frontend Integration
```
React (krow-demo, a separate repository — not part of this checkout)
|
base44Client.js the transport seam; the one frontend file allowed to change
|
HTTP client Not verified from current repository
|
Go API /api/v1, JSON, session cookie
|
PostgreSQL
```
### What is verifiable from this repository
- **API base path** — `/api/v1`. The server binds `HTTP_HOST:HTTP_PORT`, defaulting
to `127.0.0.1:8080`.
- **CORS** — an exact-match allowlist from `HTTP_CORS_ORIGINS`. In development with
the variable unset, the defaults are `http://localhost:5173`,
`http://127.0.0.1:5173` and the `vite preview` port. Unset outside development the
allowlist is empty and the middleware is not installed, which is same-origin only.
- **Transport compatibility** — field names are snake_case and identical to the
frontend's; ids are opaque strings to the client; the response envelope is one
`.data` unwrap; `meta.total` exists because the frontend's bare array shape had
nowhere to put it.
- **Preserved client interface** — the contract's acceptance criterion is that
swapping the transport inside `base44Client.js`, and changing no other frontend
file, leaves the application behaving identically. The contract states that if
implementing it required editing `krowHooks.js`, a page or a component, the
contract was wrong and got fixed first.
- **Connected entities** — the 13 resources with registered routes, plus `/me` and
`/me/preferences`.
- **The `{ persisted }` return shape** — `auth.updatePreferences()` returns
`{ user, persisted, error }` rather than a bare user, because a swallowed
`QuotaExceededError` used to lose account-authored skills silently. Over HTTP a
failed write is already a non-2xx, so the shim synthesises the success shape.
### What is not verifiable here
- `httpClient.js` — **Not verified from current repository.** The string does not
appear anywhere in this checkout.
- Any `VITE_API_*` environment variable — **Not verified from current repository.**
- The current state of the frontend's own tests — **Not verified from current
repository.** The frontend repository is not part of this checkout.
### Known integration mismatches
Recorded in `docs/api-contract.md` §12 and still accurate against the current code:
| # | Mismatch |
| --- | --- |
| 12.1 | Four flows write several records in a client-side loop with no transaction and no rollback (`useHireCandidate`, `useAssignWorkers`, `useScreenAllCandidates`, `useSubmitChallenge`). A failure halfway leaves the database inconsistent. `useAssignWorkers` issues 3n sequential round-trips for n workers. |
| 12.2 | Collection caps truncate silently. `worker-profiles` caps at 500 sorted by `-krow_score`, so profile 501 is invisible. `meta.truncated` exists to make this fixable; nothing consumes it yet. |
| 12.3 | All aggregation — funnels, KPI tiles, charts, attendance rollups, overtime — computes in the browser over the capped arrays. |
| 12.4 | Search is browser-side substring matching over a concatenated string. No endpoint implements search. |
| 12.5 | Email matching is case-insensitive server-side (`citext`) where `store.js` used `===`. Deliberate and documented as safe. |
| 12.7 | `DELETE` on a missing record returns 200, because the frontend deletes inside loops without checking and a 404 would surface an error toast where none appears today. |
| 12.8 | `InvokeLLM` and `UploadFile` remain local to the browser. Nine AI workflows run deterministically client-side. No endpoint replaces them. |
| §13.4 | `updated_date` is now always populated where `store.js` returned `undefined`; one visible effect is that `TalentDetailModal.jsx:66` shows the seeded date rather than today. |
---
## 16. Important Architectural Decisions
| Decision | Why | Current implementation | Future direction |
| --- | --- | --- | --- |
| **Go for the API** | A single static binary, a strong standard-library HTTP server, and explicit error handling for a service whose main job is correctness at a boundary | **Current.** Go 1.27, `net/http` only, three direct dependencies | — |
| **PostgreSQL, no ORM** | The contract's semantics (`NULLS LAST` both ways, the `, id` tiebreaker, shallow PATCH, idempotent DELETE) are easier to guarantee in SQL written once than to coax out of a mapper | **Current.** `internal/repo` builds every statement from a descriptor | — |
| **Descriptors generated from the live schema** | Column names, types, enum values and nullability then cannot drift from the migrations | **Current.** `make gen-resources` → `resources_gen.go` | — |
| **Policy hand-written, kept apart from descriptors** | Regenerating descriptors must never silently drop an access rule, and a new column must never grant anyone anything by accident | **Current.** `internal/domain/policy.go`, enforced by `TestEveryResourceHasAPolicy` | — |
| **One shared query builder, not fourteen repositories** | The awkward semantics get implemented once and apply identically everywhere | **Current.** `internal/repo/repo.go` | — |
| **Migrations are the source of truth** | If the database and `migrations/` disagree, `migrations/` is right; an applied migration is immutable | **Current.** Five pairs; nothing in `go-api/` issues DDL | Expand-migrate-contract as a discrete CI deploy step before the binary rolls out |
| **`org_id` on every tenant-scoped table** | Tenancy has to be a predicate on the row, not a convention in the caller | **Current.** Server-owned on every resource; `NOT NULL` even on personal definitions | — |
| **Ownership as a SQL predicate, not a filter** | `count(*)` runs over the same clause, so totals are correct and unauthorized rows are never fetched | **Current.** `builder.ownership` | — |
| **403 for role, 404 for row** | A caller must not be able to distinguish "exists and is not yours" from "does not exist" | **Current.** Role gate in the handler before any query | — |
| **API contract derived from call sites** | A table existing is never a reason for an endpoint to exist | **Current.** `docs/api-contract.md`; `badges` has no endpoints | Multi-record transactional endpoints would require touching frontend hooks, so they were deliberately left out of v1 |
| **The transport boundary is one frontend file** | The whole migration is affordable only if nothing above `base44Client.js` changes | **Current.** Envelope, snake_case fields, opaque ids, per-endpoint defaults | — |
| **Sessions are rows; only SHA-256 is stored** | A dump of the sessions table must not be replayable as a login | **Current.** Opaque 256-bit tokens, HttpOnly cookie, CHECK-pinned hash format | — |
| **Two expiries per session** | A sliding window alone can be slid forever | **Current.** 12h/24h and 30d/90d | — |
| **`users.role` is the sole authority** | `account_type` is user-writable through `PATCH /me` and cannot be an authorization field | **Current.** Documented in contract §9A.1 | — |
| **Verbatim Markdown as the authoritative artefact** | A definition must survive a round trip to a `.md` file on disk unchanged; every column is a cache of it | **Current.** `markdown` stored byte-for-byte; projections recomputed on every write | — |
| **JS/Go parser parity enforced by replay** | Two parsers that disagree fail silently in both directions | **Current.** `testdata/oracle.json` captured from the real frontend module graph; 37 corpus + 132 adversarial cases | Regenerate the oracle after any change to `src/lib/skills` or `src/lib/agents` |
| **No YAML library** | A general parser accepts a larger language than the frontend does; every extra construct is a definition the editor cannot read | **Current.** Hand-ported subset in `yaml.go` | — |
| **Personal vs organization definitions, two partial unique indexes** | Shadowing by id is the point, so the id must be unique per tier and never globally | **Current.** `..._personal_key` and `..._org_key` | — |
| **No skill versioning** | The frontend has no notion of a skill version and cannot set one; inventing it would create a column forever equal to 1 | **Current.** `skill_definitions` has no `version` column, asserted by `TestSkillStatusAndNoVersion` | Would require an explicit architecture decision |
| **Two definition tables, not one** | Agents and skills do not share a lifecycle; one table would need a union CHECK permitting nonsense states | **Current.** `agent_definitions`, `skill_definitions` | — |
| **`permissions:` parsed but not enforced** | Its semantics are not defined yet, and a half-enforced permission model is worse than none | **Current.** Parsed onto `Agent.Permissions`; no `definition_permissions` table | Define the semantics before storing them |
| **Runtime execution boundary before any executor** | Loading, scoping, eligibility and resolution are worth getting right independently of what eventually executes | **Current.** `AgentExecutor`/`SkillExecutor` interfaces; `UnavailableExecutor` refuses | Real executors attach here |
| **Health endpoint says the verdict, not the reasoning** | An unauthenticated endpoint is a public document, not an operator console | **Current.** One status field; full detail in the server log | — |
| **Firebase Auth** | — | **Future direction.** The repository contains no reference to Firebase; the current implementation is server-side sessions with argon2id | Would replace or front the current credential path |
| **Google Pub/Sub** | — | **Future direction.** No reference in the repository | — |
| **GCS-compatible object storage** | — | **Future direction.** No reference in the repository; `infrastructure/README.md` names MinIO as a "later" candidate | — |
| **Docker deployment** | — | **Future direction.** `infrastructure/` contains only a README; `Dockerfile.api` and `Dockerfile.owliver` are listed there as later work | — |
| **Redis** | Shared state for rate limiting across instances | **Future direction.** `ratelimit.go` names Redis or the database as the seam; nothing is wired | Needed before multi-instance deployment for the limiter to mean anything |
| **NATS is not part of the target architecture** | — | **Future direction / stated rule.** `infrastructure/README.md` previously listed NATS as a later docker-compose candidate; that table has been corrected. See §17.17. | Keep it out |
| **RAG / pgvector** | — | **Future direction.** `pgvector` is not installed in the local database and is named only once, in `infrastructure/README.md` | Introduce when the architecture calls for it |
---
## 17. Known Issues & Current Limitations
Everything below was checked against the current repository. Items the request asked
about that could not be verified are marked as such rather than guessed at.
### 17.1 Documentation is behind the code — resolved
**Issue as recorded.** `README.md` described the repository as being at an earlier
stage than the code is. It stated "Authorization — which roles may do what — is Phase
3D and is not implemented", "Authorization is **not** implemented. `users.role` is
carried on the identity and consulted nowhere: any signed-in user reaches every
endpoint", "38 endpoints" and "55 tests". The same staleness appeared in two source
comments: the `internal/httpserver/server.go` package documentation ("Authorization is
NOT here"), and `internal/authctx/authctx.go` ("It is NOT consulted anywhere in Phase
3C: authentication only").
**Cause.** Authorization, the definition tables, the CRUD surface and the runtime all
landed after the README was last revised.
**Current behaviour.** `internal/domain/policy.go`, `Server.authorize`,
`builder.ownership` and `internal/httpserver/rbac_test.go` (730 lines) all exist and
run. 51 routes are registered. 175 test functions run. `docs/api-contract.md` §9A
*is* current and documents the authorization contract accurately.
**Resolution.** `README.md` was revised against the code: the status paragraph, the
layout tree (which was missing `cmd/setpassword`, `internal/auth`, `internal/authctx`,
`internal/definition` and `internal/runtime`), the route count, the test count and the
authorization section. The `server.go` and `authctx.go` package comments were
corrected to describe the authorization that exists. `internal/orgctx/orgctx.go`,
which still framed itself as pre-authentication, was corrected in the same pass.
### 17.2 The definitions endpoints are not in the API contract
**Issue.** `docs/api-contract.md` calls itself "the frozen client contract" and
contains **zero** occurrences of `agent-definitions` or `skill-definitions`, yet ten
routes are registered.
**Cause.** The contract was revised for authorization (§9A) but not for the
definitions surface.
**Impact.** A frontend engineer working from the contract would not know these
endpoints exist, and their request/response shapes, filters and error codes are
documented only in Go source and tests.
**Current behaviour.** The endpoints work and are covered by 8 test functions.
**Possible future handling.** Add a definitions section to the contract.
### 17.3 A referenced document does not exist
**Issue.** `internal/definition/definition.go` states that backend-only rejection
rules "each one is listed in docs/phase-4d-parser-contract.md", and
`internal/definition/skill.go` refers to "the contract document". `docs/` contains
only `api-contract.md` and this file.
**Impact.** The two `BackendOnly` rules are discoverable only by reading
`ValidateAgent` and `ValidateSkill`.
### 17.4 The definitions surface bypasses the policy table
**Issue.** The ten definition handlers do not call `Server.authorize` and the
`policies` map has no entry for `agent-definitions` or `skill-definitions`. Role
checks are instead written inline in `internal/service/definitions.go` — the same
`if !known || role == domain.RoleTalent { return nil, domain.Forbidden() }` block
appears eight times.
**Cause.** Definitions are not `domain.Resource` values, so the descriptor-and-policy
machinery does not reach them.
**Impact.** The deny-by-default guarantee — enforced by `TestEveryResourceHasAPolicy`
— does not extend to the newest ten endpoints. A future definition-related endpoint
added without its inline check would be reachable by any signed-in user, and no test
would notice.
**Current behaviour.** Tenancy and ownership *are* enforced, in the repository
predicates, for every definition read and write. The role rule that exists —
organization-visibility definitions are writable by admin and employer only — is
applied on create, update and delete, and is covered by tests.
**Possible future handling.** Bring definitions under the policy table, or add a test
that asserts every definition route has an explicit role check.
### 17.5 The runtime package is not reachable from the running service
**Issue.** No code outside `internal/runtime` and its own test file references the
package. There is no HTTP route, no command, and no service that constructs an
`Engine`.
**Impact.** The runtime's loading, scoping, eligibility and resolution logic is
exercised only by its tests. Nothing a deployed binary does touches it.
**Current behaviour.** As stated in §7: runtime execution is an internal backend
boundary and no public runtime execution endpoint is implemented.
### 17.6 Cycle detection in dependency resolution is unreachable
**Issue.** `Loader.ResolveAgentDependencies` maintains an `inProgress` map and returns
`ErrCircularDependency`, but sets `inProgress[skillID] = true` and back to `false`
within the same loop iteration, and never recurses — skills do not resolve skills.
The error is therefore not reachable, and no test covers it.
**Cause.** The guard is shaped for a recursive resolver; the resolver is one level
deep.
**Impact.** A genuine cycle cannot exist in the current model (an agent names skills;
skills name nothing), so nothing is currently mis-handled. The code implies a
protection that is not active.
**Possible future handling.** Either make resolution recursive, at which point the
guard becomes live and needs a test, or remove it.
### 17.7 Unchecked type assertions in the runtime loader
**Issue.** `internal/runtime/loader.go` lines 67, 72, 129 and 133 perform
`rec["id"].(string)` and `rec["visibility"].(string)` without the comma-ok form. Every
other map access in the codebase, including `internal/repo/repo.go:346`, uses the
defensive form.
**Impact.** A nil or non-string value would panic. The panic would be caught by
`recoverer` and become a logged 500 — but only if the runtime were reachable over
HTTP, which it is not.
**Current behaviour.** The repository projects both columns as `::text`, so the
assertion holds in practice.
### 17.8 CORS does not permit credentials
**Issue.** `internal/httpserver/cors.go` never sets
`Access-Control-Allow-Credentials: true`. Its own comment anticipates this — "so
adding credentials later does not require rewriting this" — but the header was not
added when cookie authentication landed.
**Impact.** A browser at `http://localhost:5173` calling the API at
`http://127.0.0.1:8080` with `credentials: 'include'` will have the response blocked
by the browser, even though the server answered correctly. This affects exactly the
cross-origin development setup the CORS middleware exists for. A same-origin
deployment is unaffected.
**Current behaviour.** Five CORS tests exist and pass; none asserts anything about
credentials.
**Possible future handling.** Set the header for allowlisted origins, and add a test.
### 17.9 Login rate limiting does not survive scale-out
Covered in §13 under limitations: per-process, in-memory, `RemoteAddr`-based, pruned
rather than capped. The source documents all three limitations itself. **Impact:**
the effective limit multiplies by instance count, resets on restart, and degrades to
a global budget behind a reverse proxy. **Possible future handling:** move the
limiter behind shared state before deploying more than one instance.
### 17.10 The repository has almost no history
**Issue as recorded.** `git log` reported `fatal: your current branch 'main' does not
have any commits yet`, and every file was untracked.
**Current behaviour.** The tree is now committed: a single commit on `main`
(`7d12ebe`, "first commit") holds the whole repository.
**Remaining impact.** One commit is not history. There is still no blame, no
incremental recovery point, and no record of when any of the work described in §2
happened. The chronology in this document was reconstructed from migration headers,
package documentation and the contract, not from version control, and that remains
the only source for it.
### 17.11 A minor documentation defect in the source — resolved
`internal/httpserver/api.go` — the doc comment describing `decodeBody` sat
immediately above `decodeInto`, so `decodeInto` carried two doc comments and
`decodeBody` carried none. The `decodeBody` comment has been moved to sit above
`decodeBody`.
### 17.12 Badge endpoint mismatch — verified
**Issue.** The `badges` table exists and is seeded with 4 records, but there is no
endpoint for it.
**Cause.** Recorded in contract §2: the frontend's `useBadges` hook has **zero
consumers**, and every badge the UI renders comes from
`worker_profiles.earned_badges`. The table exists; nothing reads it.
**Impact.** A frontend `Badge.list()` call would receive a 404. `README.md` states
that this call "has 404ed since Phase 2C".
**Current behaviour.** `badges` declares `Ops: 0`, so no route is registered, and its
policy is written out as an explicit empty policy so the resource is deliberately
closed rather than merely forgotten.
**Possible future handling.** Add the endpoint if a consumer appears, or drop the
table.
### 17.13 Job Posting 422 — mechanism verified, no specific defect recorded
**Issue as raised.** A "Job Posting 422 validation issue".
**What the repository shows.** No artifact — comment, test, TODO or contract note —
records a specific job-posting defect. *Not verified from current repository.*
**What can be stated factually.** `POST /api/v1/job-postings` answers 422 in exactly
the documented cases, and two of them are easy to hit from a client that was written
against the in-browser store:
1. **An unknown body field.** Contract §13.5 made unknown fields a 422 rather than a
silent drop, deliberately. Any field the frontend sends that has no column
produces `validation_failed` with the field named in `details`.
2. **An invalid enum value.** `job_postings` has three enum columns —
`english_required`, `status` and `priority` — and a value outside the declared set
is a 422 naming the permitted values.
`title` is the only column marked `Required` on this resource, so a missing-required
422 would name `title` specifically.
**Possible future handling.** If the symptom is reproducible, the `details` map in
the 422 body names the offending field, which is enough to decide whether the column
is absent from the schema (as `interview_id`, `training_outline` and
`score_breakdown` once were) or the value is wrong.
### 17.14 Shift snapshot / date behaviour — verified
Not a defect; a documented design constraint. Shift records are generated rather than
snapshotted because their dates are anchored to *now*, and the frontend windows every
collection on `created_date`. A frozen snapshot would read as permanently empty a
fortnight later. The consequence, recorded as U1 in the contract, is that attendance
and overtime will be **empty in any real deployment** until a rostering source exists.
Re-seeding on a later day prunes stale rows to keep the rolling window convergent.
### 17.15 `resetDemoData` — not present
The string `resetDemoData` does not appear anywhere in this repository. *Not verified
from current repository.* It is presumably a frontend concern.
### 17.16 Frontend test drift — not verifiable here
The frontend repository is not part of this checkout. *Not verified from current
repository.*
### 17.17 NATS appears in the repository despite the stated direction
**Issue.** The direction stated for this document is that NATS is not part of the
target architecture. `infrastructure/README.md` currently lists "Redis, NATS, MinIO"
as candidates for a later `docker-compose.dev.yml`, and `README.md` mentions NATS
among the things that do not exist yet.
**Impact.** A reader of `infrastructure/README.md` would take NATS to be planned.
**Resolution.** The rule stands, so the table was amended: `infrastructure/README.md`
no longer lists NATS as a candidate and says explicitly that it is not part of the
target architecture. `README.md` no longer names it either.
### 17.18 Deliberate contract behaviours that read as defects
These are recorded in `docs/api-contract.md` §12 with reasons and are not bugs:
multi-record writes are not transactional (§12.1); collection caps truncate silently
(§12.2); all aggregation is client-side (§12.3); search is browser-side (§12.4); email
comparison is case-insensitive server-side (§12.5); `DELETE` on a missing record
returns 200 (§12.7); LLM and file-upload integrations remain in the browser (§12.8).
---
## 18. Future Technical Direction
**Nothing in this section is implemented.** Each item is labelled with what the
repository actually says about it, so the two are never confused.
### Agent Runtime / Owliver
The intended direction is a Python "Owliver" service alongside the Go API, with a
subagent bridge and orchestration, and real execution behind the existing
`AgentExecutor` / `SkillExecutor` interfaces.
*What the repository says:* `README.md` states "a Python Owliver service to follow"
and "No Owliver". `internal/runtime/executor.go` names "real AI / Owliver / LangGraph
executors" as what would replace `UnavailableExecutor`.
`infrastructure/README.md` lists `Dockerfile.owliver` as later work. No Python source,
no bridge and no orchestration code exists here.
### AI Runtime
An LLM provider abstraction, with an appropriate provider per workload.
*What the repository says:* nothing. No provider name — OpenAI, Anthropic, Gemini,
Vertex — appears anywhere in this checkout. *Not verified from current repository.*
The attachment point is the two executor interfaces.
### RAG
`pgvector`, embeddings and retrieval over an agent's knowledge corpus.
*What the repository says:* `pgvector` is named exactly once, in
`infrastructure/README.md`, as part of a later dev compose file. It is **not
installed** in the local database (only `citext` and `plpgsql` are). Migration 000005
explicitly declines to create an `agent_knowledge` table because "no corpus exists".
Contract §11 lists RAG and embeddings as D7–D10, "none touch v1".
### Tools
FastMCP and a tool-execution layer.
*What the repository says:* nothing. Neither "MCP" nor "FastMCP" appears anywhere.
*Not verified from current repository.* Note that `internal/definition/yaml.go`
states as a design property that "there is no code path from a definition to
execution of any kind" — introducing tool execution changes that property
deliberately and should be done knowingly.
### Infrastructure
- **Google Pub/Sub** — no reference in the repository. *Not verified from current
repository.*
- **GCS-compatible object storage** — no reference. `infrastructure/README.md` names
MinIO as a later compose candidate.
- **Docker** — `infrastructure/` contains only a README, deliberately empty. It lists
`docker-compose.dev.yml`, `Dockerfile.api`, `Dockerfile.owliver` and
`otel-collector.yaml` as later work, with the stated reason that adding any of them
before its phase would be speculative.
- **Redis** — named in `ratelimit.go` as the seam for shared limiter state, and in
`infrastructure/README.md` as a later compose candidate. Nothing is wired.
- **Observability** — `otel-collector.yaml` is listed as later work. Today there is
structured JSON logging via `log/slog` and nothing else: no metrics, no tracing, no
exporter.
- **Production hardening** — no TLS termination, secret management, key rotation,
backup policy or deployment manifest exists in this repository.
### NATS
**NATS is not part of the target architecture.** `infrastructure/README.md` no longer
lists it as a candidate; see §17.17.
### Nearest-term work implied by the code itself
Not a roadmap, but the things the repository points at:
- Multi-record transactional endpoints (`POST /job-applications/{id}/hire`,
`POST /job-postings/{id}/assignments`), which the contract defers because they
require touching frontend hooks.
- Server-side aggregation, which contract §12.3 identifies as the first thing that
breaks as data grows — `shift_records` at 500 rows.
- Consuming `meta.truncated`, which exists so §12.2 is fixable without a contract
change.
- Defining the semantics of the `permissions:` block before storing or enforcing it.
- Shared state for the login limiter before deploying more than one instance.
---
## 19. What We Have Built So Far
For a team audience, in plain terms.
**Where we started.** A React demo whose data lived in a browser tab. `store.js` held
arrays, `seed.js` filled them at boot, and every hook called through one file:
`base44Client.js`. Nothing survived a refresh, everyone saw the same data, and no
rule existed that the client did not enforce on itself.
**Where we are.** A Go service in front of PostgreSQL that the same React app talks to
through the same one file.
| What we built | What it is practically worth |
| --- | --- |
| **A Go API over PostgreSQL** | Data survives. Two people see the same records. A query that takes 200 ms in the browser over 500 rows takes a millisecond in an index. |
| **A written contract derived from call sites** | Nobody had to guess what the frontend needed, and no endpoint exists that nothing calls. The frontend migration cost exactly one file. |
| **Migrations as the source of truth** | The schema has a history, a rollback, and a review surface. No table has ever been created by hand. |
| **Generated resource descriptors** | Column names, types and enums cannot drift from the database, because they are read out of it. |
| **A seeder that runs the frontend's own seed** | The demo dataset in PostgreSQL is the demo dataset the frontend ships, not a transcription of it. Re-seeding is safe and converges. |
| **Session authentication** | There is a real identity behind every request, held in an HttpOnly cookie the page cannot read, backed by a row we can revoke. |
| **Role-based authorization with deny-by-default** | Who may do what is written down in one table, and a resource nobody has written a rule for is unreachable rather than open. |
| **Tenant and ownership isolation in SQL** | A row from another organization, or another worker, is never fetched — so it cannot leak through a count, a total, or a bug in a later loop. |
| **Agent and Skill authoring in the database** | Authored definitions moved out of a jsonb blob in user preferences into real tables with an owner, a tenant, a size bound and a query surface. |
| **A Go parser that provably matches the JS one** | An author sees the same acceptance, the same rejection and the same wording in the editor and from the API — asserted against the real frontend parser's captured output over 37 shipped definitions and 132 adversarial cases. |
| **A runtime boundary** | Loading, tenant scoping, status eligibility and dependency resolution are written and tested, so when execution arrives it plugs into an interface rather than reopening those questions. |
| **A test suite that runs against a real database** | 175 test functions, 920 test entries, a disposable migrated-and-seeded database per test process, and ordering guarantees verified to fail when removed. |
**What is deliberately not built yet.** Real AI execution. Any public runtime
endpoint. Server-side aggregation. Transactional multi-record flows. RAG. Tooling.
Anything in `infrastructure/`.
---
## 20. Current Backend Snapshot
**Current Database:** PostgreSQL 18.6 locally; one schema (`public`); five migration
pairs, applied version 5, not dirty; 20 application tables plus `schema_migrations`;
16 enum types; extensions `citext` and `plpgsql`; UUID primary keys with a nullable
unique `legacy_id`; `timestamptz` throughout.
**Current API:** Go 1.27, `net/http` only, base path `/api/v1`, 51 registered routes
(34 resource + 4 `/me` + 2 auth + 10 definitions + `/health`). JSON envelope with
`data` and, on collections, `meta`. Two filter operators, `NULLS LAST` in both
directions with an `id` tiebreaker, `limit`/`offset` with per-endpoint defaults and a
1000 cap, and nine error codes.
**Current Authentication:** email + password with argon2id at OWASP parameters;
opaque 256-bit session tokens stored only as SHA-256; HttpOnly, `SameSite=Lax`
cookies with `Secure` outside development; sliding expiry with an absolute ceiling
(12h/24h, or 30d/90d with remember-me); a 15-minute expired-session sweeper; failed
attempts limited per email and per address; uniform 401s with decoy hashing to resist
enumeration. Passwords enter only through `cmd/setpassword`.
**Current RBAC:** three roles from `users.role` — `admin`, `employer`, `talent`.
Deny-by-default policy table keyed by URL path. Role gate in the handler answering
403 before any query. Eight server-owned identity columns. `PATCH /me` writes only
`full_name` and `account_type`.
**Current Tenant Model:** one organization per user, taken from the user's row and
never from the request. `org_id` predicate on every read and write. `courses` and
`learning_paths` also match `org_id IS NULL`, the shared platform library. Talent
callers carry a second ownership predicate in the same `WHERE` clause; a row outside
either predicate answers 404.
**Current Agent System:** Markdown with YAML frontmatter, stored verbatim in
`agent_definitions`; six projected columns; `draft`/`published`/`archived` plus a
monotonic integer version; personal and organization tiers with per-tier uniqueness;
full CRUD over five routes.
**Current Skill System:** the same shape in `skill_definitions`, minus version;
`active`/`inactive`; full CRUD over five routes. `ui:` and `owliver:` frontmatter
blocks are recorded as deferred rather than validated.
**Current Runtime:** an internal boundary. Loads by UUID or `definition_id` within
the caller's tenant and ownership scope, re-parses and re-validates the stored
Markdown, applies status eligibility, resolves an agent's skill dependencies with
personal-over-organization shadowing and deduplication, and dispatches to an executor
interface. The only executor implementation refuses execution. No HTTP route reaches
it.
**Current Frontend Integration:** the frontend repository is separate and unmodified.
The seam is `base44Client.js`; the contract's acceptance criterion is that swapping
the transport inside it changes nothing above it. CORS allowlist defaults to the Vite
dev server on both hostnames in development, and to empty elsewhere.
**Current Testing:** 175 test functions, 920 test entries executed, no failures and
none skipped on a machine with PostgreSQL reachable. `gofmt` reports no files,
`go vet` reports no diagnostics, `go build ./...` builds. Database-backed tests use a
disposable `krow_backend_autotest_<pid>` database per test process.
**Current Known Limitations:** no AI execution and no public runtime endpoint; login
rate limiting is per-process and in-memory; CORS does not permit credentials;
definitions endpoints bypass the policy table; the runtime package is unreachable
from the running service; the definitions endpoints are absent from the API contract;
attendance data is empty without a rostering source; the repository has a single
commit and so no usable history.
**Current Development Focus:** the most recent work in the tree is the definition
system — schema, parser conformance, CRUD — and the runtime boundary that sits on top
of it. The open seams the code itself points at are real executors behind the
executor interfaces, and a public surface for the runtime if one is wanted.
---
## 21. Critical Engineering Rules
These are the invariants the current code depends on. Breaking any of them silently
is how this codebase would stop being trustworthy.
1. **Do not casually modify migrations.** Once a migration has run anywhere other
than your own machine it is immutable. Change it and every database that already
applied it diverges from every one that has not. Write the next one instead —
that is exactly what 000002 and 000003 are.
2. **Do not bypass tenant scoping.** `org_id` belongs in the `WHERE` clause of every
read and write. It is not a filter applied afterwards, and it is not optional on
personal rows.
3. **Do not trust a client-provided `org_id`.** It comes from the user's row, which
comes from the session row, which comes from a cookie value the client cannot
forge without already holding it.
4. **Do not trust a client-provided `owner_user_id`,** or any of the eight
server-owned identity columns. They are the columns every ownership rule rests on;
a caller who could set them could defeat the rule with the same request it
constrains.
5. **Do not bypass parser validation.** `ValidateAgent` / `ValidateSkill` before
`ParseAgent` / `ParseSkill` before persistence, on create and on any update that
touches `markdown`.
6. **Do not mutate stored Markdown.** The bytes in are the bytes stored. Normalization
reads; it never rewrites. Three tests assert this from three directions.
7. **Do not invent Agent or Skill fields.** Every field exists because the frontend
parser has it. A field the Go parser knows and the JS parser does not is a
definition the editor cannot read.
8. **Do not add skill versioning** without an explicit architecture decision. The
frontend has no notion of it and cannot set one.
9. **Do not bypass RBAC.** The role gate runs in the handler before any query, and
the row predicate runs in SQL. Both, every time.
10. **Do not expose cross-tenant resource existence.** 403 means "your role"; 404
means "not yours or not there". Never a message that distinguishes the two.
11. **Do not introduce NATS.**
12. **Do not introduce RAG or `pgvector`** before the architecture calls for it. There
is no corpus, and migration 000005 declines to create a table for one.
13. **Do not introduce MCP or tool execution prematurely.** The parser's stated
property is that no code path leads from a definition to execution of any kind;
changing that should be a decision, not a side effect.
14. **Do not introduce Redis prematurely.** The limiter names it as the seam; wire it
when there is more than one instance, not before.
15. **Do not modify the frontend during backend-only work** unless it is explicitly
required. The whole transport migration is affordable because exactly one
frontend file changes.
16. **Do not fake AI execution.** `UnavailableExecutor` refuses honestly. A stub that
returns plausible text would be worse than one that returns an error.
17. **Do not silently change API contracts.** `docs/api-contract.md` is the client's
contract. Divergences from it have been reported and written down every time —
§13.4 and §13.5 exist for that reason.
18. **Do not let the descriptors and the policy table drift.** `resources_gen.go` is
generated; `policy.go` is hand-written. That separation is why regenerating one
cannot silently drop a rule in the other, and `TestEveryResourceHasAPolicy` is
what keeps it honest.
19. **Do not put reconnaissance in a public response.** `/health` answers the verdict;
the reasoning goes to the log.
20. **Do not log a password, a session token, or a password hash.** Nothing in
`internal/auth` logs at all, by design.
---
## 22. Final Team Summary
### What the team should know
- **We started with a browser-only demo.** Data lived in arrays in a tab; there was
no identity, no tenancy, and no rule the client did not enforce on itself.
- **Go and PostgreSQL were introduced to move persistence, identity and enforcement
to a server** without rewriting the frontend. The acceptance criterion was written
down first: swapping the transport inside one frontend file must change nothing
above it. That held.
- **The architecture is a plain layered service** — handler, service, repository, pgx,
PostgreSQL — with no ORM and no web framework. Three direct Go dependencies.
- **The schema lives in `migrations/` and nowhere else.** Five migration pairs;
nothing in the Go code issues DDL. Resource descriptors are *generated* from the
live schema so they cannot drift.
- **The API contract was derived from frontend call sites, not designed.** An
operation with no call site gets no endpoint — which is why `badges` has none and
`DELETE /job-postings/{id}` answers 405.
- **Authentication is server-side sessions:** argon2id passwords, opaque 256-bit
tokens, only SHA-256 stored, HttpOnly cookies, sliding expiry with an absolute
ceiling, and login rate limiting per email and per address.
- **Authorization is deny-by-default.** Three roles from `users.role`. The role gate
answers 403 before any query; tenancy and ownership are SQL predicates, so a row
outside them answers 404 and cannot be distinguished from one that does not exist.
- **Agent and Skill definitions moved out of a jsonb preferences blob into real
tables** with an owner, a tenant, a size bound and a query surface. The Markdown is
the authoritative artefact; every column is derived from it and recomputed on write.
- **The Go parser provably matches the JavaScript one** — replayed against the real
frontend parser's captured output over 37 shipped definitions and 132 adversarial
cases, with mutation checks so the tests cannot pass against a broken parser.
- **The runtime boundary exists; execution does not.** Loading, tenant scoping,
status eligibility and dependency resolution with personal-over-organization
shadowing are written and tested. The only executor refuses. No HTTP route reaches
the runtime at all.
- **Verification is real:** 175 test functions, 920 test entries, against a disposable
migrated-and-seeded PostgreSQL database per test process. `gofmt`, `go vet` and
`go build` are clean.
- **The main current limitations** are: no AI execution and no runtime endpoint;
login rate limiting that does not survive scale-out; CORS that does not yet permit
credentials; ten definition endpoints that sit outside the policy table and outside
the written contract; and a repository with a single commit and so no usable
history.
- **The direction** is real execution behind the existing executor interfaces
(Owliver / an LLM provider abstraction), then retrieval, then tooling, then
deployment infrastructure — none of which exists here today. **NATS is not part of
the target architecture**, and the documentation no longer implies otherwise.
- **The rules that must not be broken** are in §21. The two that matter most in daily
work: never trust a client-supplied identity or tenant, and never rewrite stored
Markdown.
---
*Generated from the repository state on 2026-08-24. Every count, version and
behaviour above was read from the current checkout or from read-only queries against
the local development database. Nothing in the repository was modified to produce
this document.*
*Revised on 2026-08-24 by a structure-and-dead-code cleanup pass, which changed no
schema, no migration, no database row and no runtime behaviour. What it did change is
recorded in §17.1, §17.10, §17.11 and §17.17: documentation and source comments that
had fallen behind the code were corrected. The counts above were re-verified against
the checkout and are unchanged.*