Files
krow_backend/docs/KROW_BACKEND_COMPLETE_SUMMARY.md
2026-08-25 16:37:05 +05:30

133 KiB
Raw Permalink Blame History

Krow Backend — Complete Technical Summary

Audience: backend, frontend, AI/agent and DevOps engineers, and engineering leads. Scope: what exists in the krow-backend repository today, how it got here, and where it is heading.

How this document was produced. Every factual claim below was read out of the repository at /Users/apple/Krow/krow-backend as it currently stands — Go source, go.mod, migrations, Makefile, README.md, docs/api-contract.md, tests and fixtures — plus read-only queries against the local development database. Where an earlier report or an in-repository comment disagrees with the code, the code is treated as authoritative and the disagreement is recorded in §17.

Anything that could not be checked against this repository is marked "Not verified from current repository." Nothing here is inferred from a system that is not present in this checkout.

This is a technical history and current-state document. It does not certify phases.


1. Executive Summary

What the Krow backend is

Krow is a hiring and workforce platform: job postings, applications, AI screening interviews, staff records, worker profiles, assignments, attendance, a training library, evidence of work, and an audit log. On top of that sits an authoring system for Agents and Skills — Markdown documents with YAML frontmatter that describe what the product's assistant can be offered on each page.

krow-backend is the Go HTTP API and PostgreSQL database behind that product. The Krow frontend lives in a separate repository (krow-demo) and is not vendored, copied or modified here.

Why the backend was introduced

The frontend shipped first, as a demo. Its data layer was a client-side store: src/api/store.js held arrays in memory, src/api/seed.js filled them at boot, and src/api/base44Client.js was the seam every hook called through. That arrangement has three properties that stop being acceptable the moment more than one person uses the product:

  • Data does not survive. It lives in the browser tab.
  • There is no tenancy and no identity. Every viewer sees the same array.
  • Nothing is enforceable. Any rule the client does not apply does not exist.

The backend exists to move persistence, identity, tenancy and enforcement to a server, without rewriting the frontend above the transport seam. That constraint is written into docs/api-contract.md as its acceptance criterion:

Replacing the transport inside src/api/base44Client.js — and changing no other frontend file — must leave the application behaving identically.

Every design decision downstream follows from that sentence.

The transformation

OLD                                CURRENT

React                              React
  |                                  |
base44Client                       base44Client
  |                                  |
localStorage / mock data           HTTP
                                     |
                                   Go API
                                     |
                                   PostgreSQL

The shape above the seam is unchanged. What moved is everything below it: the records now live in PostgreSQL, the organization a request runs as comes from a server-side session rather than from nowhere, and the rules about who may read or write what are applied in SQL and in the handler rather than in the browser.

The Agent / Skill path

Agents and Skills are authored as Markdown. The backend parses them, validates them against the same rules the frontend editor applies, projects a few queryable columns out of them, and stores the Markdown itself verbatim:

Markdown
   |
Parser              (internal/definition — YAML subset + frontmatter, ported from JS)
   |
Validator           (ValidateAgent / ValidateSkill — frontend messages, plus DB bounds)
   |
PostgreSQL          (agent_definitions / skill_definitions; markdown stored byte-for-byte)
   |
CRUD API            (/api/v1/agent-definitions, /api/v1/skill-definitions)
   |
Runtime Loader      (internal/runtime — loads by id or definition_id, re-parses)
   |
Dependency Resolution  (agent's `skills:` list resolved within tenant + ownership scope)
   |
Execution Boundary  (AgentExecutor / SkillExecutor interfaces)

What the runtime can and cannot do today

It can: load an authored agent or skill from the database by UUID or by definition_id; re-parse and re-validate its Markdown; refuse it if its status makes it ineligible (a draft or archived agent, an inactive skill); resolve an agent's declared skill dependencies within the caller's organization and ownership scope, preferring a personal definition over an organization one with the same definition_id; deduplicate the resolved list; and hand the resolved agent to an executor through a typed interface.

It cannot: execute anything. The only executor implementation in the repository is UnavailableExecutor, which returns ErrExecutorUnavailable for both agents and skills. There is no LLM client, no prompt assembly, no tool dispatch and no external AI call anywhere in go-api/.

It is also not reachable over HTTP. No route in internal/httpserver references the runtime package; a repository-wide search for runtime. outside the package itself and its tests returns nothing. The runtime is an internal boundary exercised only by its own test suite.

Current architectural direction

The backend is a layered, no-ORM Go service where the shape of every resource is generated from the live database schema and the rules about every resource are hand-written next to it. Authorization is deny-by-default. Tenancy and ownership are SQL predicates rather than post-fetch filters. The definition system treats the authored Markdown as the authoritative artefact and every column derived from it as a cache. The runtime establishes where AI execution will attach, without attaching it.


2. Development History

The repository documents its own work in phases, in migration headers, package documentation and docs/api-contract.md. What follows describes what each phase introduced and what of it is in the code today.

Phase 1 — Backend Foundation

What was built. The Go module, the PostgreSQL connection, the migration discipline and the initial schema.

  • Go module — github.com/krow/krow-backend/go-api, go 1.27. Three direct dependencies and nothing else: github.com/jackc/pgx/v5, golang.org/x/crypto, golang.org/x/term. No web framework, no ORM, no YAML library, no test framework.
  • Database access — internal/db owns a pgxpool.Pool. db.Open deliberately acquires and pings once at startup, because pgxpool.New is lazy and would otherwise defer a misconfiguration to the first request.
  • Configuration — internal/config loads from the environment with typed defaults and a validate() pass that refuses: an unknown APP_ENV; a port out of range; a DATABASE_SCHEMA that is a PostgreSQL system schema or starts with pg_; MIN_IDLE_CONNS above MAX_OPEN_CONNS; DATABASE_SSLMODE=disable when APP_ENV=production; and a CORS entry of * or one without a scheme.
  • Health — GET /health, and db.Check behind it.
  • Migrations — migrations/ is the source of truth for the schema. Nothing in go-api/ issues DDL, and no ORM generates any. The Makefile wraps the golang-migrate CLI (README records 4.19.1).

Implementation decisions worth carrying forward.

  • uuid primary keys plus a nullable unique legacy_id text, so the frontend's existing string ids (jobposting_m1a2b3c001) survive as data without becoming the addressable id.
  • timestamptz throughout; created_date / updated_date keep the frontend's own names, because the frontend windows every collection on created_date and its default sort string is -created_date.
  • Native enums for closed vocabularies, text + CHECK where the set is still moving.
  • Email columns are citext, and stay populated even where a UUID foreign key also exists, because the frontend joins workers by email today.
  • Guards against destructive change: there is deliberately no make migrate-drop target; make migrate-down refuses unless APP_ENV=development; down migrations name every object they drop, with no DROP SCHEMA and no CASCADE on a schema.

Phase 2A — Architecture Decisions

Four questions were raised, investigated against the frontend, and recorded in docs/api-contract.md §11 as deferred with reasons. All four are still recorded there as deferred.

Ref Question Recorded outcome
D2 Is company a real entity? It is free text on job_postings, displayed and read through a fallback chain, but never grouped by id, joined, or given a route. Kept as company text NOT NULL DEFAULT ''. No clients table, no endpoint.
D3 Are assigned / rejected pipeline stages? The frontend's STAGE_ORDER excludes both from every funnel count via indexOf(...) >= from. The enum carries all seven values; the API stores and returns status verbatim and never filters, reinterprets or normalises it. The funnel exclusion is recorded as a live frontend behaviour that Phase 2 does not touch.
D6 Do the unmounted pages come back? Fourteen page files plus a layout are imported by App.jsx and mounted on no route. Three operations are reachable only through them. All three endpoints were included anyway and marked unreachable-today, because the calling code exists and excluding them would make the transport shim need conditional methods.
U1 Where do shift records come from? ShiftRecord has exactly one frontend seam operation: .list('-created_date', 500). No create, no update, no delete. The contract specifies GET /api/v1/shift-records only — no POST — because there is no call site to derive a write contract from.

Organization / tenant modelling. organizations and user_preferences are the two tables in migration 000001 that do not correspond to an entity the frontend reads through base44Client.js; both are documented in the migration itself. Every tenant-scoped table carries org_id. Two tables — courses and learning_paths — are organization-nullable: a NULL org_id means the shared platform library, visible to every tenant.

Phase 2B — API Contract

docs/api-contract.md (1,060 lines) was written before implementation, derived entirely from the frontend repository and the Phase 1 schema. Its stated rule: every endpoint exists because a call site exists, and every semantic was read out of src/api/store.js rather than designed.

  • Operation matrix. An operation with no call site gets no endpoint. DELETE /job-postings/{id} and POST /shift-records are absent for that reason, and the router answers 405 for them because the path pattern exists under another method.
  • REST shape. Base path /api/v1; kebab-case plural resource names, with mass nouns staying singular (/staff, /evidence, /user-activity); PATCH with shallow merge and no PUT.
  • Field naming. snake_case, identical to the frontend's own field names — no renaming, no camelCase conversion. The single documented exception is /me/preferences.
  • Envelope. Responses are wrapped in {"data": …}, with meta on collections. The frontend's methods return bare values; the shim unwraps with one .data. What the envelope buys is meta.total, which the bare shape has nowhere to put.
  • Filtering (§6). Two operators only: equality for a single value, membership for a repeated parameter. Every non-reserved query parameter is a field filter.
  • Sorting (§7). NULLS LAST in both directions, because the frontend's comparator returns before it can place a null; and , id appended to every ORDER BY, because JavaScript's sort is not stable in the way the UI relies on.
  • Pagination (§8). limit / offset, with per-endpoint defaults taken from the literal arguments at each frontend call site.
  • Errors (§5). A single envelope shape with a code, a message and a details map.

Phase 2C — Backend Implementation

Layers. internal/domain (descriptors and errors) → internal/service (parsing, validation, contract semantics) → internal/repo (all SQL) → pgx.

Generated descriptors. internal/domain/resources_gen.go is produced by scripts/gen_resources.py (make gen-resources), which reads column names, types, enum values and nullability out of information_schema. The hand-maintained part is the per-resource metadata — path, default sort, default limit, supported operations, required fields — which comes from the contract. Column definitions therefore cannot drift from the migrations.

One query builder, not fourteen repositories. internal/repo/repo.go builds every statement from a *domain.Resource: column lists are explicit, every value is a bind parameter cast to its declared type, and no identifier ever comes from user input — a filter or sort name is resolved to a real *domain.Column before any SQL is assembled. The contract's awkward semantics are then implemented once and apply identically everywhere.

Seed system. seed/fixtures/seed.json is generated by executing the frontend's src/api/seed.js through Vite and serialising what it exports, so ids, dates, numbers and enum values arrive exactly as the demo has them with no transcription step. internal/seeder loads it inside a single transaction with explicit upserts. Shift records are the exception and are generated by a Go port of attendanceSeed.js — see §6.

Bugs and gaps this phase exposed, all recorded in docs/api-contract.md §13:

  • Three columns had no home (§13.1). job_applications.interview_id, courses.training_outline and worker_profiles.score_breakdown are written or read by the frontend and were absent from migration 000001. Migration 000001 was not edited; the columns were added in 000002. interview_id is deliberately not a foreign key, because a seeded application references an interview id for which no record exists, and nulling it would have transformed source data.
  • A constraint that rejected real data (§13.2). job_applications_screened_consistent from 000001 rejected 9 of 24 seeded applications — precisely the 9 AI-scored ones. screened_at is never read or written by the frontend; the constraint came from a blueprint rather than from evidence. Dropped in 000003; the column is retained, nullable and unused. All 53 CHECK constraints were then evaluated against all 245 seeded records; the other 52 hold.
  • Unknown fields are rejected, not ignored (§13.5). A body field that maps to no column answers 422 with the field named in details. The reasoning is explicit: silently ignoring them is exactly how the three columns above would have been lost. Server-owned fields (id, org_id, created_date, updated_date, legacy_id) remain ignored rather than rejected.
  • Three deliberate divergences from store.js (§13.4): email matching is case-insensitive because the columns are citext; updated_date is always populated because the column is NOT NULL and the seeder sets it to created_date; and _order, a positional index the frontend used while building its course list, is dropped.

Phase 2D — Frontend Transport Integration

This phase belongs to the frontend repository, which is not present in this checkout. What is verifiable here is the backend side of the seam:

  • CORS (internal/httpserver/cors.go). Origins are matched exactly against an allowlist and echoed back one at a time; * is rejected at config validation. A request whose Origin is not on the list is served normally with no CORS headers — the API does not refuse it, the browser simply will not hand the response to the page. That distinction is deliberate so that curl, health checkers and server-to-server callers, which send no Origin, are unaffected. With an empty allowlist the middleware is not installed at all.
  • Development defaults. With APP_ENV=development and HTTP_CORS_ORIGINS unset, the allowlist defaults to http://localhost:5173, http://127.0.0.1:5173 and the vite preview port — both hostnames, because a browser treats them as different origins. Unset anywhere else it defaults to empty, meaning same-origin only.
  • JSON everywhere. net/http's own plain-text 404 and 405 replies are intercepted and rewritten into the error envelope, so every response from the API is JSON.

The frontend files named in the contract — base44Client.js, krowHooks.js, store.js — appear only as references inside this repository's documentation and comments. httpClient.js and any VITE_API_* environment variable are not referenced anywhere in this repository: Not verified from current repository.

Known frontend/backend mismatches recorded in the contract §12: multi-record writes are not transactional; collection caps truncate silently; all aggregation is client-side; search is browser-side substring matching; DELETE on a missing record returns 200. See §17.

Phase 3 — Authentication & Authorization

Split across several steps in the repository's own narrative; what exists today is described in §8. Chronologically:

3B — schema and primitives. Migration 000004 adds two facts and nothing else: a globally unique index on users.email (because the login form supplies no organization, so an email must resolve to exactly one user), and a sessions table. The migration explicitly declines to add roles, permissions, organization_members, credentials or refresh_tokens tables. internal/auth was written in the same step and deliberately knows nothing about HTTP: password hashing (password.go), token generation and hashing (token.go), the session lifecycle (session.go), the PostgreSQL stores (store.go, users.go) and credential verification (credentials.go).

3C — the HTTP surface. POST /api/v1/auth/login, POST /api/v1/auth/logout, and the authenticate middleware. This step removed devOrgMiddleware, which had put a fixed organization on every request with no credential behind it. Because the organization had always been threaded explicitly through the service and repository boundaries rather than defaulted inside the SQL, that swap changed one line and nothing below it.

3D — authorization. internal/domain/policy.go — a hand-written table, kept separate from the generated descriptors precisely so that regenerating descriptors can never silently drop an access rule, and a new column can never grant anyone anything by accident. The handler gate is Server.authorize in internal/httpserver/api.go; the row predicates are in builder.ownership in internal/repo/repo.go. docs/api-contract.md §9A documents the resulting contract.

Health information exposure reduction. GET /health is public, so its body was reduced to a single field: {"status": "ok" | "degraded" | "unavailable"}. The PostgreSQL version, database name, schema name, applied migration version, table count, connection error text, environment and uptime were removed from the response and moved to the server log — debug when healthy, warn otherwise. db.Check itself is unchanged and still gathers all of it.

Future authentication direction is discussed in §18 and is separated there from what the code does.

Phase 4A — Agent & Skill Architecture

Agents and Skills are configuration written as Markdown with YAML frontmatter. Three tiers exist, and the reasoning for keeping them apart is recorded in migration 000005:

Tier Where it lives Rows in this database
shipped src/agents/**/*.md, src/skills/**/*.md in the frontend repo, versioned in Git, bundled at build time none — they are code, and putting them in a table would trade git log, code review and atomic deploy for nothing
organization authored in the app, shared across one tenant agent_definitions / skill_definitions, visibility = 'organization'
personal authored in the app, private to one user agent_definitions / skill_definitions, visibility = 'personal'

Before this work, the last two lived in user_preferences.extra — a jsonb blob with no owner, no tenancy, no size bound, no server-side validation and no query surface, returned in full by GET /api/v1/me on every page load.

Phase 4B — Definition Contract

A definition is a Markdown document whose leading fenced block is YAML frontmatter and whose remainder is the body. The parser contract is stated in internal/definition/definition.go:

  • The document layer — byte-order mark, line endings, leading blank lines, fence recognition, body extraction (frontmatter.go).
  • The YAML subset — block maps and sequences, scalars, quoting, comments (yaml.go).
  • Definition-level normalization and validation — id, name, description, status, version, pages, icons, reasoning, permissions, starters, knowledge, subagents.

Those cover every column migration 000005 projects out of a definition: definition_id, status, version, name, description, pages.

Field-by-field detail for both kinds is in §9.

Phase 4C — Definition Database

Migration 000005 creates agent_definitions and skill_definitions — two tables, deliberately, not one. The reasoning is recorded in the migration: agents have draft | published | archived plus a monotonic integer version; skills have active | inactive and no version at all, because the frontend has no notion of a skill version. One table would need a union CHECK permitting "version 5, status inactive", and a version column forever equal to 1 for half the rows.

The migration also records what it deliberately does not create: definition_versions (nothing retains prior Markdown and no rollback feature exists to serve), definition_permissions (the permissions: block stays inside the Markdown, parsed and unenforced), agent_skills / agent_subagents (the skills: list names ids in a namespace that includes shipped definitions, which have no rows here, so a join table would need foreign keys to rows that do not exist), agent_knowledge, and conversations.

Schema detail is in §5.

Phase 4D — Parser Compatibility

Two parsers read the same definition: the JavaScript in the frontend's src/lib/skills and src/lib/agents, and the Go in internal/definition. If they disagree, one of two silent failures occurs: a definition the editor accepts and the API rejects looks valid while it is being written and fails when it is saved; or a definition the API accepts and the editor rejects is stored and then cannot be rendered by the product that owns it.

The Go side is therefore a port, not an independent implementation, and compatibility is enforced by replay rather than by description. Detail is in §10.

Phase 4E — Agent & Skill CRUD

Ten HTTP routes — list, create, get, update, delete for each of agents and skills — plus internal/service/definitions.go and internal/repo/definitions.go. Detail is in §11.

Phase 4F — Runtime Boundary

internal/runtime introduces the runtime representations (runtime.Agent, runtime.Skill), the execution value types (ExecutionInput, ExecutionResult, RuntimeError), a Loader that reads a stored definition and re-parses it, status eligibility rules, dependency resolution with personal-over-organization shadowing, the AgentExecutor / SkillExecutor interfaces, an Engine that ties them together, and UnavailableExecutor.

The current runtime establishes the execution boundary. UnavailableExecutor returns ErrExecutorUnavailable from both ExecuteAgent and ExecuteSkill, and performs no AI execution of any kind. Detail is in §12.


3. Current Architecture

Request path

React (krow-demo — separate repository)
   |
base44Client.js                     the one frontend file allowed to change
   |
HTTP  (JSON, /api/v1, cookie-bearing)
   |
=====================  krow-backend  =====================
   |
Go HTTP server            net/http.Server + ServeMux
   |
requestLogger             method, path, status, duration
   |
cors                      installed only when an allowlist is configured
   |
recoverer                 a panic becomes a logged 500, not a dropped connection
   |
authenticate              cookie -> session row -> user row -> Identity + OrgID
   |
jsonErrors                rewrites the mux's plain-text 404/405 into the envelope
   |
HTTP handlers             role gate (403), body decode, path values
   |
Service layer             query parsing, validation, contract semantics
   |
Repository layer          all SQL; org + ownership predicates; bind parameters only
   |
pgx / pgxpool
   |
PostgreSQL

Definition and runtime path

Agent/Skill Markdown            authored in the app, submitted as a JSON string
   |
Parser                          internal/definition — frontmatter + YAML subset
   |
Validator                       ValidateAgent / ValidateSkill
   |
Definition Projection           id, status, version, name, description, pages
   |
PostgreSQL                      agent_definitions / skill_definitions
                                markdown stored verbatim; columns derived from it
   |
Definition CRUD                 /api/v1/agent-definitions, /api/v1/skill-definitions
   |
Runtime Loader                  internal/runtime — load by uuid or definition_id,
                                re-parse, re-validate
   |
Dependency Resolution           agent.Skills -> LoadSkill each, tenant + ownership
                                scoped, personal shadows organization, deduplicated
   |
Executor Boundary               AgentExecutor / SkillExecutor interfaces
                                (only UnavailableExecutor exists)

Layer responsibilities

Layer Package Owns Deliberately does not own
Entrypoint cmd/api config load, pool, logger, signal handling, graceful shutdown, session sweeper goroutine routing, business rules
Config internal/config environment loading, typed defaults, validation secrets (nothing is baked in)
Database internal/db pool lifecycle, startup ping, Check for /health any DDL
HTTP internal/httpserver routing, middleware chain, role gate, response envelopes, cookies, rate limiting SQL
Auth primitives internal/auth argon2id, token generation/hashing, session lifecycle, credential verification HTTP, cookies, logging (nothing in the package logs)
Identity context internal/authctx the authenticated Identity on the request context any setter that takes a user id from a client
Org context internal/orgctx the organization a request runs as deciding which organization that is
Service internal/service query-string parsing, body validation, contract semantics statement construction
Domain internal/domain resource descriptors (generated), the authorization policy table (hand-written), error constructors, record types I/O
Repository internal/repo every SQL statement, organization scope, ownership predicates, derived columns deciding whether a role may act
Definitions internal/definition frontmatter, the YAML subset, agent/skill normalization and validation storage, HTTP, and the ui:/owliver: blocks
Runtime internal/runtime loading, eligibility, dependency resolution, the executor interfaces executing anything
Seeder internal/seeder fixture loading, deterministic ids, upsert, shift generation and pruning schema changes
Test harness internal/testutil a disposable migrated + seeded database per test process touching the development database

Where authorization happens, and why it is split. The role check runs in the handler before any query, so it always answers 403 and can never reveal whether a row exists. Organization scope and ownership are SQL predicates in the WHERE clause, so a row outside them is simply absent and answers 404. A caller cannot distinguish "exists and is not yours" from "does not exist". Because the predicate is in SQL, count(*) runs over the same clause — so meta.total for a talent caller is their own count, not the organization's.


4. Technology Stack

Current implementation — verified from this repository

Concern What is used Evidence
Language Go, module directive go 1.27; local toolchain go1.27.0 darwin/arm64 go-api/go.mod, go version
PostgreSQL driver github.com/jackc/pgx/v5 v5.10.0 (with pgxpool) go-api/go.mod
Password hashing golang.org/x/crypto v0.42.0 (argon2) go-api/go.mod, internal/auth/password.go
Terminal input golang.org/x/term v0.35.0 (hidden password prompt) go-api/go.mod, cmd/setpassword
Indirect deps pgpassfile v1.0.0, pgservicefile, puddle/v2 v2.2.2, x/sync v0.17.0, x/sys v0.37.0, x/text v0.29.0 go-api/go.mod
HTTP server Go standard library only — net/http.Server and http.ServeMux with method+pattern routes ("GET /api/v1/job-postings", "{id}" path values). No third-party router or framework. internal/httpserver/server.go, api.go
Logging log/slog, JSON handler to stdout, level from LOG_LEVEL cmd/api/main.go
Migrations golang-migrate CLI, wrapped by the Makefile; -seq numbering Makefile, README.md (records version 4.19.1)
PostgreSQL 18.6 (Homebrew) on the local development machine read-only SHOW server_version
Applied migration version 5, not dirty, in the local development database read-only SELECT version, dirty FROM schema_migrations
Extensions installed citext, plpgsql. pgvector is not installed. read-only SELECT extname FROM pg_extension
Testing Go's own testing package. No assertion library, no mocking framework, no test containers. Database-backed tests build a disposable database per test process. internal/testutil/db.go
Parser Hand-written, no YAML dependency — a port of the frontend's yaml.js. Conformance is asserted against a captured JS oracle fixture. internal/definition/yaml.go, conformance_test.go
Runtime Go, in-process, no external calls. Executor is an interface with one stub implementation. internal/runtime
Code generation scripts/gen_resources.py (Python 3, shells out to psql) Makefile target gen-resources
Oracle capture scripts/oracle.mjs + scripts/cases.mjs (Node, runs the frontend module graph through Vite) scripts/

Frontend transport relationship. The frontend repository is not part of this checkout. This backend serves the contract in docs/api-contract.md; the frontend consumes it through base44Client.js. No frontend source is vendored, copied or modified here.

Future / planned direction — not implemented

Nothing in this list is present in the repository. See §18 for detail and §17 for what the repository does and does not say about each.

  • Owliver (the planned Python service), LangGraph, or any orchestration layer
  • Any LLM provider client
  • RAG, embeddings, pgvector
  • FastMCP or any tool-execution layer
  • Redis, or any shared cache/rate-limit store
  • Docker images or compose files
  • Object storage
  • OpenTelemetry or any tracing/metrics exporter

5. PostgreSQL & Database Architecture

Database and schema

  • Local development database name and credentials come from environment variables; .env is gitignored and .env.example contains only placeholders and safe local defaults. No password, key or token appears in any tracked file.
  • The application owns exactly one schema, defaulting to public, and config.validate() refuses to point it at a PostgreSQL system schema.
  • migrations/ holds five up/down pairs. The local development database reports applied version 5, not dirty.

Migrations

File Introduces
000001_initial_schema 17 tables, 16 enum types, the citext extension, and the index and constraint set
000002_application_interview_id job_applications.interview_id, courses.training_outline, worker_profiles.score_breakdown
000003_drop_screened_consistent_check drops job_applications_screened_consistent; keeps the unused screened_at column
000004_auth_sessions global unique index on users.email; the sessions table
000005_agent_skill_definitions agent_definitions, skill_definitions

Every migration has a matching down file, and a test asserts that (TestEveryMigrationHasADownFile). Two further tests assert that 000004 and 000005 are reversible, and one asserts that 000005 adds exactly two tables and no others.

Migration rules recorded in README.md: files are the source of truth; an applied migration is immutable (write the next one instead); staging and production run the same files as a discrete deploy step before the new binary rolls out, which means every migration must be backwards-compatible with the currently-running version — expand, migrate, contract, never in one migration.

Tables present in the local public schema

Twenty application tables plus schema_migrations (21 base tables, verified read-only):

organizations       users               user_preferences    sessions
job_postings        job_applications    ai_interviews       staff
worker_profiles     assignments         shift_records       evidence
user_activity       courses             learning_paths      role_categories
certifications      badges              agent_definitions   skill_definitions

Important tables in plain terms

Table What it holds Notable
organizations the tenant every tenant-scoped table references it
users people who can sign in role (admin/employer/talent) is the authorization field; account_type is a display attribute; password_hash is nullable and NULL until set; email is citext and globally unique since 000004
sessions server-side sessions stores only SHA-256 of the token; two expiries; ON DELETE CASCADE from users
user_preferences per-user settings three real columns plus an extra jsonb blob
job_postings the shop window status enum draft/active/paused/closed; created_by is server-owned; GIN index on skill_requirements
job_applications applications status enum carries seven values; email is citext; interview_id is a deliberate soft reference with no FK
ai_interviews screening interview results references an application; ownership is by reference, not by column
worker_profiles a worker's own profile user_id is server-owned for talent callers; GIN index on completed_courses
staff the employment record of the workforce operator-facing: endorsements, review dates, reviewer names
assignments who is on which position a GiST index over the assignment period
shift_records attendance read-only over the API; the seeder owns these rows
evidence proof of work submitted by a worker, verified by the organization
user_activity the audit log append-only by schema; four server-derived identity columns
courses, learning_paths the training library organization-nullable: NULL org_id is the shared platform library
badges badge definitions the table exists; no endpoint reads it (see §17)
agent_definitions, skill_definitions authored definitions see below

Conventions

  • uuid primary keys (gen_random_uuid()), plus a nullable unique legacy_id text.
  • timestamptz throughout; created_date / updated_date keep their frontend names.
  • Native enums for closed vocabularies (16 types in 000001), text + CHECK where the set is still moving. The definition tables use text + CHECK for both visibility and status.
  • Email columns are citext.
  • 53 CHECK constraints were introduced in 000001; one was dropped in 000003.

Ownership and visibility model on the definition tables

agent_definitions / skill_definitions

  org_id          NOT NULL, even for a personal definition
                  (a user belongs to exactly one organization, so the org
                   predicate applies to every read whether ownership does or not)

  visibility      'personal' | 'organization'          CHECK
  owner_user_id   set if and only if visibility='personal'   CHECK
                  ((visibility = 'personal') = (owner_user_id IS NOT NULL))
                  ON DELETE CASCADE — a personal definition dies with its owner

  created_by      always the author, for attribution
                  ON DELETE SET NULL — a shared definition survives its author

Uniqueness is per tier, by two partial unique indexes — not one global unique constraint, because shadowing by id is the point:

..._personal_key   UNIQUE (owner_user_id, definition_id) WHERE visibility='personal'
..._org_key        UNIQUE (org_id,        definition_id) WHERE visibility='organization'

Other constraints on both tables: definition_id ~ '^[a-z0-9][a-z0-9-]*$'; length(markdown) BETWEEN 1 AND 65536; agent status IN ('draft','published','archived') and version >= 1; skill status IN ('active','inactive') and no version column.

Indexes: (org_id, visibility) for tier listing (which also covers the org_id FK), (owner_user_id) for a user's own definitions and the cascade check, and a partial index for the live rows — (org_id, visibility) WHERE status='published' for agents, and the equivalent active-only index for skills. There is deliberately no index on created_by, because it is attribution only.

Verified seed counts

Counted directly from seed/fixtures/seed.json in this repository:

Entity Records in fixture
Course 40
JobApplication 24
UserActivity 15
RoleCategory 9
WorkerProfile 9
JobPosting 8
Certification 8
AIInterview 4
Badge 4
Staff 3
Evidence 3
LearningPath 2
User 1
Assignment 0 (empty in the source by design)

ShiftRecord is not in the fixture — it is generated. docs/api-contract.md §13.7 records the generated figure as 115 records (2 absent, 1 no-show, 5 late, 107 present) with the qualifier that the set is anchored to the day the seeder runs.

docs/api-contract.md §13.7 also records that "6 job postings" in an earlier brief was the active count; the fixture carries 8 (6 active, 1 paused, 1 closed).


6. Seed & Data Strategy

Source

seed/fixtures/seed.json is generated, not written: the frontend's src/api/seed.js is executed through Vite and what it exports is serialised. The reason is stated in internal/seeder/seeder.go: ids, dates, numbers and enum values then arrive exactly as the demo has them, with no transcription step and nothing to drift.

Shift records — generated, not snapshotted

internal/seeder/shifts.go is a Go port of the frontend's attendanceSeed.js. Shift records are the one collection whose dates are anchored to now rather than to a fixed calendar: the frontend's dataResolver.inPeriod windows every collection on created_date, so a frozen snapshot would read as permanently empty a fortnight later, and the attendance and overtime features would have nothing to show.

Nothing in the generator is random. The distribution is deterministic given the date the seeder runs, over a 56-day window, across a three-person roster designed so the demo has a control, an anomaly and a trend:

Person Pattern
Marco Rivera the control — reliable, weekend event overtime
Marcus Williams attendance degrading over the last fortnight (the anomaly)
Chef Antoine Dubois present throughout, overtime climbing week on week (the trend)

UUID strategy

seeder.DeterministicUUID derives a UUID v5 from the source id over a fixed namespace (krow-seed-v1...), implemented with SHA-1 and the RFC 4122 version and variant bits. Changing the namespace re-keys the entire dataset, so it is a constant rather than configuration. Foreign key values in the fixture are run through the same derivation, so a reference resolves to the row it named.

Idempotency

Explicit upsert inside one transaction. Every record's primary key is derived from its source id, so re-running targets exactly the same rows and ON CONFLICT (id) DO UPDATE restores each one to its seeded values. Consequences, stated in the package documentation:

  • Records created through the API survive a re-seed — nothing is deleted.
  • A column the fixture does not carry is left as it is.
  • Entities are inserted in a fixed order chosen so every foreign key is satisfied by the time it is referenced.

The one exception: pruning

Shift records are also pruned (pruneShiftRecords). Upsert alone cannot converge a rolling window: yesterday's generated set and today's overlap but are not the same set, so without a prune, re-seeding on a later day would leave stale rows behind. The prune is scoped to the seeded organization, and the number pruned is reported in the seeder's result rather than being silent — a delete during a seed should not be something you have to read the source to discover.

Four tests cover exactly this behaviour: TestReseedOnALaterDayLeavesNoStaleShifts, TestReseedIsConvergentAcrossAWeek, TestReseedSameInstantPrunesNothing, and TestPruneIsScopedToTheSeededOrganization.

Limitations

  • The demo user's password_hash is NULL after seeding. Nothing in the fixture, the migrations or this repository contains, generates or defaults a password. A password enters the system only through cmd/setpassword, typed by a person or piped on stdin.
  • Assignment is empty in the source and stays empty.
  • Attendance and overtime are among the richest features in the product and will be empty in any real deployment until a rostering source exists — an import, an integration, or a scheduling UI, none of which exist. This is recorded as U1 in the contract.

7. API Architecture

Conventions

Aspect Behaviour
Base path /api/v1
Resource naming kebab-case plural; mass nouns singular (/staff, /evidence, /user-activity)
Identifiers uuid in the path; the frontend's original string ids live in legacy_id
Content type application/json; charset=utf-8 both ways
Field naming snake_case, identical to the frontend's names. The single exception is /me/preferences, which is camelCase
Dates ISO 8601 with offset, rendered by the database as YYYY-MM-DDTHH:MM:SS.mmmZ
Partial updates PATCH, shallow merge. There is no PUT
Reserved query params sort, limit, offset. Every other parameter is a field filter
Request body cap 4 MiB (maxBodyBytes)
Cache headers every JSON response sets Cache-Control: no-store

Response envelopes

Success — a record:

{ "data": { "id": "…", "title": "…", "created_date": "…" } }

Success — a collection:

{
  "data": [ … ],
  "meta": { "total": 24, "limit": 100, "offset": 0, "returned": 24, "truncated": false }
}

meta.truncated is total > offset + returned. It exists so that the silent truncation recorded in contract §12.2 is fixable without another contract change.

Failure:

{ "error": { "code": "validation_failed", "message": "…", "details": { "field": "…" } } }

details is always present, as an object, even when empty.

Filtering, sorting, pagination

  • Filtering — a single value is equality (?status=hired); a repeated parameter is membership (?status=hired&status=interview). A filter name is resolved to a real column before any SQL is assembled; an unknown name answers 400. Array and JSON columns are not filterable and say so.
  • Sorting — ?sort=field ascending, ?sort=-field descending, ?sort= (empty) means unordered. Every ordered query renders `ORDER BY NULLS LAST, .id` — nulls last in **both** directions, and the id tiebreaker always.
  • Pagination — limit and offset, both non-negative integers. limit is clamped to MaxLimit = 1000. Each resource's default limit is the literal argument at its frontend call site.
  • Error codes and HTTP status

    Code HTTP When
    invalid_query 400 a query parameter is malformed or names an unknown/unfilterable/unsortable field
    unauthorized 401 no valid session
    forbidden 403 the caller's role does not permit the operation
    not_found 404 no such row within the caller's organization and ownership scope
    method_not_allowed 405 the path exists but not under this method
    conflict 409 a uniqueness constraint was violated
    validation_failed 422 a body field is unknown, blank, null on a NOT NULL column, an invalid enum value, or a required field is absent
    rate_limited 429 too many failed sign-in attempts; accompanied by Retry-After
    internal 500 anything else; detail goes to the log, never to the client

    A resource with no item route at all — assignments, which is only listed and created — answers 404 rather than 405, because 405 requires the path pattern to exist under some other method.

    Endpoint groups

    Routes registered by the server: 51. routeResources 34 + routeMe 4 + routeAuth 2 + routeDefinitions 10 + GET /health 1.

    docs/api-contract.md §2 enumerates 38 endpoints; those are exactly the 34 resource routes plus the 4 /me routes. The 2 auth routes are specified in §9 and the 10 definition routes are not in the contract document at all (see §17).

    Health

    Method Path Notes
    GET /health Public. Body is one field: {"status": …}. "ok" with 200; "degraded" with 200 when the database is up but unmigrated or left dirty; "unavailable" with 503 when it is unreachable, so a load balancer can act on the status code alone. Nothing about the server, database or schema appears in the body.

    Authentication

    Method Path Notes
    POST /api/v1/auth/login Public. Body {"email", "password", "remember_me"}. 200 returns the user in the same shape as GET /me and sets the session cookie. 422 when a field is absent. 429 when rate limited. Every credential failure is the same 401.
    POST /api/v1/auth/logout Public and idempotent. Revokes the session behind the cookie if there is one, expires the cookie, and answers 200 with {"data":{"status":"signed_out"}} — including when the cookie is absent, stale, or was never valid.

    Current user

    Method Path Notes
    GET /api/v1/me The user behind the session cookie, with preferences embedded
    PATCH /api/v1/me Shallow merge. Only full_name and account_type are writable
    GET /api/v1/me/preferences camelCase keys
    PATCH /api/v1/me/preferences Shallow merge, returns the merged object

    Core application entities

    Registered only where the resource declares the operation, so an unsupported one is answered by the mux rather than by a handler that has to know to refuse.

    Resource Registered operations
    job-postings list, get, create, update
    job-applications list, create, update, delete
    ai-interviews list, create
    staff list, create, update
    worker-profiles list, create, update
    courses list, get, create, update
    learning-paths list
    role-categories list, create
    certifications list, create, delete
    user-activity list, create
    evidence list, create, update
    assignments list, create
    shift-records list
    badges none — Ops: 0

    Agent Definitions

    Method Path
    GET /api/v1/agent-definitions
    POST /api/v1/agent-definitions
    GET /api/v1/agent-definitions/{id}
    PATCH /api/v1/agent-definitions/{id}
    DELETE /api/v1/agent-definitions/{id}

    Skill Definitions

    Method Path
    GET /api/v1/skill-definitions
    POST /api/v1/skill-definitions
    GET /api/v1/skill-definitions/{id}
    PATCH /api/v1/skill-definitions/{id}
    DELETE /api/v1/skill-definitions/{id}

    Accepted query parameters on both collections: visibility, status, definition_id, sort, limit, offset. Any other parameter answers 400. Default limit 100; default sort -created_date. Sortable fields: created_date, updated_date, name, definition_id, status, version.

    Runtime

    Runtime execution is currently an internal backend boundary; no public runtime execution endpoint is implemented. No route in internal/httpserver references the runtime package.


    8. Authentication, RBAC & Multi-Tenancy

    Users

    users carries email (citext, globally unique since migration 000004), password_hash (nullable, and NULL until a password is set), role, account_type, status, org_id and last_login_at. User.IsActive() is status == "active"; User.CanAuthenticate() additionally requires a non-empty password hash.

    Password handling

    • argon2id via golang.org/x/crypto/argon2, with the OWASP-recommended parameters: 64 MiB memory, 3 iterations, 4 lanes, 16-byte salt, 32-byte key.
    • Hashes are stored as a complete self-describing PHC record, so verification reads the parameters out of the stored string rather than assuming today's defaults — a future cost increase does not invalidate existing hashes. NeedsRehash exists and is tested.
    • Minimum length 12 bytes (a byte floor, because that is what the KDF consumes), maximum 1024 bytes (so an unbounded body cannot be hashed at 64 MiB per attempt).
    • Nothing in internal/auth logs, and neither a password, a token nor a hash is ever returned in an error or formatted into a string.
    • Passwords enter the system only through cmd/setpassword. There is deliberately no -password flag — a password in argv is visible through ps and lands in shell history — so the routes are an interactive hidden prompt (confirmed twice) or stdin.

    Sessions

    • A session is a row, not a token payload. The token is 32 random bytes from crypto/rand, encoded as 43 characters of unpadded base64url. The database stores only its lowercase hex SHA-256, pinned by a CHECK constraint — a raw token fails that pattern, so the catastrophic-and-silent mistake of storing the secret cannot pass.
    • Plain SHA-256 rather than argon2 is deliberate and the reasoning is recorded: a session token is 256 uniformly random bits, so there is no dictionary to run against it and no work factor worth paying on every request.
    • Two expiries. expires_at slides forward as the session is used; absolute_expires_at is fixed at creation and never moves. Without the second, the first could be slid indefinitely.
    Session kind Idle lifetime Absolute cap
    normal 12 hours 24 hours
    remember_me 30 days 90 days
    • Sliding only happens once a session is past a threshold of its idle window, so a page firing ten requests does not fire ten UPDATEs. The idle window is recovered from the row rather than from the policy, so a session keeps the lifetime it was issued under.
    • A background sweeper deletes expired rows every 15 minutes, running once immediately at startup. Sweeping is housekeeping, not correctness: Authenticate already refuses an expired session and deletes the row as it finds it.
    • ON DELETE CASCADE from users, so a deleted user cannot leave a live session behind.

    krow_session, Path=/, HttpOnly, SameSite=Lax, Max-Age matching the session's lifetime, and Secure whenever APP_ENV != development. The __Host- prefix was considered and declined because it requires Secure, which cannot be set over plain HTTP on localhost; the hardening is done by attributes instead, where it can be conditional. SameSite=Lax rather than Strict (which would drop the cookie on any cross-site navigation) or None (which would require Secure and send the cookie on cross-site POSTs).

    The token is never in a response body. It goes out in a Set-Cookie header and nowhere else.

    Login flow

    1. Parse and validate the shape of the request. A missing field is a malformed request (422), not a failed login.
    2. Check both rate limiters, before any expensive work — an attacker must not be able to make the server hash on their behalf.
    3. Verify the credentials. auth.Credentials.Verify takes the same measurable time whether the email exists or not: an unknown email, a user with no password set, and a wrong password all pay a full argon2id derivation against a lazily-built decoy hash. Account status is checked after the password, so a suspended account is indistinguishable from a wrong password in both answer and timing.
    4. Issue the session and set the cookie. A correct password clears the email's failure counter; the address counter is left alone.
    5. last_login_at is stamped best-effort, after the session exists — a failure there is a lost diagnostic, not a reason to refuse a sign-in that already succeeded.

    Every credential failure produces one identical 401. The reason (no_such_user, no_password, bad_password, not_active) goes to the log at warn, with the email — which was already in the request — and never the password.

    Rate limiting is a fixed-window counter of failed attempts held in this process's memory: 5 per email and 20 per client address per 15-minute window. Two limiters rather than one, because the budgets are deliberately different sizes — an email is one account, while an address may be a whole office behind NAT. Its limitations are documented in the source itself and repeated in §17.

    Authentication middleware

    publicPaths = { /health, /api/v1/auth/login, /api/v1/auth/logout }
    

    An allowlist, not a list of protected prefixes, so the failure mode of forgetting to update it is a route that refuses everyone rather than one that serves everyone. It sits below the router, so even the mux's own 404 is behind it.

    For every other request:

    cookie -> Manager.Authenticate(token)
                hash the token, look up the row, refuse if expired (and delete it),
                slide the expiry if due
           -> users.FindByID(session.UserID)
                re-read on EVERY request, not cached in the session row, so suspending
                an account takes effect on its next request
           -> refuse and revoke if the user is not active
           -> authctx.Identity{UserID, OrgID, Email, FullName, Role, AccountType,
                               Status, SessionID, ExpiresAt}  on the context
           -> orgctx.With(ctx, user.OrgID)
    

    Nothing in the request influences any of it. The middleware reads no body, no query string and no header other than Cookie. A user_id or org_id in a body or query string is ignored.

    Roles and the authority

    users.role, and only users.role. Three values, fixed by the users_role_check constraint. An unrecognised role authorizes nothing.

    Role Who
    admin runs the platform for the organization
    employer runs the organization's hiring and workforce
    talent a worker, acting for themselves

    account_type is not an authorization field. It is a display attribute the user may change on themselves through PATCH /me, and nothing in the API reads it to make a decision.

    Access resolution

            User  (session -> users row)
              |
              +--> Organization   org_id, from the user's row, never from the request
              |         |
              |         +--> org predicate on every read and write
              |              (courses / learning_paths also match org_id IS NULL —
              |               the shared platform library)
              |
              +--> Role  (admin | employer | talent)
                        |
                        +--> handler gate: may this role perform this operation?
                        |         no  -> 403, before any query runs
                        |
                        +--> row predicate: which rows may this role see?
                                  admin, employer -> the whole organization
                                  talent          -> an extra WHERE clause
                                                     -> outside it: 404
    

    Permission matrix

    R = read (list/get), C = create, U = update, D = delete. own means restricted by a SQL predicate.

    Resource Admin Employer Talent
    /me, /me/preferences R U R U R U
    job-postings R C U R C U R (active only)
    job-applications R C U D R C U D R C (own)
    ai-interviews R C R C R C (own)
    staff R C U R C U —
    worker-profiles R C U R C U R C U (own)
    assignments R C R C R (own)
    shift-records R R R (own)
    courses R C U R R
    learning-paths R R R
    role-categories R C R C R
    certifications R C D R C R
    user-activity R C R C R (own) C
    evidence R C U R C U R C (own)
    badges — — —

    Two admin-only operations, each with a reason beyond seniority: POST/PATCH /courses, because a course with a NULL org_id is the shared library and a write there can reach beyond the writer's own tenant; and DELETE /certifications/{id}, because deleting one changes what every existing posting that required it means.

    Ownership predicates (talent callers)

    Resource Predicate
    worker-profiles user_id = <session user id>
    job-applications email = <session email>
    assignments worker_email = <session email>
    shift-records worker_email = <session email>
    evidence worker_email = <session email>
    user-activity user_email = <session email>
    ai-interviews application_id IN (SELECT id FROM job_applications WHERE org_id = … AND email = <session email>)
    job-postings status = 'active' — visibility rather than ownership

    The ai-interviews case also has a write guard: an insert whose ownership is expressed by reference has the reference checked against the caller's own applications, and a reference that is not theirs answers the same 404 a nonexistent application would.

    Server-owned identity columns

    Six columns are filled from the session and ignored if present in a request body, because they are the columns every ownership rule rests on:

    Column Filled with When
    job_postings.created_by session user id always
    user_activity.user_id session user id always
    user_activity.user_email session email always
    user_activity.user_name session full name always
    user_activity.account_type session account type always
    worker_profiles.user_id session user id talent callers only

    Two more are overridden for talent callers and left writable for operators: job_applications.email and evidence.worker_email. The distinction is between who acted and who the row is about — when an admin creates a candidate's profile or files an application on their behalf, the subject is the candidate, not the operator.

    PATCH /me has its own allowlist: only full_name and account_type are writable. role used to be in that map, which would have been a one-line privilege-escalation path the instant sessions existed. Attempts to write a server-owned user field are ignored and logged, so a client asking to change its own role is visible even though the answer is no.

    Cross-user and cross-tenant protection

    Both are predicates, not filters, so a row outside them is never fetched. Tests covering this: TestCrossOrganizationIsolation, TestTalentSeesOnlyTheirOwnRecords, TestTenancyFollowsTheSession, TestIdentityCannotBeSuppliedByTheRequest, TestForbiddenVersusNotFound, TestOrganizationScoping, TestTalentCannotInterviewForAnotherApplication.

    Deny by default

    internal/domain/policy.go maps a URL path to a *Policy. A resource with no entry keeps a nil policy and permits nothing, to anyone — a table added to the schema tomorrow is unreachable until someone writes down who may reach it. TestEveryResourceHasAPolicy makes the omission loud. badges has an empty policy written out explicitly, so the resource is deliberately closed rather than merely forgotten.

    What is not implemented

    No permissions table, no policy engine, no per-record ACL, no role hierarchy, no delegation, no refresh tokens, no multi-factor, no SSO, no OAuth, no password reset flow, no self-service registration, and no email delivery of any kind.

    Firebase Auth is an architectural direction, not the current implementation. The string "Firebase" does not appear anywhere in this repository — Not verified from current repository as a recorded plan; it is recorded here because it was named as a direction for this document.


    9. Agent & Skill System

    The three representations

    Representation Where Purpose
    Source-controlled (shipped) src/agents/**/*.md, src/skills/**/*.md in the frontend repository product source, versioned in Git, bundled at build time. No rows in this database.
    PostgreSQL authored agent_definitions, skill_definitions authored in the app; the Markdown is the authoritative artefact, the columns are projections
    Runtime runtime.Agent, runtime.Skill the parsed definition plus its database identity and ownership, prepared for execution

    The three tiers share one id namespace, and shadowing across them is the point: a personal definition may carry the same definition_id as an organization one, which may carry the same id as a shipped one.

    Agent — fields parsed by internal/definition

    definition.Agent:

    Field Type Notes
    ID string from id:; falls back through name and then the origin path
    Name string falls back to Untitled agent, which the validator then refuses
    Description string
    Status string draft / published / archived
    Version int
    Pages []string canonical surface ids — mapped through the frontend's canonicalPage
    Icon string checked against a closed vocabulary (10 values in the captured oracle)
    Reasoning string fast / balanced / deep
    Trigger string
    WebSearch bool
    Skills []string ids of skills the agent depends on
    Subagents []string
    Starters []Starter {Label, Prompt} — conversation starters
    Knowledge []Knowledge {ID, Label, Kind, Body, URL}; kinds note / link / skill-reference
    Permissions Permissions {Owner, Access, People[]}; access all / specific; roles manager / editor / viewer. Parsed and not enforced
    Instructions string the body's ## Instructions section — prose belongs under a heading
    Errors []string what this definition lost on the way in, carried on the record rather than thrown
    Body string the Markdown after the frontmatter

    runtime.Agent additionally carries DatabaseID, Visibility, OwnerUserID, ResolvedSkills and RawMarkdown.

    Skill — fields parsed by internal/definition

    definition.Skill:

    Field Type Notes
    ID string
    Name string falls back to Untitled skill
    Description string
    Status string active / inactive
    Pages []string as the author wrote them, not canonicalised
    Kind string
    Category string
    Actions []string
    Triggers []string
    DeclaredTriggers bool whether the author declared any
    Prompt *string nullable
    SkillID *string nullable
    Levels []Level {Level, Label, Summary}, read from the body's own headings
    Body string Markdown after the frontmatter, trimmed
    Deferred []string frontmatter blocks this package does not check — see §10

    runtime.Skill additionally carries DatabaseID, Visibility, OwnerUserID and RawMarkdown.

    One asymmetry that must not be "fixed"

    Agent.Pages holds canonical surface ids; Skill.Pages holds the strings the author wrote. That is not an oversight: the frontend's normalizeAgent maps every page through canonicalPage and its parseSkill does not. Making the two agree in Go would make each one disagree with its own editor. The comment in the source says so explicitly.

    Vocabularies (captured from the frontend)

    18 pages with aliases; agent statuses draft/published/archived; reasoning fast/balanced/deep; 10 icons; knowledge kinds note/link/skill-reference; access all/specific; permission roles manager/editor/viewer. TestVocabularyMatchesFrontend asserts the Go tables against the captured ones.

    Backend-only bounds

    Two rules the frontend editor does not have, both flagged BackendOnly so they are distinguishable from shared rules:

    • MaxMarkdownLength = 65536 characters (PostgreSQL length() counts characters), matching the markdown_size CHECK.
    • MaxVersion = 2147483647, the range of agent_definitions.version as a PostgreSQL integer.

    Both turn a constraint violation into a message an author can act on.


    10. Parser & Definition Compatibility

    Why the two parsers must agree

    Frontend JS parser              (src/lib/skills, src/lib/agents)
            |
       compatibility contract       (internal/definition + testdata/oracle.json)
            |
    Backend Go parser               (internal/definition)
    

    A definition is authored in the browser and stored by the server, so both parsers see it. Disagreement is silent in both directions: a definition the editor accepts and the API rejects fails at save time after looking valid; a definition the API accepts and the editor rejects is stored and then cannot be rendered.

    The supported YAML subset

    internal/definition/yaml.go is a line-for-line port of the frontend's yaml.js.

    Supported, and nothing else: block maps and block sequences nested to any depth; scalars (strings, integers, floats, booleans, null); quoted strings for values containing : or #; - key: value (a mapping whose first key sits on the dash); # comments and blank lines.

    Not supported: anchors, aliases, merge keys, multi-document files, flow mappings, flow sequences, block scalars, tags. These are not silently half-read — an unparseable line is an error carrying its 1-based line number within the frontmatter, worded exactly as the frontend words it, because an author who sees one message in the editor and another from the API is being told about two different problems.

    No YAML dependency, deliberately. A general library would accept a much larger language than the frontend does, and every construct it accepted and the frontend did not would be a definition the backend stores and the editor cannot read.

    Nothing here evaluates anything. There is no reflection, no template, and no code path from a definition to execution of any kind. A definition is configuration, and the parser is the boundary that keeps it configuration.

    Document layer

    frontmatter.go handles byte-order mark stripping, CRLF normalisation, leading blank lines, fence recognition (including a trailing tab after the closing fence), and body extraction. A second --- in the document is body, not a new frontmatter block.

    Validation and normalization

    • Everything is optional. A definition declaring only an id and a name normalizes to a working agent with documented defaults.
    • Nothing unknown survives — statuses, reasoning modes, pages, icons, knowledge kinds and permission roles are checked against closed tables, and an unrecognised value is a named error rather than a dropped key.
    • What validates is kept. One bad entry costs its author that entry and a message, never the rest of the file (Agent.Errors).
    • An agent with no skills is deliberately not refused, because several product pages have no assistant skills and answer from their own responder.
    • The rejection messages reuse the frontend's own wording wherever the rule is shared, including its quirks — ValidateAgent performs a literal strings.Contains(raw, "name:") substring test because the JavaScript does.

    Raw Markdown is never rewritten

    Normalization reads; it does not rewrite what is stored. The Markdown handed in is the Markdown that goes to the database, byte for byte. TestParsingDoesNotMutateSource, TestNormalizationIsNotStorage and TestMarkdownIsStoredVerbatim assert it from three directions.

    Deferred blocks

    A skill may carry a ui: block (declarative page sections) or an owliver: block (assistant capabilities). Validating those means reproducing roughly 1,500 lines of closed vocabulary describing what the frontend can render — placements, data sources, section types, periods — none of which the backend stores, projects or acts on.

    The package therefore does not check them. It records their presence on Skill.Deferred, so the gap is a value a caller can see rather than an assumption. The one consequence is stated exactly in the source: skill-examples/board-invalid-context.md is rejected by the frontend on a rule about which placement may supply which data source, and accepted here. It is the only definition in the corpus where the two disagree, and the conformance suite asserts that it stays the only one.

    Visibility is deliberately absent from the parser: the frontend ignores a visibility: key in a definition entirely. It is a storage tier chosen by the request and checked by a database constraint. A definition cannot name its own tenancy.

    The conformance corpus

    internal/definition/testdata/oracle.json (9,633 lines) is not hand-written. It is captured by scripts/oracle.mjs, which loads the real frontend module graph through Vite — import.meta.glob, the @/ alias and raw Markdown loading behave exactly as they do in the app — and records what the JavaScript parser did with every definition. Counted from the fixture in this repository:

    Set Entries Breakdown Accepted by JS
    corpus (shipped definitions) 37 28 skills, 9 agents 36 accepted, 1 rejected
    cases (adversarial) 132 87 skills, 45 agents 86 accepted, 46 rejected

    The adversarial cases come from scripts/cases.mjs, which stores raw bytes as an author could actually produce them — nothing is normalized on the way in, because the point is what the two parsers do with the awkward form.

    The assertion is therefore not "Go agrees with a description of the frontend" but "Go agrees with the frontend", replayed. A frontend change that alters parsing fails these tests, which is the intent: the contract cannot drift silently in either direction.

    Mutation checks

    TestMutationsWouldBeCaught exists because tests that pass against a broken parser are not tests. Each entry is a plausible mistake in the Go package, and each must be caught by a real definition changing its meaning rather than by an assertion written to notice it. The mutations covered include: dropping the BOM strip; dropping CR normalisation (so candidates\r stops being a page); trimming the closing fence too eagerly or not at all; treating a second --- as frontmatter; treating a # inside a word as a comment; failing to treat a spaced # as a comment; letting a colon inside quotes split the value; taking the first rather than the last of a duplicate key; accepting ragged indentation instead of reporting its line; and canonicalising a skill's page name.

    Parser behaviours worth knowing

    • pages: candidates (a bare string, not a sequence) is accepted for an agent and refused for a skill, because normalizeAgent coerces and parseSkill requires a real sequence.
    • A definition with no id:, no name: and a valid pages: list gets the id custom on the frontend and is accepted; the Go parser derives the same id from the same default parameter (AuthoredPath = "custom") so that it does not refuse a definition the editor accepts.

    11. Agent & Skill CRUD

    Create flow

    POST /api/v1/agent-definitions   { "markdown": "...", "visibility": "personal" }
       |
    authctx.MustFrom(ctx)           identity from the session
       |
    body["markdown"] present?       absent -> 422 "Paste or upload a Markdown definition."
       |
    ValidateAgent(markdown)         frontend rules + backend bounds -> 422 with the message
       |
    visibility resolved             defaults to "personal"; must be personal|organization
       |
    organization? role gate         talent (or an unrecognised role) -> 403
       |
    ParseAgent(markdown)            -> id, status, version, name, description, pages
       |
    AgentInsertInput                org_id = session org
                                    created_by = session user
                                    owner_user_id = session user  (personal only)
                                    markdown = the caller's bytes, unchanged
       |
    INSERT ... RETURNING            partial unique index -> 409 on a duplicate id per tier
       |
    { "data": { … the stored row … } }
    

    Skills follow the same path through ValidateSkill / ParseSkill, minus version.

    List, get, update, delete

    Operation Behaviour
    GET collection Scoped to the caller's organization and ownership: visibility='organization' OR (visibility='personal' AND owner_user_id = <session user>). Optional visibility filter narrows to one tier. Returns the standard page envelope with meta.
    GET /{id} Same predicate. A non-UUID id, or a row outside the predicate, answers 404 — the caller cannot distinguish the two.
    PATCH /{id} Reads the row first (within scope). If the stored row is organization, a talent caller is refused 403. visibility cannot be changed after creation — an attempt answers 422 immutable. If markdown is supplied it is re-validated, re-parsed, and every projection is recomputed from it. Otherwise a bare status change is accepted against the closed status list.
    DELETE /{id} Idempotent, matching the contract's §12.7 semantics: a non-UUID id or an already-absent row answers 200 with {"data":{"id":"…"}}. An organization row still enforces the role gate before deleting.

    Scope, ownership and tenancy

    • Personal scope — visibility='personal', owner_user_id set to the session user by the server. A caller never supplies it.
    • Organization scope — visibility='organization', owner_user_id NULL, writable by admin and employer only.
    • Tenant isolation — org_id = <session org> is on every read, update and delete, including for personal definitions, whose org_id is NOT NULL for exactly this reason.
    • Duplicate handling — the two partial unique indexes make an id unique per owner or per tenant, never globally. A collision surfaces as conflict / 409 through the repository's error translation.
    • Server-owned fields — id, org_id, owner_user_id, created_by, created_date, updated_date, and every projection (definition_id, status, version, name, description, pages). The only client-supplied values are markdown, visibility (at creation), and a bare status on PATCH.
    • Projection synchronization — a PATCH that changes markdown recomputes every projected column in the same statement, so a column can never describe a different document than the one stored.
    • Verbatim Markdown — the bytes the caller sent are the bytes stored.

    Tests covering this surface

    TestAgentCreate, TestSkillCreate, TestDefinitionsList, TestDefinitionsGet, TestDefinitionsPatch, TestDefinitionsDelete, TestSecurityAndSQLInjection, TestFullCRUDFlowAndProjections (internal/httpserver/definitions_api_test.go, 848 lines).


    12. Runtime Architecture

    What the package contains

    Type Role
    Loader loads a stored definition by UUID or definition_id, re-parses and re-validates it, and returns a runtime representation
    runtime.Agent parsed agent + DatabaseID, Visibility, OwnerUserID, ResolvedSkills, RawMarkdown
    runtime.Skill parsed skill + the same identity and ownership fields
    ExecutionInput {Identity, TargetID, Input, Parameters, Context}
    ExecutionResult {Success, Output, AgentID, AgentVersion, ResolvedSkills, Error}
    RuntimeError structured {Code, Message, Target, Cause} with Unwrap
    AgentExecutor / SkillExecutor the execution boundary interfaces
    UnavailableExecutor the only implementation; refuses execution
    Engine ties loader + executors together; RunAgent, RunSkill

    Typed errors

    ErrNotFound, ErrUnauthorized, ErrInvalidDefinition, ErrDraftAgent, ErrArchivedAgent, ErrInactiveSkill, ErrNotExecutable, ErrDependencyMissing, ErrDependencyInactive, ErrCircularDependency, ErrExecutorUnavailable.

    Loading

    LoadAgent / LoadSkill accept either a UUID (matched by regex) or a definition_id, and dispatch to the repository accordingly. The stored markdown is then re-validated and re-parsed — the runtime does not trust the projected columns for anything but identity, visibility and ownership. A row whose markdown is absent or empty answers ErrInvalidDefinition.

    Both lookups go through DefinitionsRepo, so tenant and ownership scoping applies to the runtime exactly as it does to the CRUD API: org_id = <session org> and visibility='organization' OR (visibility='personal' AND owner_user_id = <session user>).

    Status eligibility

    Agent                                    Skill
      published  -> eligible                   active    -> eligible
      draft      -> ErrDraftAgent              inactive  -> ErrInactiveSkill
      archived   -> ErrArchivedAgent           other     -> ErrNotExecutable
      other      -> ErrNotExecutable
    

    LoadExecutableAgent applies the agent rule and then resolves dependencies; LoadExecutableSkill applies the skill rule.

    Dependency resolution

    Agent
      |
    agent.Skills  (ids from the `skills:` frontmatter)
      |
    for each id, in declared order:
      |
      +-- already resolved?  -> skip                     (deterministic deduplication)
      |
      +-- LoadSkill(ident, id)                           tenant + ownership scoped
      |     |
      |     +-- by definition_id: ORDER BY
      |           CASE WHEN visibility='personal' THEN 1 ELSE 2 END
      |           LIMIT 1                                 personal shadows organization
      |     |
      |     +-- not found -> ErrDependencyMissing
      |
      +-- status != 'active' -> ErrDependencyInactive
      |
      v
    agent.ResolvedSkills  ([]*Skill, in declared order)
    

    An agent with no skills: entries resolves to an empty list rather than an error.

    Cross-tenant dependency protection falls out of the repository predicate: a skill belonging to another organization is simply not found, so the dependency resolves to ErrDependencyMissing rather than to another tenant's definition. TestRuntime_DependencyResolution covers the cross-org case explicitly, and TestRuntime_PersonalSkillShadowing covers the shadowing precedence.

    Cycle detection is present in the code (inProgress map, ErrCircularDependency) but is not reachable in the current design — see §17.

    The executor boundary

    type AgentExecutor interface {
        ExecuteAgent(ctx, agent *Agent, input ExecutionInput) (*ExecutionResult, error)
    }
    type SkillExecutor interface {
        ExecuteSkill(ctx, skill *Skill, input ExecutionInput) (*ExecutionResult, error)
    }
    

    NewEngine defaults both to UnavailableExecutor and accepts WithAgentExecutor / WithSkillExecutor overrides. UnavailableExecutor.ExecuteAgent returns a result carrying the agent id, version and resolved skill ids, with Success: false and Error: ErrExecutorUnavailable, and returns that error.

    Real LLM execution is not part of the current executor boundary. There is no LLM client, no prompt assembly, no tool dispatch and no external call anywhere in go-api/. The boundary exists so that adding one is an implementation of an interface rather than a change to the loading, scoping and eligibility rules.

    Reachability

    Nine test functions in internal/runtime/runtime_test.go (890 lines) exercise the package: loader, status eligibility, version semantics, dependency resolution, executor boundary, personal skill shadowing, malformed Markdown, and skill execution. Outside those tests, nothing in the repository calls into the package — see §17.


    13. Security Architecture

    Controls that are present and verifiable

    Control How it is implemented
    SQL parameterization Every value reaches PostgreSQL as a bind parameter cast to its declared type. No identifier ever comes from user input: a filter or sort name is resolved to a *domain.Column before any SQL is assembled, and an unresolved name is rejected. TestSecurityAndSQLInjection covers the definitions surface.
    Deny-by-default authorization A resource with no policy permits nothing to anyone. TestEveryResourceHasAPolicy, TestNilPolicyDeniesEverything, TestUnknownRoleIsDenied.
    Tenant isolation org_id = <session org> is a predicate on every read and write, taken from the user's row and never from the request.
    Ownership checks Predicates in the same WHERE clause, so count(*) matches what the caller may see and no unauthorized row is ever fetched.
    Server-owned fields Eight identity columns are filled from the session and override anything in the body; PATCH /me writes only two fields, and attempts on the rest are ignored and logged.
    Password handling argon2id at OWASP parameters, self-describing PHC records, a 12-byte floor and a 1024-byte ceiling, no password ever logged or echoed, and no -password CLI flag.
    Credential-failure uniformity One 401 for every failure mode, and a decoy argon2id derivation so an unknown email costs the same time as a known one. Account status is checked after the password.
    Sessions 256-bit random opaque tokens; only SHA-256 stored, pinned by a CHECK; sliding plus absolute expiry; per-request user re-read so suspension takes effect immediately; expired sessions deleted on contact and swept every 15 minutes.
    Cookies HttpOnly, SameSite=Lax, Path=/, Secure outside development, Max-Age matching the session. The token never appears in a response body.
    Rate limiting Failed sign-in attempts counted per email (5) and per client address (20) in a 15-minute window, checked before any hashing.
    CORS Exact-match allowlist, echoed one origin at a time, * rejected at config load, Vary: Origin always set, and the middleware not installed at all when the allowlist is empty.
    Input validation Unknown body fields answer 422 with the field named; enum values are checked against the column's declared set; NOT NULL columns refuse an explicit null; required fields are enforced on create; request bodies are capped at 4 MiB.
    Health endpoint exposure One status word and nothing else. Version, database name, schema, migration version, table count, error text, environment and uptime all moved to the server log. TestHealthLeaksNoInfrastructure, TestHealthUnavailableSaysNothingAboutWhy.
    Error exposure An unrecognised error is flattened to {"code":"internal","message":"internal error"}; the detail goes to the log. A 403 names neither the caller's role nor the roles that would have worked.
    Existence non-disclosure 403 for role, 404 for row. A caller cannot tell "exists and is not yours" from "does not exist". TestForbiddenVersusNotFound.
    Panic containment recoverer turns a panic into a logged 500 rather than a dropped connection.
    Config guards System schemas rejected; sslmode=disable rejected in production; CORS * rejected.
    Operational guards No make migrate-drop; migrate-down gated on APP_ENV=development; test databases are named krow_backend_autotest_<pid> and never touch the development database.
    Secret hygiene .env is gitignored; .env.example carries no real values; the connection string is redacted in logs (DBConfig.Redacted).

    Honest limitations

    These are the repository's own words where it states them, and direct observation where it does not:

    • Rate limiting is per-process and in-memory. Two API instances behind a load balancer each allow the full budget, so the effective limit is the limit times the instance count, and a restart clears every counter. Multi-instance deployment needs shared state.
    • The client address is RemoteAddr. Behind a reverse proxy every request appears to come from the proxy, so the per-address budget becomes global. Reading X-Forwarded-For instead would be worse until a trusted-proxy list exists, because a client can send that header itself and mint a fresh budget per request.
    • The limiter map is bounded by pruning, not by a hard cap, so a flood from many distinct addresses grows it until the next prune.
    • There is no CSRF token. The defence is SameSite=Lax, which does not send the cookie on cross-site POST/PATCH/DELETE. That is adequate for the current same-origin and localhost-development posture and would need re-examination alongside any change to the cookie's SameSite value.
    • CORS never sets Access-Control-Allow-Credentials. See §17.
    • The permissions: block in a definition is parsed and not enforced. No definition_permissions table exists.
    • No rate limiting exists on any endpoint other than login.
    • No audit trail of authorization refusals beyond the application log.
    • No secret management, key rotation, or encryption at rest is configured in this repository; those are deployment concerns and infrastructure/ is empty.

    Nothing here should be read as a claim that the service is production-hardened. It is a claim about which controls exist and which do not.


    14. Testing & Verification

    Toolchain checks — run against this repository

    Check Command Result observed
    Formatting gofmt -l ./cmd ./internal no files listed
    Static analysis go vet ./... no diagnostics
    Build go build ./... builds
    Tests go test ./... -count=1 every package's tests ran without failure

    Test counts — counted from this repository

    Measure Count
    Top-level Test… functions 175
    Total test entries executed, including subtests 920
    Failures 0
    Skipped 0 (PostgreSQL was reachable on the machine used)

    Per package:

    Package Top-level test functions Lines of test code (largest files)
    internal/httpserver 78 api_test.go 1,114; auth_test.go 926; definitions_api_test.go 848; rbac_test.go 730
    internal/auth 26 session_test.go 488; schema_test.go 376; password_test.go 254
    internal/domain 22 definitions_schema_test.go 748; policy_test.go 191
    internal/seeder 18 seeder_test.go 300; shifts_convergence_test.go 178
    internal/definition 13 conformance_test.go 928
    internal/runtime 9 runtime_test.go 890
    internal/service 9 service_test.go 217

    cmd/api, cmd/seed, cmd/setpassword, internal/authctx, internal/config, internal/db, internal/orgctx, internal/repo and internal/testutil have no test files of their own; internal/repo is exercised through the service and httpserver suites.

    Database-backed tests

    internal/testutil builds a disposable database per test process: dropped, recreated, migrated and seeded, named krow_backend_autotest_<pid> so that concurrently-running test packages cannot collide, and named distinctly enough that it cannot be confused with a real database. The package documentation states that nothing in it ever connects to, reads or drops the development database. Tests skip rather than fail when PostgreSQL is unreachable.

    What the important tests prove

    Area Representative tests What they establish
    Contract semantics TestNullsSortLastInBothDirections, TestStableSortWithIDTiebreaker, TestEndpointSpecificDefaults, TestLimitAndTruncationMeta, TestOffsetPaging The two ordering guarantees hold, and each endpoint carries its own defaults. The README records that the ordering tests were verified to fail when the guarantee is removed.
    Validation TestCreateRejectsUnknownFields, TestCreateRejectsMissingRequiredAndBlank, TestCreateRejectsInvalidEnum, TestCreateIgnoresServerOwnedFields Unknown fields are a 422 rather than silent data loss; server-owned fields are ignored.
    Idempotent delete TestDeleteIsIdempotent 200 whether or not a row matched, as §12.7 requires.
    Seed fidelity TestSeedMatchesFixtureCounts, TestSeedPreservesSourceValues, TestSeedRegressionAnchors The database is compared against the fixture field-by-field, not against numbers typed into a test.
    Seed convergence TestSeedIsIdempotent, TestReseedOnALaterDayLeavesNoStaleShifts, TestReseedIsConvergentAcrossAWeek, TestPruneIsScopedToTheSeededOrganization Re-seeding converges the rolling shift window without touching other organizations or API-created rows.
    Password and session TestDefaultPasswordParamsMeetOWASP, TestSessionCannotOutliveItsAbsoluteDeadline, TestDatabaseRefusesARawToken, TestUserDeletionCascadesToSessions The cost parameters, the absolute cap, the CHECK that refuses a raw token, and the cascade all hold at the database level.
    Authentication behaviour TestLoginFailuresAreIndistinguishable, TestLoginSetsHardenedCookieAndNeverReturnsTheToken, TestCookieIsSecureOutsideDevelopment, TestSessionSlidesButNotForever, TestSuspendedUserIsRejectedMidSession, TestLoginRateLimitIsPerEmailAndPerAddress Enumeration resistance, cookie hardening, sliding-with-a-ceiling, mid-session suspension, and both limiter dimensions.
    Authorization TestRoleMatrix, TestTalentSeesOnlyTheirOwnRecords, TestCrossOrganizationIsolation, TestForbiddenVersusNotFound, TestServerOwnedIdentityCannotBeSupplied, TestTalentCannotInterviewForAnotherApplication, TestTalentSeesOnlyActivePostings, TestExpiredSessionIsRefusedBeforeRoleCheck The full matrix, the row predicates, the 403/404 distinction, and that identity cannot be supplied by the request.
    Policy invariants TestEveryResourceHasAPolicy, TestDerivedColumnsAreReadOnlyOrTalentScoped, TestOnlyTalentIsRowScoped, TestTalentScopesNameRealColumns Regenerating the descriptors cannot silently drop a rule, and a scope cannot name a column that does not exist.
    Schema invariants TestDefinitionTablesShape, TestUniquenessPerTier, TestForeignKeysAndDeleteBehaviour, TestMigrationAddsExactlyTwoTables, TestMigration000005IsReversible, TestEveryMigrationHasADownFile The definition schema is asserted against the live database, including the deferred tables that must not exist.
    Parser conformance TestFrontmatterTreeParity, TestAcceptanceParity, TestRejectionMessageParity, TestProjectionParity, TestAdversarialCoverage, TestDeferredBlocksAreReported Go matches the captured JS output over 37 shipped definitions and 132 adversarial cases, including the exact rejection wording.
    Parser anti-regression TestMutationsWouldBeCaught, TestParsingDoesNotMutateSource, TestNormalizationIsNotStorage Plausible parser mistakes are caught by real definitions changing meaning, and normalization never rewrites what is stored.
    Runtime TestRuntime_StatusEligibility, TestRuntime_DependencyResolution, TestRuntime_PersonalSkillShadowing, TestRuntime_ExecutorBoundary, TestRuntime_MalformedMarkdown Eligibility, resolution across tenants and tiers, shadowing precedence, and that the boundary refuses rather than pretends.
    Health TestHealthEndpoint, TestHealthLeaksNoInfrastructure, TestHealthUnavailableSaysNothingAboutWhy The public body carries a verdict and no reconnaissance.
    CORS TestCORSAllowsConfiguredOrigin, TestCORSPreflight, TestCORSRefusesUnknownOrigin, TestCORSIgnoresRequestsWithoutOrigin, TestCORSOffByDefault Exact matching, and that an origin-less caller is unaffected.

    15. Frontend Integration

    React  (krow-demo, a separate repository — not part of this checkout)
       |
    base44Client.js          the transport seam; the one frontend file allowed to change
       |
    HTTP client              Not verified from current repository
       |
    Go API                   /api/v1, JSON, session cookie
       |
    PostgreSQL
    

    What is verifiable from this repository

    • API base path — /api/v1. The server binds HTTP_HOST:HTTP_PORT, defaulting to 127.0.0.1:8080.
    • CORS — an exact-match allowlist from HTTP_CORS_ORIGINS. In development with the variable unset, the defaults are http://localhost:5173, http://127.0.0.1:5173 and the vite preview port. Unset outside development the allowlist is empty and the middleware is not installed, which is same-origin only.
    • Transport compatibility — field names are snake_case and identical to the frontend's; ids are opaque strings to the client; the response envelope is one .data unwrap; meta.total exists because the frontend's bare array shape had nowhere to put it.
    • Preserved client interface — the contract's acceptance criterion is that swapping the transport inside base44Client.js, and changing no other frontend file, leaves the application behaving identically. The contract states that if implementing it required editing krowHooks.js, a page or a component, the contract was wrong and got fixed first.
    • Connected entities — the 13 resources with registered routes, plus /me and /me/preferences.
    • The { persisted } return shape — auth.updatePreferences() returns { user, persisted, error } rather than a bare user, because a swallowed QuotaExceededError used to lose account-authored skills silently. Over HTTP a failed write is already a non-2xx, so the shim synthesises the success shape.

    What is not verifiable here

    • httpClient.js — Not verified from current repository. The string does not appear anywhere in this checkout.
    • Any VITE_API_* environment variable — Not verified from current repository.
    • The current state of the frontend's own tests — Not verified from current repository. The frontend repository is not part of this checkout.

    Known integration mismatches

    Recorded in docs/api-contract.md §12 and still accurate against the current code:

    # Mismatch
    12.1 Four flows write several records in a client-side loop with no transaction and no rollback (useHireCandidate, useAssignWorkers, useScreenAllCandidates, useSubmitChallenge). A failure halfway leaves the database inconsistent. useAssignWorkers issues 3n sequential round-trips for n workers.
    12.2 Collection caps truncate silently. worker-profiles caps at 500 sorted by -krow_score, so profile 501 is invisible. meta.truncated exists to make this fixable; nothing consumes it yet.
    12.3 All aggregation — funnels, KPI tiles, charts, attendance rollups, overtime — computes in the browser over the capped arrays.
    12.4 Search is browser-side substring matching over a concatenated string. No endpoint implements search.
    12.5 Email matching is case-insensitive server-side (citext) where store.js used ===. Deliberate and documented as safe.
    12.7 DELETE on a missing record returns 200, because the frontend deletes inside loops without checking and a 404 would surface an error toast where none appears today.
    12.8 InvokeLLM and UploadFile remain local to the browser. Nine AI workflows run deterministically client-side. No endpoint replaces them.
    §13.4 updated_date is now always populated where store.js returned undefined; one visible effect is that TalentDetailModal.jsx:66 shows the seeded date rather than today.

    16. Important Architectural Decisions

    Decision Why Current implementation Future direction
    Go for the API A single static binary, a strong standard-library HTTP server, and explicit error handling for a service whose main job is correctness at a boundary Current. Go 1.27, net/http only, three direct dependencies —
    PostgreSQL, no ORM The contract's semantics (NULLS LAST both ways, the , id tiebreaker, shallow PATCH, idempotent DELETE) are easier to guarantee in SQL written once than to coax out of a mapper Current. internal/repo builds every statement from a descriptor —
    Descriptors generated from the live schema Column names, types, enum values and nullability then cannot drift from the migrations Current. make gen-resources → resources_gen.go —
    Policy hand-written, kept apart from descriptors Regenerating descriptors must never silently drop an access rule, and a new column must never grant anyone anything by accident Current. internal/domain/policy.go, enforced by TestEveryResourceHasAPolicy —
    One shared query builder, not fourteen repositories The awkward semantics get implemented once and apply identically everywhere Current. internal/repo/repo.go —
    Migrations are the source of truth If the database and migrations/ disagree, migrations/ is right; an applied migration is immutable Current. Five pairs; nothing in go-api/ issues DDL Expand-migrate-contract as a discrete CI deploy step before the binary rolls out
    org_id on every tenant-scoped table Tenancy has to be a predicate on the row, not a convention in the caller Current. Server-owned on every resource; NOT NULL even on personal definitions —
    Ownership as a SQL predicate, not a filter count(*) runs over the same clause, so totals are correct and unauthorized rows are never fetched Current. builder.ownership —
    403 for role, 404 for row A caller must not be able to distinguish "exists and is not yours" from "does not exist" Current. Role gate in the handler before any query —
    API contract derived from call sites A table existing is never a reason for an endpoint to exist Current. docs/api-contract.md; badges has no endpoints Multi-record transactional endpoints would require touching frontend hooks, so they were deliberately left out of v1
    The transport boundary is one frontend file The whole migration is affordable only if nothing above base44Client.js changes Current. Envelope, snake_case fields, opaque ids, per-endpoint defaults —
    Sessions are rows; only SHA-256 is stored A dump of the sessions table must not be replayable as a login Current. Opaque 256-bit tokens, HttpOnly cookie, CHECK-pinned hash format —
    Two expiries per session A sliding window alone can be slid forever Current. 12h/24h and 30d/90d —
    users.role is the sole authority account_type is user-writable through PATCH /me and cannot be an authorization field Current. Documented in contract §9A.1 —
    Verbatim Markdown as the authoritative artefact A definition must survive a round trip to a .md file on disk unchanged; every column is a cache of it Current. markdown stored byte-for-byte; projections recomputed on every write —
    JS/Go parser parity enforced by replay Two parsers that disagree fail silently in both directions Current. testdata/oracle.json captured from the real frontend module graph; 37 corpus + 132 adversarial cases Regenerate the oracle after any change to src/lib/skills or src/lib/agents
    No YAML library A general parser accepts a larger language than the frontend does; every extra construct is a definition the editor cannot read Current. Hand-ported subset in yaml.go —
    Personal vs organization definitions, two partial unique indexes Shadowing by id is the point, so the id must be unique per tier and never globally Current. ..._personal_key and ..._org_key —
    No skill versioning The frontend has no notion of a skill version and cannot set one; inventing it would create a column forever equal to 1 Current. skill_definitions has no version column, asserted by TestSkillStatusAndNoVersion Would require an explicit architecture decision
    Two definition tables, not one Agents and skills do not share a lifecycle; one table would need a union CHECK permitting nonsense states Current. agent_definitions, skill_definitions —
    permissions: parsed but not enforced Its semantics are not defined yet, and a half-enforced permission model is worse than none Current. Parsed onto Agent.Permissions; no definition_permissions table Define the semantics before storing them
    Runtime execution boundary before any executor Loading, scoping, eligibility and resolution are worth getting right independently of what eventually executes Current. AgentExecutor/SkillExecutor interfaces; UnavailableExecutor refuses Real executors attach here
    Health endpoint says the verdict, not the reasoning An unauthenticated endpoint is a public document, not an operator console Current. One status field; full detail in the server log —
    Firebase Auth — Future direction. The repository contains no reference to Firebase; the current implementation is server-side sessions with argon2id Would replace or front the current credential path
    Google Pub/Sub — Future direction. No reference in the repository —
    GCS-compatible object storage — Future direction. No reference in the repository; infrastructure/README.md names MinIO as a "later" candidate —
    Docker deployment — Future direction. infrastructure/ contains only a README; Dockerfile.api and Dockerfile.owliver are listed there as later work —
    Redis Shared state for rate limiting across instances Future direction. ratelimit.go names Redis or the database as the seam; nothing is wired Needed before multi-instance deployment for the limiter to mean anything
    NATS is not part of the target architecture — Future direction / stated rule. infrastructure/README.md previously listed NATS as a later docker-compose candidate; that table has been corrected. See §17.17. Keep it out
    RAG / pgvector — Future direction. pgvector is not installed in the local database and is named only once, in infrastructure/README.md Introduce when the architecture calls for it

    17. Known Issues & Current Limitations

    Everything below was checked against the current repository. Items the request asked about that could not be verified are marked as such rather than guessed at.

    17.1 Documentation is behind the code — resolved

    Issue as recorded. README.md described the repository as being at an earlier stage than the code is. It stated "Authorization — which roles may do what — is Phase 3D and is not implemented", "Authorization is not implemented. users.role is carried on the identity and consulted nowhere: any signed-in user reaches every endpoint", "38 endpoints" and "55 tests". The same staleness appeared in two source comments: the internal/httpserver/server.go package documentation ("Authorization is NOT here"), and internal/authctx/authctx.go ("It is NOT consulted anywhere in Phase 3C: authentication only").

    Cause. Authorization, the definition tables, the CRUD surface and the runtime all landed after the README was last revised.

    Current behaviour. internal/domain/policy.go, Server.authorize, builder.ownership and internal/httpserver/rbac_test.go (730 lines) all exist and run. 51 routes are registered. 175 test functions run. docs/api-contract.md §9A is current and documents the authorization contract accurately.

    Resolution. README.md was revised against the code: the status paragraph, the layout tree (which was missing cmd/setpassword, internal/auth, internal/authctx, internal/definition and internal/runtime), the route count, the test count and the authorization section. The server.go and authctx.go package comments were corrected to describe the authorization that exists. internal/orgctx/orgctx.go, which still framed itself as pre-authentication, was corrected in the same pass.

    17.2 The definitions endpoints are not in the API contract

    Issue. docs/api-contract.md calls itself "the frozen client contract" and contains zero occurrences of agent-definitions or skill-definitions, yet ten routes are registered.

    Cause. The contract was revised for authorization (§9A) but not for the definitions surface.

    Impact. A frontend engineer working from the contract would not know these endpoints exist, and their request/response shapes, filters and error codes are documented only in Go source and tests.

    Current behaviour. The endpoints work and are covered by 8 test functions.

    Possible future handling. Add a definitions section to the contract.

    17.3 A referenced document does not exist

    Issue. internal/definition/definition.go states that backend-only rejection rules "each one is listed in docs/phase-4d-parser-contract.md", and internal/definition/skill.go refers to "the contract document". docs/ contains only api-contract.md and this file.

    Impact. The two BackendOnly rules are discoverable only by reading ValidateAgent and ValidateSkill.

    17.4 The definitions surface bypasses the policy table

    Issue. The ten definition handlers do not call Server.authorize and the policies map has no entry for agent-definitions or skill-definitions. Role checks are instead written inline in internal/service/definitions.go — the same if !known || role == domain.RoleTalent { return nil, domain.Forbidden() } block appears eight times.

    Cause. Definitions are not domain.Resource values, so the descriptor-and-policy machinery does not reach them.

    Impact. The deny-by-default guarantee — enforced by TestEveryResourceHasAPolicy — does not extend to the newest ten endpoints. A future definition-related endpoint added without its inline check would be reachable by any signed-in user, and no test would notice.

    Current behaviour. Tenancy and ownership are enforced, in the repository predicates, for every definition read and write. The role rule that exists — organization-visibility definitions are writable by admin and employer only — is applied on create, update and delete, and is covered by tests.

    Possible future handling. Bring definitions under the policy table, or add a test that asserts every definition route has an explicit role check.

    17.5 The runtime package is not reachable from the running service

    Issue. No code outside internal/runtime and its own test file references the package. There is no HTTP route, no command, and no service that constructs an Engine.

    Impact. The runtime's loading, scoping, eligibility and resolution logic is exercised only by its tests. Nothing a deployed binary does touches it.

    Current behaviour. As stated in §7: runtime execution is an internal backend boundary and no public runtime execution endpoint is implemented.

    17.6 Cycle detection in dependency resolution is unreachable

    Issue. Loader.ResolveAgentDependencies maintains an inProgress map and returns ErrCircularDependency, but sets inProgress[skillID] = true and back to false within the same loop iteration, and never recurses — skills do not resolve skills. The error is therefore not reachable, and no test covers it.

    Cause. The guard is shaped for a recursive resolver; the resolver is one level deep.

    Impact. A genuine cycle cannot exist in the current model (an agent names skills; skills name nothing), so nothing is currently mis-handled. The code implies a protection that is not active.

    Possible future handling. Either make resolution recursive, at which point the guard becomes live and needs a test, or remove it.

    17.7 Unchecked type assertions in the runtime loader

    Issue. internal/runtime/loader.go lines 67, 72, 129 and 133 perform rec["id"].(string) and rec["visibility"].(string) without the comma-ok form. Every other map access in the codebase, including internal/repo/repo.go:346, uses the defensive form.

    Impact. A nil or non-string value would panic. The panic would be caught by recoverer and become a logged 500 — but only if the runtime were reachable over HTTP, which it is not.

    Current behaviour. The repository projects both columns as ::text, so the assertion holds in practice.

    17.8 CORS does not permit credentials

    Issue. internal/httpserver/cors.go never sets Access-Control-Allow-Credentials: true. Its own comment anticipates this — "so adding credentials later does not require rewriting this" — but the header was not added when cookie authentication landed.

    Impact. A browser at http://localhost:5173 calling the API at http://127.0.0.1:8080 with credentials: 'include' will have the response blocked by the browser, even though the server answered correctly. This affects exactly the cross-origin development setup the CORS middleware exists for. A same-origin deployment is unaffected.

    Current behaviour. Five CORS tests exist and pass; none asserts anything about credentials.

    Possible future handling. Set the header for allowlisted origins, and add a test.

    17.9 Login rate limiting does not survive scale-out

    Covered in §13 under limitations: per-process, in-memory, RemoteAddr-based, pruned rather than capped. The source documents all three limitations itself. Impact: the effective limit multiplies by instance count, resets on restart, and degrades to a global budget behind a reverse proxy. Possible future handling: move the limiter behind shared state before deploying more than one instance.

    17.10 The repository has almost no history

    Issue as recorded. git log reported fatal: your current branch 'main' does not have any commits yet, and every file was untracked.

    Current behaviour. The tree is now committed: a single commit on main (7d12ebe, "first commit") holds the whole repository.

    Remaining impact. One commit is not history. There is still no blame, no incremental recovery point, and no record of when any of the work described in §2 happened. The chronology in this document was reconstructed from migration headers, package documentation and the contract, not from version control, and that remains the only source for it.

    17.11 A minor documentation defect in the source — resolved

    internal/httpserver/api.go — the doc comment describing decodeBody sat immediately above decodeInto, so decodeInto carried two doc comments and decodeBody carried none. The decodeBody comment has been moved to sit above decodeBody.

    17.12 Badge endpoint mismatch — verified

    Issue. The badges table exists and is seeded with 4 records, but there is no endpoint for it.

    Cause. Recorded in contract §2: the frontend's useBadges hook has zero consumers, and every badge the UI renders comes from worker_profiles.earned_badges. The table exists; nothing reads it.

    Impact. A frontend Badge.list() call would receive a 404. README.md states that this call "has 404ed since Phase 2C".

    Current behaviour. badges declares Ops: 0, so no route is registered, and its policy is written out as an explicit empty policy so the resource is deliberately closed rather than merely forgotten.

    Possible future handling. Add the endpoint if a consumer appears, or drop the table.

    17.13 Job Posting 422 — mechanism verified, no specific defect recorded

    Issue as raised. A "Job Posting 422 validation issue".

    What the repository shows. No artifact — comment, test, TODO or contract note — records a specific job-posting defect. Not verified from current repository.

    What can be stated factually. POST /api/v1/job-postings answers 422 in exactly the documented cases, and two of them are easy to hit from a client that was written against the in-browser store:

    1. An unknown body field. Contract §13.5 made unknown fields a 422 rather than a silent drop, deliberately. Any field the frontend sends that has no column produces validation_failed with the field named in details.
    2. An invalid enum value. job_postings has three enum columns — english_required, status and priority — and a value outside the declared set is a 422 naming the permitted values.

    title is the only column marked Required on this resource, so a missing-required 422 would name title specifically.

    Possible future handling. If the symptom is reproducible, the details map in the 422 body names the offending field, which is enough to decide whether the column is absent from the schema (as interview_id, training_outline and score_breakdown once were) or the value is wrong.

    17.14 Shift snapshot / date behaviour — verified

    Not a defect; a documented design constraint. Shift records are generated rather than snapshotted because their dates are anchored to now, and the frontend windows every collection on created_date. A frozen snapshot would read as permanently empty a fortnight later. The consequence, recorded as U1 in the contract, is that attendance and overtime will be empty in any real deployment until a rostering source exists. Re-seeding on a later day prunes stale rows to keep the rolling window convergent.

    17.15 resetDemoData — not present

    The string resetDemoData does not appear anywhere in this repository. Not verified from current repository. It is presumably a frontend concern.

    17.16 Frontend test drift — not verifiable here

    The frontend repository is not part of this checkout. Not verified from current repository.

    17.17 NATS appears in the repository despite the stated direction

    Issue. The direction stated for this document is that NATS is not part of the target architecture. infrastructure/README.md currently lists "Redis, NATS, MinIO" as candidates for a later docker-compose.dev.yml, and README.md mentions NATS among the things that do not exist yet.

    Impact. A reader of infrastructure/README.md would take NATS to be planned.

    Resolution. The rule stands, so the table was amended: infrastructure/README.md no longer lists NATS as a candidate and says explicitly that it is not part of the target architecture. README.md no longer names it either.

    17.18 Deliberate contract behaviours that read as defects

    These are recorded in docs/api-contract.md §12 with reasons and are not bugs: multi-record writes are not transactional (§12.1); collection caps truncate silently (§12.2); all aggregation is client-side (§12.3); search is browser-side (§12.4); email comparison is case-insensitive server-side (§12.5); DELETE on a missing record returns 200 (§12.7); LLM and file-upload integrations remain in the browser (§12.8).


    18. Future Technical Direction

    Nothing in this section is implemented. Each item is labelled with what the repository actually says about it, so the two are never confused.

    Agent Runtime / Owliver

    The intended direction is a Python "Owliver" service alongside the Go API, with a subagent bridge and orchestration, and real execution behind the existing AgentExecutor / SkillExecutor interfaces.

    What the repository says: README.md states "a Python Owliver service to follow" and "No Owliver". internal/runtime/executor.go names "real AI / Owliver / LangGraph executors" as what would replace UnavailableExecutor. infrastructure/README.md lists Dockerfile.owliver as later work. No Python source, no bridge and no orchestration code exists here.

    AI Runtime

    An LLM provider abstraction, with an appropriate provider per workload.

    What the repository says: nothing. No provider name — OpenAI, Anthropic, Gemini, Vertex — appears anywhere in this checkout. Not verified from current repository. The attachment point is the two executor interfaces.

    RAG

    pgvector, embeddings and retrieval over an agent's knowledge corpus.

    What the repository says: pgvector is named exactly once, in infrastructure/README.md, as part of a later dev compose file. It is not installed in the local database (only citext and plpgsql are). Migration 000005 explicitly declines to create an agent_knowledge table because "no corpus exists". Contract §11 lists RAG and embeddings as D7–D10, "none touch v1".

    Tools

    FastMCP and a tool-execution layer.

    What the repository says: nothing. Neither "MCP" nor "FastMCP" appears anywhere. Not verified from current repository. Note that internal/definition/yaml.go states as a design property that "there is no code path from a definition to execution of any kind" — introducing tool execution changes that property deliberately and should be done knowingly.

    Infrastructure

    • Google Pub/Sub — no reference in the repository. Not verified from current repository.
    • GCS-compatible object storage — no reference. infrastructure/README.md names MinIO as a later compose candidate.
    • Docker — infrastructure/ contains only a README, deliberately empty. It lists docker-compose.dev.yml, Dockerfile.api, Dockerfile.owliver and otel-collector.yaml as later work, with the stated reason that adding any of them before its phase would be speculative.
    • Redis — named in ratelimit.go as the seam for shared limiter state, and in infrastructure/README.md as a later compose candidate. Nothing is wired.
    • Observability — otel-collector.yaml is listed as later work. Today there is structured JSON logging via log/slog and nothing else: no metrics, no tracing, no exporter.
    • Production hardening — no TLS termination, secret management, key rotation, backup policy or deployment manifest exists in this repository.

    NATS

    NATS is not part of the target architecture. infrastructure/README.md no longer lists it as a candidate; see §17.17.

    Nearest-term work implied by the code itself

    Not a roadmap, but the things the repository points at:

    • Multi-record transactional endpoints (POST /job-applications/{id}/hire, POST /job-postings/{id}/assignments), which the contract defers because they require touching frontend hooks.
    • Server-side aggregation, which contract §12.3 identifies as the first thing that breaks as data grows — shift_records at 500 rows.
    • Consuming meta.truncated, which exists so §12.2 is fixable without a contract change.
    • Defining the semantics of the permissions: block before storing or enforcing it.
    • Shared state for the login limiter before deploying more than one instance.

    19. What We Have Built So Far

    For a team audience, in plain terms.

    Where we started. A React demo whose data lived in a browser tab. store.js held arrays, seed.js filled them at boot, and every hook called through one file: base44Client.js. Nothing survived a refresh, everyone saw the same data, and no rule existed that the client did not enforce on itself.

    Where we are. A Go service in front of PostgreSQL that the same React app talks to through the same one file.

    What we built What it is practically worth
    A Go API over PostgreSQL Data survives. Two people see the same records. A query that takes 200 ms in the browser over 500 rows takes a millisecond in an index.
    A written contract derived from call sites Nobody had to guess what the frontend needed, and no endpoint exists that nothing calls. The frontend migration cost exactly one file.
    Migrations as the source of truth The schema has a history, a rollback, and a review surface. No table has ever been created by hand.
    Generated resource descriptors Column names, types and enums cannot drift from the database, because they are read out of it.
    A seeder that runs the frontend's own seed The demo dataset in PostgreSQL is the demo dataset the frontend ships, not a transcription of it. Re-seeding is safe and converges.
    Session authentication There is a real identity behind every request, held in an HttpOnly cookie the page cannot read, backed by a row we can revoke.
    Role-based authorization with deny-by-default Who may do what is written down in one table, and a resource nobody has written a rule for is unreachable rather than open.
    Tenant and ownership isolation in SQL A row from another organization, or another worker, is never fetched — so it cannot leak through a count, a total, or a bug in a later loop.
    Agent and Skill authoring in the database Authored definitions moved out of a jsonb blob in user preferences into real tables with an owner, a tenant, a size bound and a query surface.
    A Go parser that provably matches the JS one An author sees the same acceptance, the same rejection and the same wording in the editor and from the API — asserted against the real frontend parser's captured output over 37 shipped definitions and 132 adversarial cases.
    A runtime boundary Loading, tenant scoping, status eligibility and dependency resolution are written and tested, so when execution arrives it plugs into an interface rather than reopening those questions.
    A test suite that runs against a real database 175 test functions, 920 test entries, a disposable migrated-and-seeded database per test process, and ordering guarantees verified to fail when removed.

    What is deliberately not built yet. Real AI execution. Any public runtime endpoint. Server-side aggregation. Transactional multi-record flows. RAG. Tooling. Anything in infrastructure/.


    20. Current Backend Snapshot

    Current Database: PostgreSQL 18.6 locally; one schema (public); five migration pairs, applied version 5, not dirty; 20 application tables plus schema_migrations; 16 enum types; extensions citext and plpgsql; UUID primary keys with a nullable unique legacy_id; timestamptz throughout.

    Current API: Go 1.27, net/http only, base path /api/v1, 51 registered routes (34 resource + 4 /me + 2 auth + 10 definitions + /health). JSON envelope with data and, on collections, meta. Two filter operators, NULLS LAST in both directions with an id tiebreaker, limit/offset with per-endpoint defaults and a 1000 cap, and nine error codes.

    Current Authentication: email + password with argon2id at OWASP parameters; opaque 256-bit session tokens stored only as SHA-256; HttpOnly, SameSite=Lax cookies with Secure outside development; sliding expiry with an absolute ceiling (12h/24h, or 30d/90d with remember-me); a 15-minute expired-session sweeper; failed attempts limited per email and per address; uniform 401s with decoy hashing to resist enumeration. Passwords enter only through cmd/setpassword.

    Current RBAC: three roles from users.role — admin, employer, talent. Deny-by-default policy table keyed by URL path. Role gate in the handler answering 403 before any query. Eight server-owned identity columns. PATCH /me writes only full_name and account_type.

    Current Tenant Model: one organization per user, taken from the user's row and never from the request. org_id predicate on every read and write. courses and learning_paths also match org_id IS NULL, the shared platform library. Talent callers carry a second ownership predicate in the same WHERE clause; a row outside either predicate answers 404.

    Current Agent System: Markdown with YAML frontmatter, stored verbatim in agent_definitions; six projected columns; draft/published/archived plus a monotonic integer version; personal and organization tiers with per-tier uniqueness; full CRUD over five routes.

    Current Skill System: the same shape in skill_definitions, minus version; active/inactive; full CRUD over five routes. ui: and owliver: frontmatter blocks are recorded as deferred rather than validated.

    Current Runtime: an internal boundary. Loads by UUID or definition_id within the caller's tenant and ownership scope, re-parses and re-validates the stored Markdown, applies status eligibility, resolves an agent's skill dependencies with personal-over-organization shadowing and deduplication, and dispatches to an executor interface. The only executor implementation refuses execution. No HTTP route reaches it.

    Current Frontend Integration: the frontend repository is separate and unmodified. The seam is base44Client.js; the contract's acceptance criterion is that swapping the transport inside it changes nothing above it. CORS allowlist defaults to the Vite dev server on both hostnames in development, and to empty elsewhere.

    Current Testing: 175 test functions, 920 test entries executed, no failures and none skipped on a machine with PostgreSQL reachable. gofmt reports no files, go vet reports no diagnostics, go build ./... builds. Database-backed tests use a disposable krow_backend_autotest_<pid> database per test process.

    Current Known Limitations: no AI execution and no public runtime endpoint; login rate limiting is per-process and in-memory; CORS does not permit credentials; definitions endpoints bypass the policy table; the runtime package is unreachable from the running service; the definitions endpoints are absent from the API contract; attendance data is empty without a rostering source; the repository has a single commit and so no usable history.

    Current Development Focus: the most recent work in the tree is the definition system — schema, parser conformance, CRUD — and the runtime boundary that sits on top of it. The open seams the code itself points at are real executors behind the executor interfaces, and a public surface for the runtime if one is wanted.


    21. Critical Engineering Rules

    These are the invariants the current code depends on. Breaking any of them silently is how this codebase would stop being trustworthy.

    1. Do not casually modify migrations. Once a migration has run anywhere other than your own machine it is immutable. Change it and every database that already applied it diverges from every one that has not. Write the next one instead — that is exactly what 000002 and 000003 are.
    2. Do not bypass tenant scoping. org_id belongs in the WHERE clause of every read and write. It is not a filter applied afterwards, and it is not optional on personal rows.
    3. Do not trust a client-provided org_id. It comes from the user's row, which comes from the session row, which comes from a cookie value the client cannot forge without already holding it.
    4. Do not trust a client-provided owner_user_id, or any of the eight server-owned identity columns. They are the columns every ownership rule rests on; a caller who could set them could defeat the rule with the same request it constrains.
    5. Do not bypass parser validation. ValidateAgent / ValidateSkill before ParseAgent / ParseSkill before persistence, on create and on any update that touches markdown.
    6. Do not mutate stored Markdown. The bytes in are the bytes stored. Normalization reads; it never rewrites. Three tests assert this from three directions.
    7. Do not invent Agent or Skill fields. Every field exists because the frontend parser has it. A field the Go parser knows and the JS parser does not is a definition the editor cannot read.
    8. Do not add skill versioning without an explicit architecture decision. The frontend has no notion of it and cannot set one.
    9. Do not bypass RBAC. The role gate runs in the handler before any query, and the row predicate runs in SQL. Both, every time.
    10. Do not expose cross-tenant resource existence. 403 means "your role"; 404 means "not yours or not there". Never a message that distinguishes the two.
    11. Do not introduce NATS.
    12. Do not introduce RAG or pgvector before the architecture calls for it. There is no corpus, and migration 000005 declines to create a table for one.
    13. Do not introduce MCP or tool execution prematurely. The parser's stated property is that no code path leads from a definition to execution of any kind; changing that should be a decision, not a side effect.
    14. Do not introduce Redis prematurely. The limiter names it as the seam; wire it when there is more than one instance, not before.
    15. Do not modify the frontend during backend-only work unless it is explicitly required. The whole transport migration is affordable because exactly one frontend file changes.
    16. Do not fake AI execution. UnavailableExecutor refuses honestly. A stub that returns plausible text would be worse than one that returns an error.
    17. Do not silently change API contracts. docs/api-contract.md is the client's contract. Divergences from it have been reported and written down every time — §13.4 and §13.5 exist for that reason.
    18. Do not let the descriptors and the policy table drift. resources_gen.go is generated; policy.go is hand-written. That separation is why regenerating one cannot silently drop a rule in the other, and TestEveryResourceHasAPolicy is what keeps it honest.
    19. Do not put reconnaissance in a public response. /health answers the verdict; the reasoning goes to the log.
    20. Do not log a password, a session token, or a password hash. Nothing in internal/auth logs at all, by design.

    22. Final Team Summary

    What the team should know

    • We started with a browser-only demo. Data lived in arrays in a tab; there was no identity, no tenancy, and no rule the client did not enforce on itself.
    • Go and PostgreSQL were introduced to move persistence, identity and enforcement to a server without rewriting the frontend. The acceptance criterion was written down first: swapping the transport inside one frontend file must change nothing above it. That held.
    • The architecture is a plain layered service — handler, service, repository, pgx, PostgreSQL — with no ORM and no web framework. Three direct Go dependencies.
    • The schema lives in migrations/ and nowhere else. Five migration pairs; nothing in the Go code issues DDL. Resource descriptors are generated from the live schema so they cannot drift.
    • The API contract was derived from frontend call sites, not designed. An operation with no call site gets no endpoint — which is why badges has none and DELETE /job-postings/{id} answers 405.
    • Authentication is server-side sessions: argon2id passwords, opaque 256-bit tokens, only SHA-256 stored, HttpOnly cookies, sliding expiry with an absolute ceiling, and login rate limiting per email and per address.
    • Authorization is deny-by-default. Three roles from users.role. The role gate answers 403 before any query; tenancy and ownership are SQL predicates, so a row outside them answers 404 and cannot be distinguished from one that does not exist.
    • Agent and Skill definitions moved out of a jsonb preferences blob into real tables with an owner, a tenant, a size bound and a query surface. The Markdown is the authoritative artefact; every column is derived from it and recomputed on write.
    • The Go parser provably matches the JavaScript one — replayed against the real frontend parser's captured output over 37 shipped definitions and 132 adversarial cases, with mutation checks so the tests cannot pass against a broken parser.
    • The runtime boundary exists; execution does not. Loading, tenant scoping, status eligibility and dependency resolution with personal-over-organization shadowing are written and tested. The only executor refuses. No HTTP route reaches the runtime at all.
    • Verification is real: 175 test functions, 920 test entries, against a disposable migrated-and-seeded PostgreSQL database per test process. gofmt, go vet and go build are clean.
    • The main current limitations are: no AI execution and no runtime endpoint; login rate limiting that does not survive scale-out; CORS that does not yet permit credentials; ten definition endpoints that sit outside the policy table and outside the written contract; and a repository with a single commit and so no usable history.
    • The direction is real execution behind the existing executor interfaces (Owliver / an LLM provider abstraction), then retrieval, then tooling, then deployment infrastructure — none of which exists here today. NATS is not part of the target architecture, and the documentation no longer implies otherwise.
    • The rules that must not be broken are in §21. The two that matter most in daily work: never trust a client-supplied identity or tenant, and never rewrite stored Markdown.

    Generated from the repository state on 2026-08-24. Every count, version and behaviour above was read from the current checkout or from read-only queries against the local development database. Nothing in the repository was modified to produce this document.

    Revised on 2026-08-24 by a structure-and-dead-code cleanup pass, which changed no schema, no migration, no database row and no runtime behaviour. What it did change is recorded in §17.1, §17.10, §17.11 and §17.17: documentation and source comments that had fallen behind the code were corrected. The counts above were re-verified against the checkout and are unchanged.