Preserve CLAUDE.md and add a handover document
Neither survived a machine change. CLAUDE.md sat in the directory ABOVE both repositories, which is not a git repository at all, so the governing document for the project existed on exactly one laptop. It is now in this repository; place a copy at the parent level on a new machine, where it covers both. docs/handover.md records what CLAUDE.md does not: what was decided and why, what is deployed and how to verify it, and the conventions that produce confident wrong numbers rather than errors — a score of 0 meaning "not rated", screened_at being vestigial, shift data anchored to today. Written because Claude Code's own memory is per-machine and keyed to the absolute path of the checkout: it does not sync, and a different path on a new machine reads a different folder. A file in the repository travels with the code and is useful to a person besides. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
This commit is contained in:
270
CLAUDE.md
Normal file
270
CLAUDE.md
Normal file
@@ -0,0 +1,270 @@
|
|||||||
|
# CLAUDE.md
|
||||||
|
|
||||||
|
## Project instructions for Claude Code. Read this fully before writing any code in this repo.
|
||||||
|
|
||||||
|
## 1. What this project is
|
||||||
|
|
||||||
|
A multi-tenant **agent platform**: infrastructure that lets agents be _defined_, _permissioned_, _executed_, and _evaluated_. It is not a chatbot and it is not a single agent.
|
||||||
|
The platform provides six layers. Everything you build belongs to exactly one:
|
||||||
|
| Layer | Owns | Directory |
|
||||||
|
|---|---|---|
|
||||||
|
| Surfaces | how humans invoke agents (chat, mentions, triggers, API) | `src/surfaces/` |
|
||||||
|
| Orchestration runtime | the agent loop, delegation, streaming, budgets | `src/runtime/` |
|
||||||
|
| Agent registry | agent specs, versioning, sharing, resolution | `src/registry/` |
|
||||||
|
| Tool layer | MCP servers, tool schemas, confirmation gates | `src/tools/` |
|
||||||
|
| Knowledge layer | ingest, ACL-tagged chunks, hybrid retrieval | `src/knowledge/` |
|
||||||
|
| Model gateway | model routing, budgets, fallback, token accounting | `src/gateway/` |
|
||||||
|
If a change touches more than two layers, stop and describe the plan before writing code.
|
||||||
|
**Fill this in before starting:**
|
||||||
|
|
||||||
|
```
|
||||||
|
PROJECT_NAME: Krow
|
||||||
|
DOMAIN: Hospitality and event workforce operations — staffing open shifts,
|
||||||
|
screening and hiring candidates, tracking attendance and overtime,
|
||||||
|
and answering from the organisation's own policy documents.
|
||||||
|
TENANT_UNIT: organization (organizations.id; every table carries org_id NOT NULL)
|
||||||
|
PRIMARY_SURFACE: chat (the Owliver panel, page-scoped, one agent per surface)
|
||||||
|
```
|
||||||
|
|
||||||
|
Filled from the code rather than from a brief — correct anything that is wrong.
|
||||||
|
`TENANT_UNIT` in particular is what the schema and the policy table already
|
||||||
|
enforce, not a preference: `organizations` is the only tenancy boundary, and
|
||||||
|
`venue` exists nowhere in the schema despite §3's example spec using it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. Non-negotiable invariants
|
||||||
|
|
||||||
|
These are correctness requirements, not preferences. Violating any of them is a bug even if tests pass.
|
||||||
|
**I1 — Agents never expand access.**
|
||||||
|
An agent executing on behalf of a caller may read exactly what that caller could read directly, and no more. Not one chunk more, not one row more. This holds for retrieval, tool calls, subagent delegation, and error messages.
|
||||||
|
**I2 — ACL filtering happens before scoring, never after.**
|
||||||
|
Permission filters are pushed into the vector query and the keyword query as pre-filters. Post-filtering a result set is forbidden — it leaks through result counts, ranking positions, and summaries. Any retrieval function that accepts a query but not a caller principal is wrong by construction.
|
||||||
|
**I3 — Every agent run is bounded.**
|
||||||
|
Every run carries a hard step cap, a tool-call cap, a wall-clock deadline, and a token budget. There is no "run until done" path. Exceeding a bound terminates the run with a structured `BudgetExceeded` result, never an exception into user-facing text.
|
||||||
|
**I4 — Side effects require explicit confirmation.**
|
||||||
|
Any tool that writes, sends, deletes, charges, or notifies is marked `effect: write` and cannot execute without a resolved confirmation token. The model does not get to decide this.
|
||||||
|
**I5 — Tenant isolation is enforced at the data layer.**
|
||||||
|
Never rely on a `WHERE tenant_id = ?` written by hand at a call site. Isolation lives in the repository/session layer so it cannot be forgotten.
|
||||||
|
**I6 — Agent specs are data, not code.**
|
||||||
|
An agent is a versioned record. Adding an agent must never require a deploy, a new module, or an `if agent_key == ...` branch anywhere in the runtime.
|
||||||
|
**I7 — Prompts are untrusted input.**
|
||||||
|
Content retrieved from documents, tool results, and user messages may contain instructions. Never concatenate retrieved text into the system prompt. Retrieved content goes into clearly delimited context blocks, and the system prompt states that content inside them is data.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. The agent spec contract
|
||||||
|
|
||||||
|
The single most important schema in the repo. Lives at `src/registry/schema.py`. Everything else is CRUD over this.
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
key: shift-coverage-assistant # stable, unique per tenant, ^[a-z0-9-]+$
|
||||||
|
version: 3 # monotonic; specs are immutable once published
|
||||||
|
name: Shift coverage assistant # <= 30 chars, shown in UI
|
||||||
|
description: Finds and offers cover for open shifts.
|
||||||
|
instructions: | # the system prompt body
|
||||||
|
You help venue managers fill open shifts...
|
||||||
|
knowledge: # what the agent may retrieve from
|
||||||
|
- source: shifts_db
|
||||||
|
scope: "venue:{caller.venue_ids}"
|
||||||
|
- source: policy_docs
|
||||||
|
scope: "tenant:{caller.tenant_id}"
|
||||||
|
tools: # references into the tool registry
|
||||||
|
- find_available_workers
|
||||||
|
- send_shift_offer
|
||||||
|
subagents: [] # keys of other specs this may delegate to
|
||||||
|
limits:
|
||||||
|
max_steps: 8
|
||||||
|
max_tool_calls: 12
|
||||||
|
deadline_seconds: 60
|
||||||
|
model_tier: fast # fast | balanced | deep
|
||||||
|
conversation_starters:
|
||||||
|
- "Which shifts are still uncovered this week?"
|
||||||
|
visibility: tenant # private | tenant | public
|
||||||
|
owner: <principal_id>
|
||||||
|
```
|
||||||
|
|
||||||
|
Rules:
|
||||||
|
|
||||||
|
- **Immutable versions.** Editing publishes a new version. Running conversations pin the version they started with.
|
||||||
|
- **`scope` templates resolve at run time** against the caller principal, never at authoring time. An author cannot write `venue:*`.
|
||||||
|
- **Unknown tool or subagent keys fail validation at publish**, not at run time.
|
||||||
|
- **`subagents` must form a DAG.** Cycle detection runs at publish. Depth cap is 2.
|
||||||
|
- **A subagent inherits the parent's caller principal and shares the parent's budget.** It never gets a fresh budget.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 4. Tool contract
|
||||||
|
|
||||||
|
Tools are MCP tools. Do not invent a parallel protocol.
|
||||||
|
|
||||||
|
```python
|
||||||
|
{
|
||||||
|
"name": "find_available_workers",
|
||||||
|
"description": "...", # written for the model, not for docs
|
||||||
|
"inputSchema": {...}, # JSON Schema, all fields described
|
||||||
|
"effect": "read", # read | write
|
||||||
|
"requires_confirmation": False, # forced True when effect == "write"
|
||||||
|
"max_result_bytes": 262_144,
|
||||||
|
}
|
||||||
|
```
|
||||||
|
|
||||||
|
Implementation rules:
|
||||||
|
|
||||||
|
- Every handler signature is `handler(inputs, ctx)` where `ctx` carries the caller principal, tenant, run id, and remaining budget. A handler that ignores `ctx` for authorization is wrong.
|
||||||
|
- Handlers return structured data, not prose. Formatting is the model's job.
|
||||||
|
- Truncate at `max_result_bytes` and set a `truncated: true` flag. Never silently drop.
|
||||||
|
- Tool errors return `{"error": {...}}` — they do not raise. The runtime decides whether the model sees the error and retries.
|
||||||
|
- A tool description that requires the model to guess an ID it has not been given is a design bug. Add a lookup tool instead.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 5. Retrieval rules
|
||||||
|
|
||||||
|
- Hybrid: dense + BM25, fused with RRF. Do not replace this with dense-only for convenience.
|
||||||
|
- Every chunk row carries `tenant_id` and an `acl` field at write time. Chunks without ACL metadata are rejected at ingest.
|
||||||
|
- The retrieval entry point is `retrieve(query, principal, scopes, k)`. There is no overload without `principal`.
|
||||||
|
- Retrieved chunks flow to the model with source ids so the response can cite. Responses that assert facts without a retrievable citation must be marked as inference, not grounded fact — keep the two visually and structurally separate in the output payload.
|
||||||
|
- Reindex is required whenever ACL derivation logic changes. Note it in the PR.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 6. Runtime rules
|
||||||
|
|
||||||
|
The agent loop lives in `src/runtime/loop.py`. It is the highest-risk file in the repo.
|
||||||
|
|
||||||
|
- Single loop, spec-driven. No per-agent branching.
|
||||||
|
- Decrement budgets **before** dispatch, not after, so a hung tool cannot overrun.
|
||||||
|
- Stream partial assistant text as it arrives; buffer tool calls until complete.
|
||||||
|
- Termination reasons are an enum: `Completed | BudgetExceeded | Deadline | ConfirmationPending | ToolFailure | Refused`. Every run ends with exactly one.
|
||||||
|
- Persist a full trajectory per run: every message, tool call, tool result, and budget snapshot. This is what makes debugging and evals possible — it is not optional telemetry.
|
||||||
|
- Delegation is a tool call from the parent's perspective. Subagent runs get their own trajectory, linked by `parent_run_id`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 7. How to add a new agent
|
||||||
|
|
||||||
|
Adding an agent is a data change. If you find yourself editing runtime code, you have found a missing platform capability — surface that instead of special-casing.
|
||||||
|
|
||||||
|
1. Write the spec YAML in `agents/<key>.yaml`.
|
||||||
|
2. Confirm every referenced tool exists. If one is missing, build the tool first (§8).
|
||||||
|
3. Confirm every knowledge source exists and is ACL-tagged.
|
||||||
|
4. Run `make validate-agent KEY=<key>` — checks schema, tool refs, scope templates, subagent DAG.
|
||||||
|
5. Write at least 5 eval cases in `evals/<key>.yaml` (§9). This is required, not optional.
|
||||||
|
6. Run `make eval KEY=<key>` and record the baseline in the PR description.
|
||||||
|
7. Publish: `make publish-agent KEY=<key>` — assigns the next version number.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 8. How to add a new tool
|
||||||
|
|
||||||
|
1. Define the schema in `src/tools/<domain>/schema.py`.
|
||||||
|
2. Implement `handler(inputs, ctx)` in the same package. Authorize using `ctx.principal` on the first line of the handler body.
|
||||||
|
3. If `effect == "write"`, add a confirmation payload renderer describing exactly what will happen in plain language.
|
||||||
|
4. Unit test authorization first: a caller without rights must get a denial, and the denial must not reveal the existence of the resource.
|
||||||
|
5. Register in `src/tools/registry.py`.
|
||||||
|
6. Cap: 20 tools per agent spec. If an agent needs more, it should be split into a parent with subagents.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 9. Evals are part of the definition of done
|
||||||
|
|
||||||
|
No agent ships without evals. No change to the loop, retrieval, or prompt assembly merges without running the full suite.
|
||||||
|
Each eval case:
|
||||||
|
|
||||||
|
```yaml
|
||||||
|
- id: uncovered-shifts-basic
|
||||||
|
principal: fixtures/manager_two_venues.json
|
||||||
|
input: "Which shifts are uncovered this week?"
|
||||||
|
expect:
|
||||||
|
termination: Completed
|
||||||
|
tools_called: [find_open_shifts]
|
||||||
|
must_mention: ["Friday evening"]
|
||||||
|
must_not_leak: ["venue_9"] # data outside the principal's scope
|
||||||
|
max_steps: 4
|
||||||
|
```
|
||||||
|
|
||||||
|
## `must_not_leak` is mandatory on every case. Every eval doubles as a permission test.
|
||||||
|
|
||||||
|
## 10. Conventions
|
||||||
|
|
||||||
|
- Python 3.11+, FastAPI, async throughout. SQLAlchemy 2.0 style.
|
||||||
|
- Type hints on every public function. `mypy --strict` on `src/registry/` and `src/runtime/`.
|
||||||
|
- Errors: structured exception types with a `code`, never bare strings. User-facing text is derived at the surface layer, not raised from the core.
|
||||||
|
- Logging: structured JSON, always include `run_id`, `tenant_id`, `agent_key`, `agent_version`. Never log message content or retrieved chunks at INFO — that is a data leak into your log store. DEBUG only, behind a per-tenant flag.
|
||||||
|
- Config via environment, validated once at startup into a frozen settings object. No `os.getenv` at call sites.
|
||||||
|
- Migrations: Alembic, one per PR, reversible.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 11. Build order
|
||||||
|
|
||||||
|
Do not build ahead of the current phase. Each phase must be working before the next starts.
|
||||||
|
|
||||||
|
- **Phase 1 — Runtime skeleton.** Two or three hardcoded YAML specs loaded from disk. Loop, budgets, streaming, trajectory persistence. No database registry, no UI.
|
||||||
|
- **Phase 2 — Tools + knowledge.** MCP tool layer, ACL-tagged ingest, permission-aware hybrid retrieval. Evals harness alongside.
|
||||||
|
- **Phase 3 — Registry.** Specs move to the database. Versioning, publish flow, resolution by key + tenant. Still no builder UI.
|
||||||
|
- **Phase 4 — Surfaces.** Chat panel, invocation from the product, webhooks.
|
||||||
|
- **Phase 5 — Authoring UI.** Only once the spec schema has been stable for a meaningful stretch. The builder is a form generator over §3 — if it needs to be more than that, the schema is wrong.
|
||||||
|
|
||||||
|
Current phase: **Phase 4 — Surfaces.**
|
||||||
|
|
||||||
|
Phases 1, 2 and 3 are complete and verified against a live model. What remains
|
||||||
|
in Phase 3 is a publish *workflow* — approval, staged rollout — which §12 says
|
||||||
|
depends on the curated-versus-self-serve decision and is not settled.
|
||||||
|
|
||||||
|
| Layer | State |
|
||||||
|
|---|---|
|
||||||
|
| Surfaces | `POST /api/v1/agents/{id}/runs` (streams over SSE on `Accept: text/event-stream`), `GET /api/v1/runs/{id}`; the chat panel is the only answering path — the browser simulator is deleted |
|
||||||
|
| Orchestration | spec-driven loop, four bounds claimed before dispatch, six terminations, trajectories in `agent_runs` |
|
||||||
|
| Registry | 9 agents + 23 skills as rows; published versions immutable (append-only, trigger-enforced); runs pin the version they started with |
|
||||||
|
| Tools | 17, one of which writes, behind a bound single-use confirmation |
|
||||||
|
| Knowledge | ACL-tagged ingest, hybrid dense + BM25 fused with RRF, pre-filtered |
|
||||||
|
| Gateway | tier → model + effort, token accounting, refusal as an outcome |
|
||||||
|
|
||||||
|
**Deviations from this document, all deliberate and all flagged in code:**
|
||||||
|
|
||||||
|
- §3 names the retrieval block `knowledge:`. The shipped product already uses
|
||||||
|
that key for an author's free-text notes, so retrieval corpora are `sources:`.
|
||||||
|
Two meanings under one key would be resolved wrongly by whichever parser ran
|
||||||
|
second, silently. See `runtime.Agent.KnowledgeSources`.
|
||||||
|
- §5 asks for BM25. Postgres does not ship it; the keyword half is
|
||||||
|
`ts_rank_cd`, cover-density ranking. Different function, same job.
|
||||||
|
- Vectors are `real[]` with a dot-product function rather than pgvector, which
|
||||||
|
is not installed. Exact search, no ANN index, bounded by the ACL pre-filter.
|
||||||
|
The upgrade is a column type change and no logic change.
|
||||||
|
- Dense retrieval runs on a deterministic stand-in embedder until a Voyage key
|
||||||
|
exists. It is **not semantic** and refuses to run in production.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 12. Open decisions
|
||||||
|
|
||||||
|
Do not resolve these unilaterally. Flag them and ask.
|
||||||
|
|
||||||
|
- **Who authors agents?** Curated (the team ships specs) vs. self-serve (tenants author their own). Self-serve requires prompt-injection hardening at the authoring boundary, per-tenant cost caps, an approval workflow, and a sandbox — roughly 3× the platform. Current assumption: **curated**, with the registry designed so self-serve is additive later.
|
||||||
|
- **Model hosting.** Self-hosted vs. API vs. mixed by tier.
|
||||||
|
- **Confirmation UX.** Inline in-chat vs. an approval queue.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 13. Anti-patterns
|
||||||
|
|
||||||
|
Things that look like progress and are not:
|
||||||
|
|
||||||
|
- Filtering retrieval results after scoring "because it's simpler."
|
||||||
|
- A `special_cases.py` in the runtime.
|
||||||
|
- Passing the tenant id as a plain function argument through five layers.
|
||||||
|
- Letting the model choose whether a write needs confirmation.
|
||||||
|
- Fresh budgets for subagents.
|
||||||
|
- Concatenating retrieved document text into the system prompt.
|
||||||
|
- Building the authoring UI before the spec schema is stable.
|
||||||
|
- Adding an agent without evals "for now."
|
||||||
|
- Swallowing a tool error and letting the model narrate around it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 14. When stuck
|
||||||
|
|
||||||
|
If a requirement seems to demand breaking an invariant in §2, the requirement is wrong or the platform is missing a capability. Say which, and propose the platform change. Do not work around the invariant locally.
|
||||||
|
Show less
|
||||||
171
docs/handover.md
Normal file
171
docs/handover.md
Normal file
@@ -0,0 +1,171 @@
|
|||||||
|
# Handover
|
||||||
|
|
||||||
|
Written 2026-08-28, when the machine this was built on was retired.
|
||||||
|
|
||||||
|
Everything Claude Code "remembers" lives in `~/.claude/projects/<mangled-path>/`
|
||||||
|
on one machine, keyed to the absolute path of the checkout. It does not sync,
|
||||||
|
and a different path on a new machine reads a different folder. So the durable
|
||||||
|
record is this file, in the repository, where git carries it and any path works.
|
||||||
|
|
||||||
|
Read `CLAUDE.md` first — it is the governing document. This file is what it does
|
||||||
|
not say: what was decided, what is deployed, and which parts bite.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Where things stand
|
||||||
|
|
||||||
|
**Deployed.** Backend at `https://mcp.krowforce.com`, frontend at
|
||||||
|
`https://platform.krowforce.com`, Kubernetes statefulset `krow` in namespace
|
||||||
|
`krow`, pods named `krow-1` and `krow-2` (they start at 1, not 0).
|
||||||
|
|
||||||
|
ssh root@<host> -p 4422 "kubectl -n krow rollout restart statefulset/krow && \
|
||||||
|
kubectl -n krow rollout status statefulset/krow --timeout=180s"
|
||||||
|
|
||||||
|
**Verify a deployment** — 55 checks including a real agent run:
|
||||||
|
|
||||||
|
KROW_EMAIL=... KROW_PASSWORD=... make verify-deploy BASE=https://mcp.krowforce.com
|
||||||
|
|
||||||
|
Auth runs BEFORE routing, so an unauthenticated probe answers 401 for every
|
||||||
|
path including ones that do not exist. `curl` cannot tell a missing endpoint
|
||||||
|
from a guarded one; only an authenticated check can.
|
||||||
|
|
||||||
|
**Two runtime steps a deploy does not do**, both easy to forget because the API
|
||||||
|
looks healthy without them:
|
||||||
|
|
||||||
|
kubectl -n krow exec krow-1 -- importagents --org <slug> # publishes agents/ and skills/
|
||||||
|
kubectl -n krow exec krow-1 -- ingest --org <slug> # ingests knowledge/
|
||||||
|
|
||||||
|
Without the first, every Owliver question answers 404. Without the second,
|
||||||
|
retrieval finds nothing. The org slug is `krow-dev` — a hardcoded constant
|
||||||
|
(`internal/orgctx.DevOrgSlug`), not configuration.
|
||||||
|
|
||||||
|
**The endpoint count is a signal.** `GET /api/v1/version` reports it. Agent run
|
||||||
|
routes are not registered without a model credential, so 56 means no
|
||||||
|
`ANTHROPIC_API_KEY` and 58 means there is one. A keyless deployment boots
|
||||||
|
cleanly under `APP_ENV=staging` and refuses under `production`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Decisions taken, so they are not relitigated
|
||||||
|
|
||||||
|
**§12, who authors agents: self-serve, split by visibility.** `personal` agents
|
||||||
|
are authored in the UI, POST to `/api/v1/agent-definitions`, and are runnable
|
||||||
|
immediately; evals are not required for them. `organization` agents stay as
|
||||||
|
files published by `importagents` on deploy — that deploy step *is* the approval
|
||||||
|
workflow, and §9's eval requirement still applies. §12 warned self-serve needs
|
||||||
|
"3× the platform"; it does not here, because I1 means an agent runs as its
|
||||||
|
caller and cannot exceed their access, I4 means writes still need a human, and
|
||||||
|
the spec format has no `limits` block so budgets cannot be raised by an author.
|
||||||
|
|
||||||
|
**A rejected candidate counts as screened.** It ranks with `ai_screened`:
|
||||||
|
rejection overwrites the stage it came from, so shortlisted can never be
|
||||||
|
claimed. **An assigned candidate counts as hired**, matching the backend's
|
||||||
|
existing `status IN ('hired','assigned')`.
|
||||||
|
|
||||||
|
**Seeded profile scores are the formula's output**, not hand-authored narrative.
|
||||||
|
Recalculating a seeded profile is a no-op, and a skill-check enforces it.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Conventions the schema actively contradicts
|
||||||
|
|
||||||
|
These are the ones that produce confident, wrong numbers rather than an error.
|
||||||
|
|
||||||
|
**`job_applications.screened_at` is vestigial.** Nothing writes it. "Screened"
|
||||||
|
means `status <> 'applied'`, in about eight places in the frontend. Reading the
|
||||||
|
column reported 1 screened of 24 where the truth was 14.
|
||||||
|
|
||||||
|
**A score of `0` means "not rated", never "rated zero".** Every score column is
|
||||||
|
NOT NULL, so there is no null to distinguish it — that is the trap. Applies to
|
||||||
|
`ai_score`, `client_rating`, `krow_score`, `reliability_score`,
|
||||||
|
`attendance_score`, `performance_score`, `experience_years`. Aggregates need
|
||||||
|
`FILTER (WHERE col > 0)` and a stated basis count. Counting zeros once reported
|
||||||
|
"16 weak candidates averaging 28" for a pool that was 1 weak averaging 76.
|
||||||
|
|
||||||
|
Anchors to check against: applications are **9 scored averaging 76**; workers
|
||||||
|
are **5 rated of 9, client rating 4.70**.
|
||||||
|
|
||||||
|
**Genuine zeros, do not filter these:** `overtime_hours`, `minutes_late`, `xp`,
|
||||||
|
`profile_completion`, and `actual_hours` (0 only on absent/no_show shifts).
|
||||||
|
|
||||||
|
**`absent` and `no_show` are both missed shifts, but only one is a no-show.**
|
||||||
|
`attendance.js` is canonical.
|
||||||
|
|
||||||
|
**`attendance_score` defaults to 100 for display and must never be a scoring
|
||||||
|
input.** A profile with no evidence otherwise scores 12 and leaves the "not yet
|
||||||
|
scored" band.
|
||||||
|
|
||||||
|
**The stage ladder lives once**, in `krow-demo/src/lib/hiringRecords.js`. It was
|
||||||
|
three byte-identical private copies, all missing `rejected` and `assigned`,
|
||||||
|
which `indexOf` scored -1 and dropped from every bucket including `applied`.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Things that will waste your afternoon
|
||||||
|
|
||||||
|
**A space in the checkout path breaks path derivation.** This repo lives under
|
||||||
|
`Krow Project /`, and it has bitten three times: an unquoted `$(CURDIR)` in the
|
||||||
|
Makefile, `` `file://${process.argv[1]}` `` in a script guard, and
|
||||||
|
`new URL(...).pathname` in `scripts/oracle.mjs` (use `fileURLToPath`). Any new
|
||||||
|
path derivation is guilty until tested there.
|
||||||
|
|
||||||
|
**`seed.json` is generated from `krow-demo/src/api/seed.js`.** Never edit it.
|
||||||
|
`npm run seed:fixture` writes it, `npm run seed:check` verifies, and the
|
||||||
|
skill-check compares byte-for-byte. `ShiftRecord` is excluded on purpose: the Go
|
||||||
|
seeder generates it against *now*.
|
||||||
|
|
||||||
|
**Shift data is anchored to today**, so anything asserting against it is
|
||||||
|
date-dependent unless the anchor is pinned. `buildShiftsAt(anchor)` exists for
|
||||||
|
that. One detector had a Friday-and-Saturday blind spot for exactly this reason.
|
||||||
|
|
||||||
|
**Database tests skip when PostgreSQL is unreachable** (`testutil` calls
|
||||||
|
`t.Skipf`). `go test` then exits 0 having run almost nothing. CI has a guard
|
||||||
|
that fails on any skip other than `TestLive*`; keep it.
|
||||||
|
|
||||||
|
**The eval suites use a scripted model.** They prove the permission boundary,
|
||||||
|
not answer quality. `make eval-live` uses the real model and costs tokens; it
|
||||||
|
has never been run.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Still outstanding
|
||||||
|
|
||||||
|
- The knowledge corpus is 8 documents locally; production still has the 2 seeded
|
||||||
|
ones until `ingest` runs there.
|
||||||
|
- `ANTHROPIC_API_KEY` was pasted into a chat transcript and is live in a
|
||||||
|
Kubernetes Secret. Rotate it.
|
||||||
|
- Deployments report `version=dev`: the image is built without
|
||||||
|
`--build-arg VERSION`. `make docker-build` passes it.
|
||||||
|
- `APP_ENV=staging` on the deployment, so the production config guards are off.
|
||||||
|
- Skills are still stored in `user_preferences`; agents were moved to the
|
||||||
|
registry and skills were deliberately left for their own pass.
|
||||||
|
- `definition_versions` is empty — immutable versioning is built, trigger-proven,
|
||||||
|
and nothing has gone through it because `importagents` republishes in place.
|
||||||
|
- CI tests but does not deploy. The README's claim that migrations are "run by
|
||||||
|
CI against the target database" is still aspirational.
|
||||||
|
- The fixture-drift CI jobs need `FRONTEND_REPO_TOKEN` to see the sibling repo,
|
||||||
|
and fail rather than pass quietly without it.
|
||||||
|
- The remote is Gitea. These are GitHub Actions workflows; they do nothing until
|
||||||
|
a compatible runner exists.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Setting up a new machine
|
||||||
|
|
||||||
|
git clone <backend> krow-backend && git clone <frontend> krow-demo
|
||||||
|
cp krow-backend/CLAUDE.md ./claude.md # the governing doc lives above both repos
|
||||||
|
|
||||||
|
Needs: Go (see `go-api/go.mod`), Node 20, PostgreSQL, Docker, and Ollama with
|
||||||
|
`nomic-embed-text` if you want semantic retrieval locally. Then:
|
||||||
|
|
||||||
|
cd krow-backend && cp .env.example .env # fill it in; .env is gitignored
|
||||||
|
make migrate-up && make seed
|
||||||
|
make import-agents ORG=krow-dev
|
||||||
|
make ingest ORG=krow-dev
|
||||||
|
go run ./go-api/cmd/setpassword -email demo@krow.app
|
||||||
|
|
||||||
|
cd ../krow-demo && npm ci && cp .env.example .env
|
||||||
|
# VITE_AGENT_API=/api/v1 for local dev (vite proxies it);
|
||||||
|
# production passes an absolute URL as a Docker build arg instead.
|
||||||
|
|
||||||
|
`.env` files are not in git and must be carried across by hand.
|
||||||
Reference in New Issue
Block a user