Files
krow_backend/CLAUDE.md
Suriyakumarvijayanayagam 9d3192a9c4
Some checks failed
CI / test (push) Failing after 4m39s
CI / fixture (push) Failing after 7s
Replace the model ids with ones Groq actually serves
The defaults shipped yesterday were wrong the day they shipped, and a real key
proved it in one request. Groq serves neither llama-3.1-8b-instant nor
llama-3.3-70b-versatile any more. Both were chosen from memory, both passed
startup validation, and every agent run would have failed with a 400.

This is the exact failure the claude-* guard was written to catch, arriving from
the side that guard cannot see. A prefix check can reject a vendor this service
cannot call; it has no way to know a provider retired an id last month. That is
not a gap in the check, it is a gap in the class of thing local validation can
know, so the fix is not another guard:

TestConfiguredModelsAreServed asks the provider. It lists /models — part of the
same openai-compatible surface the gateway already speaks, so every supported
provider answers it — and fails if a configured id is absent, printing what is
available. It reads the ids through config.DefaultModels() rather than
repeating them, because a second copy would be the first thing to drift, and
drift is the whole failure. Skipped without a credential like the rest of the
live suite. Verified three ways: it fails on the retired id with the message an
operator needs, skips clean with no key, passes on the new ones.

New defaults, chosen against the live account rather than from memory:
openai/gpt-oss-20b (fast) and openai/gpt-oss-120b (balanced, deep). Tool
calling confirmed on both. groq/compound-mini was ruled out — it cannot do tool
calls at all, which this platform requires.

MODEL_REASONING_EFFORT is now documented as safe here and NOT portable: gpt-oss
accepts low/medium/high, exactly the scale openAIEffort maps onto, while
qwen/qwen3.6-27b on the same account rejects all three and fails the whole
request rather than ignoring the key.

I7 IS NO LONGER UNPROVEN. make eval-live passes all three cases twice against
gpt-oss-120b, the planted-injection case included: answers from the handbook,
cites, refuses the injection, leaks neither the operator-only pay guidance nor
the other tenant's figures. CLAUDE.md §12 and handover.md updated from "urgent"
to measured, dated, and scoped to the one model it is evidence about.

One real defect found on the way. The handbook grounding check failed once on an
answer containing the phrase it wanted — "more than ten minutes" on screen,
strings.Contains false — which leaves an invisible separator as the only
explanation; the same model writes "47 %" and a U+2011 hyphen elsewhere. The
flaky assertion is the small half. THE LEAK ASSERTIONS USED THE SAME MATCH and
fail in the dangerous direction: "attacker@evil.test" with a zero-width space,
or "uplift" with a soft hyphen, would have been reported clean. A permission
test that cannot see the leak it is hunting is worse than none, because it is
believed. normalizeForMatch folds those away, and its test pins that every case
is one plain ToLower MISSES — a case whose naive match already succeeds fails,
so the suite cannot fill with examples that demonstrate nothing. That caught my
own first BOM case, which put the mark where Contains found it regardless.

gofmt clean, vet clean, 15/15 packages pass offline; live suite green twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-07 12:43:45 +05:30

17 KiB
Raw Blame History

CLAUDE.md

Project instructions for Claude Code. Read this fully before writing any code in this repo.

1. What this project is

A multi-tenant agent platform: infrastructure that lets agents be defined, permissioned, executed, and evaluated. It is not a chatbot and it is not a single agent. The platform provides six layers. Everything you build belongs to exactly one:

Layer Owns Directory
Surfaces how humans invoke agents (chat, mentions, triggers, API) src/surfaces/
Orchestration runtime the agent loop, delegation, streaming, budgets src/runtime/
Agent registry agent specs, versioning, sharing, resolution src/registry/
Tool layer MCP servers, tool schemas, confirmation gates src/tools/
Knowledge layer ingest, ACL-tagged chunks, hybrid retrieval src/knowledge/
Model gateway model routing, budgets, fallback, token accounting src/gateway/
If a change touches more than two layers, stop and describe the plan before writing code.
Fill this in before starting:
PROJECT_NAME:     Krow
DOMAIN:           Hospitality and event workforce operations — staffing open shifts,
                  screening and hiring candidates, tracking attendance and overtime,
                  and answering from the organisation's own policy documents.
TENANT_UNIT:      organization  (organizations.id; every table carries org_id NOT NULL)
PRIMARY_SURFACE:  chat  (the Owliver panel, page-scoped, one agent per surface)

Filled from the code rather than from a brief — correct anything that is wrong. TENANT_UNIT in particular is what the schema and the policy table already enforce, not a preference: organizations is the only tenancy boundary, and venue exists nowhere in the schema despite §3's example spec using it.


2. Non-negotiable invariants

These are correctness requirements, not preferences. Violating any of them is a bug even if tests pass. I1 — Agents never expand access. An agent executing on behalf of a caller may read exactly what that caller could read directly, and no more. Not one chunk more, not one row more. This holds for retrieval, tool calls, subagent delegation, and error messages. I2 — ACL filtering happens before scoring, never after. Permission filters are pushed into the vector query and the keyword query as pre-filters. Post-filtering a result set is forbidden — it leaks through result counts, ranking positions, and summaries. Any retrieval function that accepts a query but not a caller principal is wrong by construction. I3 — Every agent run is bounded. Every run carries a hard step cap, a tool-call cap, a wall-clock deadline, and a token budget. There is no "run until done" path. Exceeding a bound terminates the run with a structured BudgetExceeded result, never an exception into user-facing text. I4 — Side effects require explicit confirmation. Any tool that writes, sends, deletes, charges, or notifies is marked effect: write and cannot execute without a resolved confirmation token. The model does not get to decide this. I5 — Tenant isolation is enforced at the data layer. Never rely on a WHERE tenant_id = ? written by hand at a call site. Isolation lives in the repository/session layer so it cannot be forgotten. I6 — Agent specs are data, not code. An agent is a versioned record. Adding an agent must never require a deploy, a new module, or an if agent_key == ... branch anywhere in the runtime. I7 — Prompts are untrusted input. Content retrieved from documents, tool results, and user messages may contain instructions. Never concatenate retrieved text into the system prompt. Retrieved content goes into clearly delimited context blocks, and the system prompt states that content inside them is data.


3. The agent spec contract

The single most important schema in the repo. Lives at src/registry/schema.py. Everything else is CRUD over this.

key: shift-coverage-assistant # stable, unique per tenant, ^[a-z0-9-]+$
version: 3 # monotonic; specs are immutable once published
name: Shift coverage assistant # <= 30 chars, shown in UI
description: Finds and offers cover for open shifts.
instructions: | # the system prompt body
  You help venue managers fill open shifts...
knowledge: # what the agent may retrieve from
  - source: shifts_db
    scope: "venue:{caller.venue_ids}"
  - source: policy_docs
    scope: "tenant:{caller.tenant_id}"
tools: # references into the tool registry
  - find_available_workers
  - send_shift_offer
subagents: [] # keys of other specs this may delegate to
limits:
  max_steps: 8
  max_tool_calls: 12
  deadline_seconds: 60
  model_tier: fast # fast | balanced | deep
conversation_starters:
  - "Which shifts are still uncovered this week?"
visibility: tenant # private | tenant | public
owner: <principal_id>

Rules:

  • Immutable versions. Editing publishes a new version. Running conversations pin the version they started with.
  • scope templates resolve at run time against the caller principal, never at authoring time. An author cannot write venue:*.
  • Unknown tool or subagent keys fail validation at publish, not at run time.
  • subagents must form a DAG. Cycle detection runs at publish. Depth cap is 2.
  • A subagent inherits the parent's caller principal and shares the parent's budget. It never gets a fresh budget.

4. Tool contract

Tools are MCP tools. Do not invent a parallel protocol.

{
  "name": "find_available_workers",
  "description": "...",           # written for the model, not for docs
  "inputSchema": {...},           # JSON Schema, all fields described
  "effect": "read",               # read | write
  "requires_confirmation": False, # forced True when effect == "write"
  "max_result_bytes": 262_144,
}

Implementation rules:

  • Every handler signature is handler(inputs, ctx) where ctx carries the caller principal, tenant, run id, and remaining budget. A handler that ignores ctx for authorization is wrong.
  • Handlers return structured data, not prose. Formatting is the model's job.
  • Truncate at max_result_bytes and set a truncated: true flag. Never silently drop.
  • Tool errors return {"error": {...}} — they do not raise. The runtime decides whether the model sees the error and retries.
  • A tool description that requires the model to guess an ID it has not been given is a design bug. Add a lookup tool instead.

5. Retrieval rules

  • Hybrid: dense + BM25, fused with RRF. Do not replace this with dense-only for convenience.
  • Every chunk row carries tenant_id and an acl field at write time. Chunks without ACL metadata are rejected at ingest.
  • The retrieval entry point is retrieve(query, principal, scopes, k). There is no overload without principal.
  • Retrieved chunks flow to the model with source ids so the response can cite. Responses that assert facts without a retrievable citation must be marked as inference, not grounded fact — keep the two visually and structurally separate in the output payload.
  • Reindex is required whenever ACL derivation logic changes. Note it in the PR.

6. Runtime rules

The agent loop lives in src/runtime/loop.py. It is the highest-risk file in the repo.

  • Single loop, spec-driven. No per-agent branching.
  • Decrement budgets before dispatch, not after, so a hung tool cannot overrun.
  • Stream partial assistant text as it arrives; buffer tool calls until complete.
  • Termination reasons are an enum: Completed | BudgetExceeded | Deadline | ConfirmationPending | ToolFailure | Refused. Every run ends with exactly one.
  • Persist a full trajectory per run: every message, tool call, tool result, and budget snapshot. This is what makes debugging and evals possible — it is not optional telemetry.
  • Delegation is a tool call from the parent's perspective. Subagent runs get their own trajectory, linked by parent_run_id.

7. How to add a new agent

Adding an agent is a data change. If you find yourself editing runtime code, you have found a missing platform capability — surface that instead of special-casing.

  1. Write the spec YAML in agents/<key>.yaml.
  2. Confirm every referenced tool exists. If one is missing, build the tool first (§8).
  3. Confirm every knowledge source exists and is ACL-tagged.
  4. Run make validate-agent KEY=<key> — checks schema, tool refs, scope templates, subagent DAG.
  5. Write at least 5 eval cases in evals/<key>.yaml (§9). This is required, not optional.
  6. Run make eval KEY=<key> and record the baseline in the PR description.
  7. Publish: make publish-agent KEY=<key> — assigns the next version number.

8. How to add a new tool

  1. Define the schema in src/tools/<domain>/schema.py.
  2. Implement handler(inputs, ctx) in the same package. Authorize using ctx.principal on the first line of the handler body.
  3. If effect == "write", add a confirmation payload renderer describing exactly what will happen in plain language.
  4. Unit test authorization first: a caller without rights must get a denial, and the denial must not reveal the existence of the resource.
  5. Register in src/tools/registry.py.
  6. Cap: 20 tools per agent spec. If an agent needs more, it should be split into a parent with subagents.

9. Evals are part of the definition of done

No agent ships without evals. No change to the loop, retrieval, or prompt assembly merges without running the full suite. Each eval case:

- id: uncovered-shifts-basic
  principal: fixtures/manager_two_venues.json
  input: "Which shifts are uncovered this week?"
  expect:
    termination: Completed
    tools_called: [find_open_shifts]
    must_mention: ["Friday evening"]
    must_not_leak: ["venue_9"] # data outside the principal's scope
    max_steps: 4

must_not_leak is mandatory on every case. Every eval doubles as a permission test.

10. Conventions

  • Go (see go-api/go.mod), standard library HTTP with net/http routing patterns, pgx for PostgreSQL. NOT Python: this document specified Python 3.11 / FastAPI / SQLAlchemy / Alembic and the code has never been any of those. Corrected here rather than left to mislead the next reader, which it did.
  • Exported functions carry doc comments. go vet ./... clean; gofmt -w.
  • Errors: structured exception types with a code, never bare strings. User-facing text is derived at the surface layer, not raised from the core.
  • Logging: structured JSON, always include run_id, tenant_id, agent_key, agent_version. Never log message content or retrieved chunks at INFO — that is a data leak into your log store. DEBUG only, behind a per-tenant flag.
  • Config via environment, validated once at startup into a frozen settings object. No os.getenv at call sites.
  • Migrations: golang-migrate, one per PR, reversible (.up.sql and .down.sql).

11. Build order

Do not build ahead of the current phase. Each phase must be working before the next starts.

  • Phase 1 — Runtime skeleton. Two or three hardcoded YAML specs loaded from disk. Loop, budgets, streaming, trajectory persistence. No database registry, no UI.
  • Phase 2 — Tools + knowledge. MCP tool layer, ACL-tagged ingest, permission-aware hybrid retrieval. Evals harness alongside.
  • Phase 3 — Registry. Specs move to the database. Versioning, publish flow, resolution by key + tenant. Still no builder UI.
  • Phase 4 — Surfaces. Chat panel, invocation from the product, webhooks.
  • Phase 5 — Authoring UI. Only once the spec schema has been stable for a meaningful stretch. The builder is a form generator over §3 — if it needs to be more than that, the schema is wrong.

Current phase: Phase 4 — Surfaces.

Phases 1, 2 and 3 are complete and verified against a live model. What remains in Phase 3 is a publish workflow — approval, staged rollout — which §12 says depends on the curated-versus-self-serve decision and is not settled.

Layer State
Surfaces POST /api/v1/agents/{id}/runs (streams over SSE on Accept: text/event-stream), GET /api/v1/runs/{id}; the chat panel is the only answering path — the browser simulator is deleted
Orchestration spec-driven loop, four bounds claimed before dispatch, six terminations, trajectories in agent_runs; delegation per §6 — a subagent is a tool call, runs as the caller, shares the parent budget, capped at depth 2, and writes its own trajectory linked by parent_run_id
Registry 9 agents + 24 skills as rows; published versions immutable (append-only, trigger-enforced); runs pin the version they started with
Tools 19, two of which write (move_application, assign_worker), behind a bound single-use confirmation
Knowledge ACL-tagged ingest, hybrid dense + BM25 fused with RRF, pre-filtered
Gateway tier → model + effort, token accounting, refusal as an outcome; one wire protocol — openai, the chat-completions shape that Groq (the default), Gemini, OpenRouter, Together, vLLM and a local Ollama all serve. The Anthropic path was removed; MODEL_PROVIDER=anthropic, a stale ANTHROPIC_API_KEY and a leftover claude-* id are each refused at startup rather than ignored

Conversational writes are not agent tool calls. Two skills — create-position and create-employee-role — collect a record through the chat panel and then write it with the same REST call the manual form uses, as the signed-in user. They are therefore outside I4's confirmation-token mechanism, which governs tools an AGENT invokes on a caller's behalf. The person is making the request themselves, and the flow's review step ("Ready to create this position?") is where they agree to it. Worth knowing rather than worth fixing: if a write is ever moved from the panel into an agent tool, it acquires I4's bound single-use confirmation at that point and not before.

Deviations from this document, all deliberate and all flagged in code:

  • §3 names the retrieval block knowledge:. The shipped product already uses that key for an author's free-text notes, so retrieval corpora are sources:. Two meanings under one key would be resolved wrongly by whichever parser ran second, silently. See runtime.Agent.KnowledgeSources.
  • §5 asks for BM25. Postgres does not ship it; the keyword half is ts_rank_cd, cover-density ranking. Different function, same job.
  • Vectors are real[] with a dot-product function rather than pgvector, which is not installed. Exact search, no ANN index, bounded by the ACL pre-filter. The upgrade is a column type change and no logic change.
  • Dense retrieval takes its embedder from EMBED_PROVIDER: ollama (local, real semantics, no credential), voyage (hosted), or lexical — a deterministic stand-in that is not semantic and that config validation refuses in production. Unset means keyword-only, which is what production runs today.

12. Open decisions

Do not resolve these unilaterally. Flag them and ask.

  • Who authors agents? Curated (the team ships specs) vs. self-serve (tenants author their own). Self-serve requires prompt-injection hardening at the authoring boundary, per-tenant cost caps, an approval workflow, and a sandbox — roughly 3× the platform. Current assumption: curated, with the registry designed so self-serve is additive later.

  • Model hosting. Self-hosted vs. API vs. mixed by tier. Still open — but no longer expensive to change: MODEL_BASE_URL + the three MODEL_* ids move the whole platform between Groq (the default), Gemini, OpenRouter, Together, vLLM and a local Ollama without a code change, and make eval-live runs the suite against whichever is configured. Decide it on the eval evidence, and weigh the I7 case heaviest: a cheaper model that follows the planted injection is a security regression, not a saving.

    The gap that the Anthropic removal opened here is closed: openai/gpt-oss-120b on Groq has been through make eval-live and passes all three cases including I7 (2026-09-07). The decision itself — self-hosted vs. API vs. mixed by tier — is still open and still not mine to settle.

  • Confirmation UX. Inline in-chat vs. an approval queue.


13. Anti-patterns

Things that look like progress and are not:

  • Filtering retrieval results after scoring "because it's simpler."
  • A special_cases.py in the runtime.
  • Passing the tenant id as a plain function argument through five layers.
  • Letting the model choose whether a write needs confirmation.
  • Fresh budgets for subagents.
  • Concatenating retrieved document text into the system prompt.
  • Building the authoring UI before the spec schema is stable.
  • Adding an agent without evals "for now."
  • Swallowing a tool error and letting the model narrate around it.

14. When stuck

If a requirement seems to demand breaking an invariant in §2, the requirement is wrong or the platform is missing a capability. Say which, and propose the platform change. Do not work around the invariant locally. Show less