§9 says no agent ships without evals. Eight of the nine had none: the two
other suites in evals/ are harness fixtures rather than agents in the
registry, so the rule was being met by one agent in nine.
Evals — 40 new cases, five per agent, every one carrying mustNotLeak:
- the agent is loaded from its real spec in agents/*.md rather than
written out again in Go. A hand-copied agent tests the copy: it keeps
passing after somebody edits the spec, which is the moment it most
needed to fail.
- callNamed calls the tool a case names. toolThenAnswer always called
tools[0], so seven of positions-agent's eight tools were unreachable,
and a boundary nothing calls is a boundary nothing tests.
- seedWorkspace fills BOTH tenants. A leak test against an empty second
tenant cannot fail.
Verified by breaking workersByScore's org predicate: six cases across four
agents fail with LEAKED "RIVAL".
Knowledge — six policy documents, taking the corpus from 2 to 8 (34
chunks). Three restricted to admin and employer, five tenant-wide. They
cover what the tools cannot: a tool reports how many shifts went unworked,
a policy says what cover costs inside 24 hours.
corpus_test.go treats those documents as product rather than fixtures. The
first version was tautological — it read audience: from a file and checked
that file's audience was enforced, so opening a restricted document passed.
mustNotBeTenantWide now holds that judgement apart from the files, with the
reason recorded for each.
CI — the checks this repository already had, made unskippable. testutil
calls t.Skipf on an unreachable database, so a dead service container would
produce a green build over a suite that ran almost nothing. Simulated: go
test exits 0 with 74 tests skipped, including every tenant-isolation test.
The guard exits 1 and names them, while still allowing TestLive* to skip
without a model key.
This CI tests; it does not deploy. The README's claim that migrations are
run by CI against the target database remains aspirational.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
Documents agents can read
Markdown files, one per document, ingested by:
make ingest ORG=<slug>
Each file declares who may read it in its front matter. That declaration is required — §5 says a chunk without ACL metadata is rejected at ingest, and the reason is worth stating: an empty audience is not "private", it is a row the permission filter can never match. A document that indexed to nothing looks ingested, reports a chunk count, and is silently unreachable forever.
---
source: policy_docs
audience: tenant # everyone in the organization
title: Staff Handbook
---
audience accepts:
| value | who can read it |
|---|---|
tenant |
everyone in the organization |
role:admin, role:employer, role:talent |
one role (comma-separate for several) |
email:someone@example.com |
one person, by email |
source is the corpus name. An agent spec's sources: block names which
corpora it may retrieve from, so this is part of the permission story rather
than a label: an agent granted policy_docs does not thereby gain
worker_notes.
Re-ingesting is safe. A document whose content and audience are unchanged is a no-op; changing either rewrites its chunks, because the chunks carry a copy of the audience and a permission change that did not reach them would be a permission change that did not happen.