Add evals for every shipped agent, a policy corpus, and CI
§9 says no agent ships without evals. Eight of the nine had none: the two
other suites in evals/ are harness fixtures rather than agents in the
registry, so the rule was being met by one agent in nine.
Evals — 40 new cases, five per agent, every one carrying mustNotLeak:
- the agent is loaded from its real spec in agents/*.md rather than
written out again in Go. A hand-copied agent tests the copy: it keeps
passing after somebody edits the spec, which is the moment it most
needed to fail.
- callNamed calls the tool a case names. toolThenAnswer always called
tools[0], so seven of positions-agent's eight tools were unreachable,
and a boundary nothing calls is a boundary nothing tests.
- seedWorkspace fills BOTH tenants. A leak test against an empty second
tenant cannot fail.
Verified by breaking workersByScore's org predicate: six cases across four
agents fail with LEAKED "RIVAL".
Knowledge — six policy documents, taking the corpus from 2 to 8 (34
chunks). Three restricted to admin and employer, five tenant-wide. They
cover what the tools cannot: a tool reports how many shifts went unworked,
a policy says what cover costs inside 24 hours.
corpus_test.go treats those documents as product rather than fixtures. The
first version was tautological — it read audience: from a file and checked
that file's audience was enforced, so opening a restricted document passed.
mustNotBeTenantWide now holds that judgement apart from the files, with the
reason recorded for each.
CI — the checks this repository already had, made unskippable. testutil
calls t.Skipf on an unreachable database, so a dead service container would
produce a green build over a suite that ran almost nothing. Simulated: go
test exits 0 with 74 tests skipped, including every tenant-isolation test.
The guard exits 1 and names them, while still allowing TestLive* to skip
without a model key.
This CI tests; it does not deploy. The README's claim that migrations are
run by CI against the target database remains aspirational.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
This commit is contained in:
149
.github/workflows/ci.yml
vendored
Normal file
149
.github/workflows/ci.yml
vendored
Normal file
@@ -0,0 +1,149 @@
|
||||
name: CI
|
||||
|
||||
# Every check this repository already had, run on every push instead of when
|
||||
# somebody remembers. Nothing here is new verification — it is the verification
|
||||
# that existed, made unskippable.
|
||||
|
||||
on:
|
||||
push:
|
||||
branches: [main]
|
||||
pull_request:
|
||||
workflow_dispatch:
|
||||
|
||||
concurrency:
|
||||
group: ci-${{ github.ref }}
|
||||
cancel-in-progress: true
|
||||
|
||||
jobs:
|
||||
test:
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
# The tests SKIP when PostgreSQL is unreachable — see testutil.New, which
|
||||
# calls t.Skipf rather than failing, so a developer without a database can
|
||||
# still run the non-database tests. In CI that behaviour is a trap: a broken
|
||||
# service container would produce a green build over a suite that tested
|
||||
# almost nothing. The guard at the end of this job is what closes it.
|
||||
services:
|
||||
postgres:
|
||||
image: postgres:16-alpine
|
||||
env:
|
||||
POSTGRES_PASSWORD: postgres
|
||||
POSTGRES_USER: postgres
|
||||
POSTGRES_DB: postgres
|
||||
ports: ['5432:5432']
|
||||
options: >-
|
||||
--health-cmd "pg_isready -U postgres"
|
||||
--health-interval 5s
|
||||
--health-timeout 5s
|
||||
--health-retries 20
|
||||
|
||||
env:
|
||||
DATABASE_HOST: 127.0.0.1
|
||||
DATABASE_PORT: '5432'
|
||||
DATABASE_USER: postgres
|
||||
DATABASE_PASSWORD: postgres
|
||||
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
|
||||
- uses: actions/setup-go@v5
|
||||
with:
|
||||
go-version-file: go-api/go.mod
|
||||
cache-dependency-path: go-api/go.sum
|
||||
|
||||
- name: gofmt
|
||||
working-directory: go-api
|
||||
run: |
|
||||
unformatted="$(gofmt -l ./cmd ./internal)"
|
||||
if [ -n "$unformatted" ]; then
|
||||
echo "not gofmt'd:"; echo "$unformatted"; exit 1
|
||||
fi
|
||||
|
||||
- name: go vet
|
||||
working-directory: go-api
|
||||
run: go vet ./...
|
||||
|
||||
- name: Tests
|
||||
working-directory: go-api
|
||||
# -json so the guard below can count what actually ran, and `|| true` so
|
||||
# a failure reaches that guard rather than ending the job here — the
|
||||
# guard reports which tests failed, which the raw JSON does not.
|
||||
run: go test ./... -count=1 -json > /tmp/test.json || true
|
||||
|
||||
- name: Fail if the database tests skipped
|
||||
# The point of this job. testutil skips on an unreachable database, so
|
||||
# "0 failures" is not the same as "the suite ran": a service container
|
||||
# that never came up would otherwise look identical to a passing build.
|
||||
run: |
|
||||
python3 - <<'PY'
|
||||
import json, sys
|
||||
skipped, passed, failed = [], 0, []
|
||||
for line in open('/tmp/test.json'):
|
||||
line = line.strip()
|
||||
if not line.startswith('{'):
|
||||
continue
|
||||
try:
|
||||
e = json.loads(line)
|
||||
except ValueError:
|
||||
continue
|
||||
if e.get('Action') == 'skip' and e.get('Test'):
|
||||
skipped.append(e['Test'])
|
||||
if e.get('Action') == 'pass' and e.get('Test'):
|
||||
passed += 1
|
||||
if e.get('Action') == 'fail' and e.get('Test'):
|
||||
failed.append(e['Test'])
|
||||
print(f"{passed} passed, {len(failed)} failed, {len(skipped)} skipped")
|
||||
if failed:
|
||||
print("FAILED:"); [print(" ", t) for t in failed[:40]]
|
||||
sys.exit(1)
|
||||
# A Test<Pkg>Live* / TestLive* test skips without ANTHROPIC_API_KEY, which
|
||||
# is correct here: CI should not spend tokens on every push, and the key
|
||||
# should not be present unless somebody put it there deliberately. Any
|
||||
# OTHER skip means the database was unreachable, and that is the case
|
||||
# this guard exists for — testutil calls t.Skipf rather than failing, so
|
||||
# a dead service container would otherwise look exactly like a pass.
|
||||
unexpected = [t for t in skipped if not t.startswith('TestLive')]
|
||||
if skipped:
|
||||
print("skipped:"); [print(" ", t) for t in skipped[:40]]
|
||||
if unexpected:
|
||||
print("\nThese are not live-model tests, so they skipped because the")
|
||||
print("database was unreachable. A green build over a suite that did")
|
||||
print("not run is worse than a red one.")
|
||||
sys.exit(1)
|
||||
if passed < 200:
|
||||
sys.exit(f"only {passed} tests passed; the suite is far smaller than expected — did it run?")
|
||||
PY
|
||||
|
||||
- name: Migrations are reversible
|
||||
# §10: one migration per PR, reversible. Applying and rolling back the
|
||||
# newest one is the cheapest way to find out that it is not.
|
||||
working-directory: go-api
|
||||
run: go test ./internal/repo/... -run 'Migration' -count=1
|
||||
|
||||
fixture:
|
||||
# seed.json is generated from the frontend's seed module. The two used to be
|
||||
# hand-maintained copies, and drift between them is silent: the demo and the
|
||||
# API answer the same question with different numbers.
|
||||
runs-on: ubuntu-latest
|
||||
steps:
|
||||
- uses: actions/checkout@v4
|
||||
- name: Check out the frontend beside this repo
|
||||
uses: actions/checkout@v4
|
||||
with:
|
||||
repository: ${{ github.repository_owner }}/krow-demo
|
||||
path: ../krow-demo
|
||||
# A private sibling needs a token with read access; without one this
|
||||
# job reports that it could not check, rather than passing quietly.
|
||||
token: ${{ secrets.FRONTEND_REPO_TOKEN }}
|
||||
continue-on-error: true
|
||||
- uses: actions/setup-node@v4
|
||||
with:
|
||||
node-version: '20'
|
||||
- name: seed.json matches the frontend seed module
|
||||
run: |
|
||||
if [ ! -d ../krow-demo ]; then
|
||||
echo "krow-demo is not available to this job, so the fixture could not be checked."
|
||||
echo "Set FRONTEND_REPO_TOKEN to enable it. Not passing silently."
|
||||
exit 1
|
||||
fi
|
||||
cd ../krow-demo && npm ci && npm run seed:check
|
||||
Reference in New Issue
Block a user