Commit Graph

4 Commits

Author SHA1 Message Date
54309635a3 Name one citation format, so the model stops inventing them
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 9s
ContextInstruction asked the model to cite a source id and never said how, so
it chose a different syntax on different days: a <cite> tag, a markdown link to
an empty anchor, the id narrated in a parenthesis, the <source> tag copied
straight back, and fullwidth brackets. The panel has no citation surface, so
each one arrived on a reader's screen as literal markup — on 2026-10-07
somebody read an answer carrying 【f34e8ef0-…】 twice in one sentence, and a
<br> drawn as text between two bullets.

So the instruction names ONE shape: square brackets, no HTML tags, no links, no
other kind of bracket, and never an HTML tag such as <br>. Square brackets
because that is the spelling the renderer already removes cleanly and it reads
as a reference to anyone who sees it before the strip.

The frontend still strips every spelling seen so far and that cannot be removed
— a model is free to ignore any instruction. The difference is between a rule
that holds and a rule patched after each new sighting.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-07 19:49:29 +05:30
8bc6c23770 Spend fewer tokens per run: fewer chunks, a terser shared schema, a sane result cap
The deployment's provider ceiling is 8,000 tokens a minute and a three-call run
measured 12,123, so a single question could not fit inside a minute's budget.
That is the whole of the "the model did not answer" the chat panel has been
showing: the retry loop fires three times and the provider refuses all three.

Three cuts, measured against the real corpus and the real registry:

  DefaultK 8 -> 4            ~705 -> ~352 tokens per call
  periodSchema period help   attached to THIRTEEN tools, re-sent every call
  DefaultMaxResultBytes      262_144 -> 32_768

A three-call control-center run goes from ~12,000 to ~10,700 tokens, an 11%
cut. STATED PLAINLY BECAUSE IT IS NOT ENOUGH: that is still above 8,000, and an
earlier estimate of ~7,000 was wrong. The tool catalogue is 1,312 tokens for
seven tools — about 190 each, which is JSON Schema structure rather than
padding, so trimming prose cannot reach it. The remaining lever is giving an
agent fewer tools, and that is a decision about what the agent can answer, not
a cleanup.

The result cap is the one with no downside: 256KiB let a single tool result
outweigh everything else in the prompt put together. 32KiB is ~8,000 tokens,
still more evidence than one answer needs.

DefaultK is a real trade: half the evidence behind a grounded answer. The corpus
is 43 chunks, so four is still ~10% of it per query, and the eval suites pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-06 13:10:13 +05:30
a222dcd3e4 Add evals for every shipped agent, a policy corpus, and CI
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
§9 says no agent ships without evals. Eight of the nine had none: the two
other suites in evals/ are harness fixtures rather than agents in the
registry, so the rule was being met by one agent in nine.

Evals — 40 new cases, five per agent, every one carrying mustNotLeak:

  - the agent is loaded from its real spec in agents/*.md rather than
    written out again in Go. A hand-copied agent tests the copy: it keeps
    passing after somebody edits the spec, which is the moment it most
    needed to fail.
  - callNamed calls the tool a case names. toolThenAnswer always called
    tools[0], so seven of positions-agent's eight tools were unreachable,
    and a boundary nothing calls is a boundary nothing tests.
  - seedWorkspace fills BOTH tenants. A leak test against an empty second
    tenant cannot fail.

Verified by breaking workersByScore's org predicate: six cases across four
agents fail with LEAKED "RIVAL".

Knowledge — six policy documents, taking the corpus from 2 to 8 (34
chunks). Three restricted to admin and employer, five tenant-wide. They
cover what the tools cannot: a tool reports how many shifts went unworked,
a policy says what cover costs inside 24 hours.

corpus_test.go treats those documents as product rather than fixtures. The
first version was tautological — it read audience: from a file and checked
that file's audience was enforced, so opening a restricted document passed.
mustNotBeTenantWide now holds that judgement apart from the files, with the
reason recorded for each.

CI — the checks this repository already had, made unskippable. testutil
calls t.Skipf on an unreachable database, so a dead service container would
produce a green build over a suite that ran almost nothing. Simulated: go
test exits 0 with 74 tests skipped, including every tenant-isolation test.
The guard exits 1 and names them, while still allowing TestLive* to skip
without a model key.

This CI tests; it does not deploy. The README's claim that migrations are
run by CI against the target database remains aspirational.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
2026-08-28 13:52:46 +05:30
f7df96c973 agent build 2026-08-28 12:21:44 +05:30