TWO THINGS A READER SAW TODAY.
Owliver had no memory. A run is one turn — the API takes an `input` and no
message list, and agent_runs records each run independently — which is right for
an API and wrong for a panel that looks like a conversation. "Which of those is
at risk?" arrived with no "those".
The proper fix is a `messages` array on the run request. This is not that: the
transcript already lives in the browser, so it travels inside the question until
the API grows a field for it. recall.ts is shaped like that future field so the
swap is a deletion.
BOUNDED IN TOKENS, NOT TURNS, because turns are not a unit of cost: three short
exchanges are nothing and three carrying a table each is a question that no
longer fits. The deployment allows 8,000 tokens a minute and a heavy run already
spends most of it, so recall gets a 600-token ceiling — about 7% of a minute —
each turn clipped to 400 characters, and eviction oldest-first, because dropping
the most recent exchange drops the one the follow-up is about.
The transcript is fenced and labelled as data on the same terms as retrieved
documents: an earlier answer is the model's own words, but an earlier QUESTION
is the reader's, and a reader can type anything.
CITATIONS. context.go hands the model <source id="…"> and said "cite it" without
saying how, so it invented a format per answer. The panel stripped four; a
reader got three it had never seen — the <source> tag echoed back, 【uuid】 in
fullwidth brackets, and <br> drawn as text by the Markdown renderer. All three
are stripped now, <br> becoming a real newline so bullets stay on separate
lines. The fullwidth rule matches horizontal whitespace only: \s* swallowed the
newline a <br> had just become and ran two bullets together, which the test
caught.
Verified: 1732/1732 skill-checks, clean typecheck and build. The citation rules
carry the production answer verbatim as a case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>