Commit Graph

20 Commits

Author SHA1 Message Date
6c85e4cd02 Show the passages an answer rested on, instead of deleting the citations
Some checks failed
CI / check (push) Failing after 5m4s
This morning the panel stripped citation ids because there was nowhere to put
them, and that was right at the time: an id is a thirty-six character address
with no meaning on a screen, and five of them in a sentence is the thing a
reader actually complained about. But stripping them also made a grounded
answer and an invented one look identical, which is the opposite of what
citing is for.

The run now returns its sources, so the fix is the other way round: resolve
the citation rather than remove it. An id the run actually carried becomes the
position of that passage in a list rendered under the answer — [3f8a…] reads
as [2], and the second entry is the one being pointed at.

An id the run did NOT carry is still stripped, and that distinction is the
point. A model citing something it was never given has invented an address,
and giving it a number would turn a hallucinated citation into one that looks
checkable — strictly worse than removing it. Both directions are tested.

The evidence sits after the answer and after any pending write, collapsed. It
is support rather than content: a reader who trusts the answer should not
scroll past the filing to reach what comes next, and a reader who does not
should find it where they reach for it.

A run that retrieved nothing renders no section at all rather than an empty
heading — which is most runs, since seven of nine agents answer from tools.

Verified: 1737/1737 skill-checks, clean typecheck, build and lint, and the
i18n audit still reports no missing keys in either language.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-07 20:44:27 +05:30
26f5116bb0 Remember the last few turns, and stop citation markup reaching the reader
TWO THINGS A READER SAW TODAY.

Owliver had no memory. A run is one turn — the API takes an `input` and no
message list, and agent_runs records each run independently — which is right for
an API and wrong for a panel that looks like a conversation. "Which of those is
at risk?" arrived with no "those".

The proper fix is a `messages` array on the run request. This is not that: the
transcript already lives in the browser, so it travels inside the question until
the API grows a field for it. recall.ts is shaped like that future field so the
swap is a deletion.

BOUNDED IN TOKENS, NOT TURNS, because turns are not a unit of cost: three short
exchanges are nothing and three carrying a table each is a question that no
longer fits. The deployment allows 8,000 tokens a minute and a heavy run already
spends most of it, so recall gets a 600-token ceiling — about 7% of a minute —
each turn clipped to 400 characters, and eviction oldest-first, because dropping
the most recent exchange drops the one the follow-up is about.

The transcript is fenced and labelled as data on the same terms as retrieved
documents: an earlier answer is the model's own words, but an earlier QUESTION
is the reader's, and a reader can type anything.

CITATIONS. context.go hands the model <source id="…"> and said "cite it" without
saying how, so it invented a format per answer. The panel stripped four; a
reader got three it had never seen — the <source> tag echoed back, 【uuid】 in
fullwidth brackets, and <br> drawn as text by the Markdown renderer. All three
are stripped now, <br> becoming a real newline so bullets stay on separate
lines. The fullwidth rule matches horizontal whitespace only: \s* swallowed the
newline a <br> had just become and ran two bullets together, which the test
caught.

Verified: 1732/1732 skill-checks, clean typecheck and build. The citation rules
carry the production answer verbatim as a case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-07 19:46:17 +05:30
81e5979f8f Update Owliver language and suggestions 2026-10-07 19:15:44 +05:30
8ac4e1ee88 Let the reader choose Owliver's language, and stop the suggestion row going silent
Some checks failed
CI / check (pull_request) Failing after 5m8s
Two changes to the assistant panel.

THE LANGUAGE SELECTOR. An account setting in the header menu, not a property of
the conversation: somebody who reads in Spanish reads in Spanish on every page
and after every reload, and making it per-thread would ask them to set it again
each time the panel opened. Held in localStorage and outside React, so the
Owliver panel — a different subtree from the header — sees the change without a
provider spanning both, and a second tab picks it up.

A TAG is sent, never a sentence, and the field is omitted entirely when the
choice is English. The backend maps it onto a closed set
(internal/runtime/language.go holds the same list) and an unrecognised tag
answers in English. The panel therefore cannot write prompt text from here,
which is the point: the selected string SELECTS a directive rather than
becoming one.

THE SUGGESTION ROW. It appeared exactly once per conversation and was silent
after that, whatever was asked. Every chip ever SHOWN was banned permanently in
a set that only grew; the server keeps returning the top of the same small
catalogue, so by the second turn every suggestion was already in it, the filter
emptied the list, and an empty list draws no row.

Asked and offered are not the same thing. A question this thread actually PUT
is excluded for good — it has an answer on screen and offering to repeat it is
not a follow-up. A chip merely DISPLAYED and passed over is held back only from
the turn directly after it, which is enough to stop the row redrawing verbatim
under consecutive answers; beyond that it is offerable again, because ignoring
a suggestion is not the same as having covered it.

Verified: 1710/1710 skill-checks, a clean typecheck, build and lint.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-06 15:27:02 +05:30
1ef0c6fc8a layout change ui fix
Some checks failed
CI / check (push) Failing after 5m6s
2026-09-29 16:59:13 +05:30
878a23532a test(baseline): recapture the three HTML baselines that held invented people
Commit A of the baseline remediation. The Owliver baseline is deliberately
untouched and is Commit B.

Five of the six failing checks were the `DEMO_FILL` removal. `hiringRecords`
used to pad the hires list with five invented people so Hired History read
as a history rather than as three rows, and the padding reached Analytics
too, where "total hires" counted eight against a database holding three.
Because these baselines render with queries disabled, the padding was ALL
they contained. The sixth is Candidates becoming the final-selection queue,
which changed what its empty state says.

Verified before regenerating rather than after. Hired History's tag counts,
baseline -> now: `<td` 40->0, `<span` 51->13, `<div` 78->39, `<p` 26->7,
`<svg` 20->8. Analytics lost 269 class attributes and 238 words to the same
cause; Candidates gained exactly one class, the description paragraph under
"Nobody is awaiting a decision". Every generated file was checked for the
five invented names (none), for the intended new copy (present) and for its
node identities before it was put in place.

Renamed, because a file called `pre-migration` holding post-`dca1842`
markup is a lie in the filename. `activity-page`, `candidates-analysis`,
`control-center`, `positions` and `talent-pool` still hold genuine
pre-migration markup, still pass, and keep the name that says so — so the
`PAGES` table now carries each baseline's FILENAME rather than deriving one
suffix for all five.

The migration proof was replaced, not discarded. `Hired History added only
identity wrappers` compared tag tallies to assert that migrating the page
added exactly two `<div>`s and changed nothing else. That was true, and it
was checkable only while the DATA was frozen as well: the predicate
subtracts one render from another, so removing the invented hires moved
every count and the arithmetic stopped describing wrappers. Recapturing
would not have rescued it — with the baseline equal to the render the delta
is zero and a predicate demanding two can never hold — so left in place it
would have stayed red for a new reason. Three checks read the render
directly instead and need no frozen file:

    Hired History wraps exactly the node types registered to wrap
    Hired History identities are unique and name composed nodes
    Hired History identity wrappers carry no styling

They say what the tally said: `UiTreeRenderer` encloses a type registered
`wrap: true` in `<div {...attrs}>`, every other type takes the attributes on
its own root element, and a wrapper carries identity and no styling.

Writing them first, against the OLD baseline, is what caught my own error:
the second check began as "one identity per composed node" and failed at
4 of 5. Hired History composes five nodes and `hired-extensions-top` is an
extension slot that renders nothing while no skill is attached to it, so a
one-to-one rule would have asserted that every slot is always filled. It
asserts uniqueness and no strays instead — a node addressed twice, or an
identity naming nothing the page composed, would break the editor's and
Owliver's ability to name a node. Had the baselines been recaptured first,
that mistake would have been invisible.

`carries node identity in the DOM` was documented as needing pre-migration
markup. It does not — it reads only the live render — so it is unchanged and
still passes. The dead `tally` helper is removed with its last caller.

  typecheck   0 errors
  lint        exit 0, 0 errors, 289 warnings
  npm test    1693/1693, exit 0 — was 1685/1691
  build       exit 0, bundle 74d17e2d… unchanged
  owliver     still DRIFTED — that is Commit B, untouched here

No file under `src/` changed, which is why the bundle hash cannot move.
`owliver-baseline.json` is byte-identical, the backend is untouched, and CI
configuration is unchanged. The suite total rises 1691 -> 1693 because one
assertion became three; CI asserts `passed == total` and a floor of 900,
both satisfied.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HBG1wnuRfJKCstGB8Fekr8
2026-09-20 00:52:53 +05:30
02a2ab05ef chore(ts-migration): resolve source-text reads across .js/.jsx/.ts/.tsx
The check suite consumes source files through two entirely separate channels,
and Phase 0 only hardened one of them.

`withSourceResolution` wraps `ssrLoadModule`, which covers the 159 module loads.
It does not and cannot see the other 30 sites, which read source as TEXT through
`readFileSync` to assert structural facts - "the panel imports no local
suggestion ranker", "no runtime path writes a definition". Renaming
`base44Client.js` to `.ts` is what surfaced the difference: ENOENT in the middle
of a suite that had been passing, from a line no grep for `ssrLoadModule` would
ever have found.

`resolveSourcePath()` in `ssr-resolve.mjs` reuses the existing `candidatesFor()`
ordering - the written path first, then `.ts`, then `.tsx` - and returns a path
relative to the root, so every call site keeps its `join(ROOT, ...)` as it was.
An unresolvable path comes back unchanged, which keeps the two absence
assertions honest: a check proving `store.js` is GONE still asks about the path
it means, and now also notices if the file returns under another extension.

Thirty-seven lines change, each a one-for-one replacement. No assertion text, no
record() message, no ordering, no logic.

The audit found the reads in four shapes, and two of them a path grep cannot
see:

  - direct       readFileSync(join(ROOT, 'src/x.jsx'), 'utf8')
  - via a const  const P = join(ROOT, 'src/x.jsx')
  - dir + name   ['node.js', 'patch.js'].map((f) => readFileSync(join(ROOT, 'src/lib/ui', f)))
  - path array   for (const f of files) readFileSync(join(ROOT, f))

The third and fourth hide thirteen filenames in adjacent arrays, which is why
the first estimate of this work was twenty-two files and the real number is
thirty-five.

One directory scan also filtered `/\.jsx?$/` over `src/pages/admin` and
`src/components/agents`. That one does not crash - it quietly matches nothing
once those directories are TypeScript, and the check passes having inspected an
empty set. Widened to `/\.[jt]sx?$/`. A silent shrink is worse than a failure,
and CI's FLOOR of 900 would not have caught it.

Deliberately NOT touched, because they are correctly extension-specific: the
`dist/assets` filter reads built bundles, which are `.js` whatever the source
was; `scripts/stubs/react-hot-toast.js` and `scripts/browser-flows.js` are
scripts, not migrated source; and every comment that mentions a `.js` filename
says something true about a file that still has that name.

Proven rather than assumed. `npm test` reports 1684/1691 before and after, the
same seven failures. Then `src/api/base44Client.js` - the exact file whose
rename broke the suite - was renamed to `.ts`, the suite re-run, and the check
that reads it as text passed: "the app does not import the seed fixture". The
rename was reverted and the file verified byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HBG1wnuRfJKCstGB8Fekr8
2026-09-17 22:56:45 +05:30
dca184289e feat(hiring): final-selection queue, honest seat counts, and human interviews
Authored in a parallel session alongside the TypeScript migration; committed
separately so the two never share a commit. No TypeScript migration file is
included here.

Candidates becomes the queue of hiring decisions waiting on a person, rather
than a second Talent Pool listing every application the org ever took. Final
selection is DERIVED - there is no `final_selection` value in the
`application_status` enum and none is added. The fact it reads is the existence
of an interview row, a NOT NULL foreign key, rather than
`job_applications.interview_id`, which the schema keeps as an unconstrained soft
reference precisely so it may dangle. `status = 'interview'` is set both when an
interview is arranged and when one is completed, so status alone cannot tell a
queue of people who have been interviewed from a queue of people merely booked
in.

Seats on a position are counted from the employment records instead of a stored
column. A `filled` counter would be a second source of truth, and the day it
disagreed with `staff` nothing could say which was lying. Someone who has left
frees their seat, and over-hiring floors at zero rather than going negative.

`DEMO_FILL` is gone. `hiringRecords.js` padded the hires list with five invented
people so Hired History read as a history rather than as three rows; the padding
reached Analytics too, where "total hires" counted eight against a database
holding three. Hires now come only from `staff`.

Both paths that file an application on somebody's behalf now carry
`worker_profile_id`, the link back to the talent-pool record. The column is
nullable, so omitting it saved cleanly and failed silently: the application
belonged to an email address rather than to a person, and the hire it became
could not be traced back to the profile it came from.

`HiredChronology` used to `return null` with no hires, taking the `chronology`
node identity out of the DOM with it - so on an honest empty dataset the section
could not be addressed by Owliver or the layout editor at all. It now renders an
empty state inside the section it keeps.

Nine new checks cover the above; `npm test` reports 1684/1691.

SIX SSR PARITY CHECKS FAIL ON PURPOSE, and `scripts/__baseline__/README.md`
documents each with verified tag counts. Five are the `DEMO_FILL` removal: the
Hired History and Analytics baselines were captured while the padding was in
effect and, because they render with queries disabled, the padding is all they
contain. The sixth is this change to what Candidates says. Do not regenerate
those baselines to clear them - two of the checks exist to prove the UI node
tree migration added exactly two `<div>`s, and that proof needs the baseline to
be pre-migration markup. Recapturing now would write post-migration markup into
a file named `pre-migration` and the check would compare the current render
against itself forever. The debt is held until the migration work lands, when
both files are recaptured together.

The seventh failure, the stale backend seed fixture, predates all of this.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HBG1wnuRfJKCstGB8Fekr8
2026-09-17 22:52:51 +05:30
1775395256 chore(ts-migration): establish TypeScript migration checkpoint
Phases 0-3 of the JS/JSX -> TS/TSX migration. No runtime behaviour changes:
every converted file emits byte-identical JavaScript, verified file by file.

Phase 0 - harness hardening, before any rename:
  - scripts/ssr-resolve.mjs wraps `ssrLoadModule` so the ~170 literal module
    paths in the check scripts resolve .js/.jsx/.ts/.tsx. Without it the first
    rename would have silently destroyed the 1642-check suite that guards the
    Owliver flow.
  - eslint.config.js gains a TypeScript block. Its `files` globs listed only
    {js,mjs,cjs,jsx}, so a renamed file would have dropped out of the run while
    `eslint .` went on exiting 0 - the quietest failure mode available.
  - MIGRATION_BASELINE.md records the measured starting point, including the
    pre-existing seed-fixture failure and the already-broken standalone
    owliver-baseline.mjs, so neither is later mistaken for migration damage.

Phase 1 - tsconfig.json succeeds jsconfig.json, carrying every option across at
its old value. `types` moves from [] to ["vite/client"], which fixes the eight
import.meta errors; @types/node is deliberately excluded so setTimeout stays a
number in browser code. allowJs and checkJs stay on, strict stays off.

Phase 2 - src/types/{api,entities,user}.ts. Transport envelope, error shape,
the entity-name union (the same 18 names are written down twice today, in
httpClient and base44Client, with nothing checking they agree), and the user
record. Every field transcribed from the API contract, the migrations and the
/me projection in me.go - not inferred. Entity record shapes are deliberately
absent: derivable, but nothing consumes them yet.

Phase 3 - nine leaf utilities renamed to .ts with annotations added only where
they could be established from existing usage. Entity records are typed `any`
with a comment naming what they are, rather than a guessed interface.

Verified: tsc 64 errors (65 before; one pre-existing TS2559 genuinely fixed,
none introduced), lint unchanged at 0 errors, skill-check 1641/1642 with only
the known failure, Owliver baseline section 59/59 green, production build
succeeds with the API origin inlined.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HBG1wnuRfJKCstGB8Fekr8
2026-09-11 15:52:01 +05:30
f522b6508e archive issue fix
Some checks failed
CI / check (push) Failing after 5m5s
2026-09-10 19:29:39 +05:30
6249e00a3a candidates and board ui agent issue
Some checks failed
CI / check (push) Failing after 4m58s
2026-09-05 10:46:06 +05:30
e02a0c23d4 Make a conversation a registry, and add the second one
Some checks failed
CI / check (push) Failing after 4m57s
`positionFlow.js` was five per-domain concerns in one file — a field table, an
`@`-token resolver, a sentence extractor, a commit vocabulary and a set of
outcome renderers — and only the control flow between them was general.
Everything else knew it was creating a job posting. Adding a second
conversation meant a second copy of all of it.

So the control flow is now `conversationFlow.js` and each kind of record is a
REGISTRY. Which conversation a skill runs is the skill's own `flow:` line,
resolved through a map: `routing.js` used to say `if (skill.id ===
'create-position')`, which made a second conversational skill a change to the
router rather than a file on disk — the `if agent_key == ...` shape the
platform rules out one level up. The panel's write callback is likewise a map
keyed by flow id instead of an `onCreatePosition` prop, and the outcome wording
comes off the registry, so nothing in the panel names a kind of record any
more.

`create-employee-role` is the second registry. The worker is asked for and
never assumed: a conversation that names nobody re-asks rather than falling
back to the session, because an operator records this on somebody's behalf. Its
`extract` is deliberately narrower than the posting's — "bartender, weekends,
$30/hr" settles three fields and leaves the subject alone, since guessing WHO a
record is about from a fragment is how a role gets filed against the wrong
person.

THE CONFIRMATION STEP ACCEPTED "create position" AND SILENTLY REJECTED "create
positions" — the plural the Positions page itself uses. An anchored regex missed
it, and the reader got the summary back with no indication of what was wrong
with what they said, which is indistinguishable from the screen not having
updated. Matching is now exact membership against a normalized reply, so a
vocabulary is a list of phrases somebody can read rather than an expression
somebody has to parse.

Two bugs in `extractRole`, both of which fabricated a value nobody typed on the
one field a position cannot be created without:

  - The phrase pattern marks "new" as the role by the same grammar that marks
    "sous chef", so "create new position" opened the conversation titled "New".
    The scaffolding is a PHRASE at least as often as a single word, so a
    per-word test still produced "Brand New" and "One More". Scaffolding words
    are now stripped to DECIDE whether the phrase named anything, and the
    ORIGINAL phrase is returned when it did — strip to test, never to rewrite,
    or "second chef" becomes "Chef" and the cure is worse than the bug.

  - `(?:a|an)?\s*` has no word boundary, so it matched the leading "a" of
    "another" and the capture began mid-word. That mangled scaffolding into
    "Nother New" and, worse, corrupted every role introduced with "an":
    "create an open kitchen lead position" titled the position "N Open Kitchen
    Lead". A real role, typed correctly, silently wrong. Found by mutation
    testing the first fix.

`@companies` and `@workers` resolve from data the panel already holds — postings
and profiles the API has already scoped to the caller — so neither widens
anybody's view and neither costs a request. The company list is deliberately
unsorted: `useJobPostings` asks for `-created_date`, so the clients staffed for
most recently come first, and the panel does no ranking of its own. That last
part is a rule the suite enforces structurally, and it is the right rule — a
second opinion formed in the panel outranking the server's is exactly the kind
of thing that decays quietly.

`npm test` now refuses a conversation step whose field its registry does not
define. `stepsOf` drops unknown fields, so a typo means the flow asks fewer
questions than the file lists — and a skill whose steps are ALL unknown asks
none, jumps to the summary, and offers to write an empty record. Nothing errors
and the Markdown still reads correctly. 970 checks, up from 924; the new ones
walk both conversations end to end, because a wrong answer at the confirmation
step re-renders the same summary a right answer does.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-02 15:29:48 +05:30
1a0dc7e5f1 Settle subagents: it means delegate to, not borrow skills from
Some checks failed
CI / check (push) Has been cancelled
The key had two incompatible meanings running at once. CLAUDE.md §3 defines
subagents as "keys of other specs this may DELEGATE to" and §6 as a tool call
from the parent's perspective — the subagent runs its own turn, as the same
caller, out of the parent's budget, and returns an answer. The backend
implements exactly that.

This side did something else: agentSkillIds folded one level of subagent skills
into the parent's carried set, so a parent silently gained everything its
subagents carried, and the UI described it that way — "other agents whose
skills this one may also use", "borrowing them cannot reach data this page does
not hold".

Both are defensible readings. Only one is the specification, and running both
meant krow-workforce-agent carried eight delegation tools AND the flattened
skills of those same eight agents — able to answer a question directly or to
ask an agent that had already lent it the means to answer. Two ways to do one
thing, differing in cost and in what the trajectory records.

So an agent carries what it declares. Reaching another agent is delegation,
which the runtime does with its own budget and its own trajectory.

The blast radius was one check, which is the useful part of the answer: only
"a subagent cycle terminates" depended on the folding, because that traversal
was the only thing that could loop. Nothing walks subagents here now, so the
cycle question moved to where it belongs — refused at publish by
definition.FindSubagentCycle, bounded at run time by the depth cap. The
replacement checks assert the new meaning rather than deleting the old ones,
because the previous behaviour reads as perfectly reasonable and will be
reinvented otherwise.

The `agents` parameter stays on agentSkillIds and is no longer read. Removing
it is a wider edit for no behavioural gain, and agentScopedDisabledWith exists
precisely to pass it.

924/924 checks pass; production build clean.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 15:29:09 +05:30
88c412fda6 skill issue 2026-08-28 12:22:08 +05:30
814073e97b agnets done 2026-08-28 11:13:20 +05:30
eaf08e061d agnets done 2026-08-28 11:02:02 +05:30
b2e6868824 update agents skill design 2026-08-20 18:18:10 +05:30
161b237695 update the workspace and skills flow 2026-08-20 11:04:13 +05:30
d3f7f439f6 update Markdown skills 2026-08-19 17:36:27 +05:30
7fa21a4513 fix error line 2026-08-17 17:33:56 +05:30