Commit Graph

39 Commits

Author SHA1 Message Date
797ee5f2d2 Round-trip provider metadata on tool calls; Gemini requires it
Gemini 3 models attach a thought signature to every function call
and reject the follow-up -- 400, "Function call is missing a
thought_signature in functionCall parts" -- when the assistant
message echoing that call does not carry it back. The gateway
rebuilt the assistant turn from id, name and arguments alone, so
every tool-using run on Gemini died on its second model call, after
a first call that looked perfectly healthy. Found when production
was pointed at Gemini on 2026-09-22; rolled back to Groq within
minutes.

ToolCall gains an opaque Extra field: the raw JSON of the wire's
extra_content, captured on both the streaming and non-streaming
paths and emitted verbatim on the next request. The gateway does
not read it and must not -- the point of one wire shape is that a
vendor's private fields pass through untouched. Absent stays
absent; no provider receives a null it never sent.

Also: Gemini wraps its error body in a one-element array, which
the message parser read as "no detail". The trajectory therefore
said only "the model rejected the request" where the body named
the missing signature outright. Unwrapped now, so the next
provider quirk is legible in the trajectory instead of costing a
day of proxy captures.

Verified end to end with the real gateway against real Gemini: a
three-turn tool-calling run completed and the proxy confirmed the
signature on every echoed call.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-22 13:15:39 +05:30
5166fde764 Add GatewayFailure: the provider not answering is not a tool failing
Some checks failed
CI / test (push) Failing after 4m37s
CI / fixture (push) Failing after 8s
terminationFor sent every gateway error that was not Refused or Timeout
to ToolFailure, because the enum had nowhere else to put it. On
2026-09-22 that was 131 of 318 production runs, and not one of them was
a tool failing: 56 were retired model ids, 25 an exhausted Anthropic
balance, 45 Groq's free-tier rate limit -- the only one still happening.
An operator reading the termination column saw "a tool is broken" for
two weeks while the actual answer was "we are not paying for capacity".

GatewayFailure is the seventh termination. Rate limited, request
rejected, credential refused and unreachable land there; Refused and
Deadline keep their own reasons; a non-gateway error is still the tool
layer's. A delegation whose subagent died at the gateway now carries
that reason up to the parent instead of reading as a tool call that
failed.

Migration 000016 widens the CHECK that 000006 chose precisely so this
would be a migration rather than an ALTER TYPE. Its down folds any
GatewayFailure rows back to ToolFailure BEFORE narrowing the constraint,
which is the order that works; verified up, down and up again on a
scratch database. Existing rows are left as they are -- the trajectory
entries still carry the gateway.* code for anyone reclassifying history.

The surface wording is the one termination where "try again" is honest
advice, since the dominant cause clears within a minute.

Full suite run against a real database, including the tests that skip
without one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-22 12:48:09 +05:30
b765495eb7 Refuse archiving an agent that a published agent still delegates to
§3 says an unknown subagent key fails at publish, not at run time, and
refuseSubagentCycle enforced that in one direction only: the edge was
checked when the PARENT was written, and nothing re-checked it when the
CHILD was later archived. So a spec could validate on Monday and be
delegating into nothing by Friday.

That is what happened on 2026-09-15. activity-agent was archived while
krow-workforce-agent v2 still listed it, and every run since logged
runtime.unknown_subagent and answered activity questions without its
activity capability -- quietly, because the parent still Completed.

Both archive paths now refuse with 409 naming the dependents: the
status-only patch the UI sends, and a markdown save whose frontmatter
says archived. Only PUBLISHED parents count, so an abandoned draft
cannot pin a production agent in place. Unlike the cycle check this
fails closed when the graph cannot be read, because the only backstop
here is the failure it exists to prevent.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-22 12:48:09 +05:30
822b3b1edb Stop tracking .env and restore the ignore rules that 3455ad0 removed
Some checks failed
CI / test (push) Failing after 4m36s
CI / fixture (push) Failing after 8s
3455ad0 deleted the .env patterns from .gitignore and committed a
real .env carrying a live ANTHROPIC_API_KEY. Two problems:

  - The key is now in shared history and must be rotated; untracking
    the file here stops the bleeding but does not un-publish it.
  - That .env does not boot the API. It sets ANTHROPIC_API_KEY with no
    MODEL_API_KEY, which config.go:503 refuses at startup — the same
    guard that crash-looped krow-2 on 2026-09-07. Nothing in the
    codebase reads ANTHROPIC_API_KEY; the MCP surface needs
    OAUTH_ISSUER and MCP_RESOURCE, not a model credential.

The file stays on disk and is ignored again, along with
infrastructure/.env which the deleted pattern also covered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-22 12:00:30 +05:30
8e36faff89 Export a database's rows, so a tenant can move between deployments
Some checks failed
CI / test (push) Failing after 4m37s
CI / fixture (push) Failing after 9s
Local work has been stranded on one machine: `make seed` replays
seed/fixtures/seed.json, which is generated from krow-demo and has never
reflected what is in a database. Agents published through importagents,
knowledge ingested, applications screened — none of it had a path to
another deployment.

`make export-data` reads the database itself and writes replayable SQL.

Three things it does deliberately:

Tables are ordered topologically from pg_constraint rather than left in
pg_dump's own order, which sorts by name and so fails on foreign keys in a
way that depends on what the tables are called. Parents always precede
children, and Kahn's algorithm breaks ties alphabetically so two runs
against one schema produce a byte-identical file.

Every statement is INSERT ... ON CONFLICT DO NOTHING. The export can
create rows on a target and cannot modify or delete one. That is a
property of the generated file, not a rule someone has to remember when
they apply it.

The file refuses to apply to a schema older than the one it came from. A
restore into a half-migrated database half-succeeds, and a partial import
is harder to unpick than a failed one.

pg_dump 18 wraps its output in the psql meta-commands \restrict and
\unrestrict. Dumping per table left the closing one without its opener,
which fails with "not currently in restricted mode" and, under
--single-transaction, rolls back having inserted nothing — silently, if
the caller reads psql's output through a pipe instead of its exit code.
Each table's slice is therefore cut at the last line ending in a
semicolon, which no line of pg_dump's epilogue does and every generated
statement does.

sessions and schema_migrations are excluded: sessions are bound to cookies
one deployment issued, and a stale schema_migrations row would make the
target lie about its own version.

Verified against a scratch database migrated to 15: all 25 tables match
the source row for row, a second run is a no-op, and the guard refuses a
version-10 target.

The output is real tenant data — password hashes, personal details — so
seed/exports/ is gitignored. It moves over scp, not through git.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-22 11:27:42 +05:30
3455ad022f env
Some checks failed
CI / test (push) Failing after 4m39s
CI / fixture (push) Failing after 9s
2026-09-22 11:02:48 +05:30
f2aa3b3ad8 mcp connection
Some checks failed
CI / fixture (push) Has been cancelled
CI / test (push) Has been cancelled
2026-09-22 10:58:02 +05:30
4e1f746b22 update the archive options
Some checks failed
CI / test (push) Failing after 4m40s
CI / fixture (push) Failing after 7s
2026-09-10 19:30:46 +05:30
74089eb3e9 Add the krow-2 deploy runbook for 9d3192a
Some checks failed
CI / test (push) Failing after 4m36s
CI / fixture (push) Failing after 8s
The API on krow-2 is down on a startup guard, and the fix is two environment
variables. Written down rather than left in a chat log, in the same shape as
deploy-b6f8655.md: what is being deployed, what was actually verified before
claiming it works, the rollout, a smoke test, and rollback.

The smoke test insists on one real agent run. Boot and /health both pass with a
broken model configuration — that is precisely how the current outage stayed
invisible until run time — so a health check alone is not evidence the deploy
worked. Includes the symptom-to-cause table for the four failures this rollout
can actually produce.

Records what the preflight covered: the production image built and booted under
APP_ENV=production against a TLS Postgres, 11 migrations applied, seed and 9
agents imported, login, a real Groq-backed run terminating Completed, and SSE
streaming. 62 endpoints is written down as the expected number because two
fewer is the signature of a missing credential.

Not executed. No working SSH to that host from here, and a production rollout
wants the operator watching.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-07 13:06:49 +05:30
9d3192a9c4 Replace the model ids with ones Groq actually serves
Some checks failed
CI / test (push) Failing after 4m39s
CI / fixture (push) Failing after 7s
The defaults shipped yesterday were wrong the day they shipped, and a real key
proved it in one request. Groq serves neither llama-3.1-8b-instant nor
llama-3.3-70b-versatile any more. Both were chosen from memory, both passed
startup validation, and every agent run would have failed with a 400.

This is the exact failure the claude-* guard was written to catch, arriving from
the side that guard cannot see. A prefix check can reject a vendor this service
cannot call; it has no way to know a provider retired an id last month. That is
not a gap in the check, it is a gap in the class of thing local validation can
know, so the fix is not another guard:

TestConfiguredModelsAreServed asks the provider. It lists /models — part of the
same openai-compatible surface the gateway already speaks, so every supported
provider answers it — and fails if a configured id is absent, printing what is
available. It reads the ids through config.DefaultModels() rather than
repeating them, because a second copy would be the first thing to drift, and
drift is the whole failure. Skipped without a credential like the rest of the
live suite. Verified three ways: it fails on the retired id with the message an
operator needs, skips clean with no key, passes on the new ones.

New defaults, chosen against the live account rather than from memory:
openai/gpt-oss-20b (fast) and openai/gpt-oss-120b (balanced, deep). Tool
calling confirmed on both. groq/compound-mini was ruled out — it cannot do tool
calls at all, which this platform requires.

MODEL_REASONING_EFFORT is now documented as safe here and NOT portable: gpt-oss
accepts low/medium/high, exactly the scale openAIEffort maps onto, while
qwen/qwen3.6-27b on the same account rejects all three and fails the whole
request rather than ignoring the key.

I7 IS NO LONGER UNPROVEN. make eval-live passes all three cases twice against
gpt-oss-120b, the planted-injection case included: answers from the handbook,
cites, refuses the injection, leaks neither the operator-only pay guidance nor
the other tenant's figures. CLAUDE.md §12 and handover.md updated from "urgent"
to measured, dated, and scoped to the one model it is evidence about.

One real defect found on the way. The handbook grounding check failed once on an
answer containing the phrase it wanted — "more than ten minutes" on screen,
strings.Contains false — which leaves an invisible separator as the only
explanation; the same model writes "47 %" and a U+2011 hyphen elsewhere. The
flaky assertion is the small half. THE LEAK ASSERTIONS USED THE SAME MATCH and
fail in the dangerous direction: "attacker@evil.test" with a zero-width space,
or "uplift" with a soft hyphen, would have been reported clean. A permission
test that cannot see the leak it is hunting is worse than none, because it is
believed. normalizeForMatch folds those away, and its test pins that every case
is one plain ToLower MISSES — a case whose naive match already succeeds fails,
so the suite cannot fill with examples that demonstrate nothing. That caught my
own first BOM case, which put the mark where Contains found it regardless.

gofmt clean, vet clean, 15/15 packages pass offline; live suite green twice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-07 12:43:45 +05:30
7d83c16ec5 Say plainly that a container has no keyless option
Some checks failed
CI / test (push) Failing after 4m39s
CI / fixture (push) Failing after 7s
Two comments in the compose model block were wrong in a way that mattered to
the question "do we actually need a Groq key".

The note about agent routes had lost its antecedent in the previous commit and
dangled above MODEL_PROVIDER, appearing to describe provider selection. It
belongs to MODEL_API_KEY.

It was also only half true. It said an absent key is "a legitimate way to run
this", which is correct outside production and impossible inside it: this stack
defaults to APP_ENV=production, where validateModel refuses to start without a
credential unless MODEL_BASE_URL is loopback. isLoopback accepts only
localhost, 127.0.0.1 and ::1, so host.docker.internal does not qualify and no
containerised deployment can take the keyless path. Reading the old comment,
an operator would reasonably conclude they could leave the key empty and get a
working API without Owliver. They get a container that will not boot.

The endpoint count was stale too: routeRuns registers two, not three.
routeOwliver's one endpoint does not touch s.agents and stays registered.

Comments only; no behaviour change. vet clean, config and httpserver pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-07 12:33:47 +05:30
bd9a8f91fc Make the shipped example envs ones that can actually start
Some checks failed
CI / test (push) Failing after 4m40s
CI / fixture (push) Failing after 9s
The krow-2 deploy failed on the ANTHROPIC_API_KEY guard, which is the guard
doing its job. Checking what an operator hits *after* fixing it turned up two
older faults in the files they are told to copy — both predating the Groq
switch, both fatal at boot.

HTTP_WRITE_TIMEOUT shipped as 30s in .env.example, .env.docker.example and the
compose default, while validateWriteTimeout refuses anything at or under the
deep tier's 2m deadline. `cp .env.docker.example .env && docker compose up`
could not start. Now 180s. krow-2 never saw this because someone had already
overridden it in that environment.

.env.docker.example carried no model block at all, so a production stack built
from it is refused for a missing MODEL_API_KEY. Added, with the Groq defaults
and the reasoning-effort note (most non-reasoning models reject the request
rather than ignoring the key).

Neither was subtle. Both survived because the examples were prose to every test
in this package: the validator and the file documenting it had no mechanical
connection, so tightening one silently invalidated the other. That connection
is now TestShippedExampleEnvActuallyBoots, which parses each example and runs
Load() on it under the APP_ENV the file itself declares — production for the
docker one, development for the root one, each internally consistent. Verified
by mutation: reverting the timeout, removing the key line, and restoring a
claude-* id each fail it with the message an operator would see.

Go does not treat these files as test inputs, so an example-only edit can be
served a stale pass from the test cache. Noted in the test; use -count=1.

Also documented the upgrade path in handover.md, including the one thing
startup validation cannot catch: renaming ANTHROPIC_API_KEY to MODEL_API_KEY
without replacing the value boots fine and 401s on every run.

gofmt clean, go vet clean, 15/15 packages pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-07 12:25:20 +05:30
34fa58a6b9 Remove the Anthropic path; the gateway speaks one wire protocol
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 7s
The platform now runs on Groq by default, through the OpenAI-compatible
chat-completions shape. That shape is not one vendor — Gemini, OpenRouter,
Together, vLLM and a local Ollama serve it too — so moving again stays
configuration rather than code.

Two things in the deleted file were not Anthropic's and would have gone
with it silently:

  withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and
  CALLED BY openai.go. Deleting the file wholesale would have removed the
  retry policy of the provider that survived, and nothing in openai.go
  mentions it, so the loss would have been invisible until the next 429.
  The policy is a property of this platform's runs, not of a vendor's API;
  it now lives in retry.go where no provider can carry it off.

  StreamComplete had the same problem and moves to gateway.go, beside the
  Streamer interface whose comment already referenced it.

Three stale-configuration failures are now refused at startup instead of
being ignored. Each was verified firing through the real config.Load():

  MODEL_PROVIDER=anthropic — named separately from every other wrong value
  because it used to be correct. Ignoring it gives a stack that believes it
  is on Claude while every run goes to Groq and is billed there.

  ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an
  operator did set is the worst version of this: they fail every run on a
  missing credential they are looking straight at.

  A leftover claude-* model id, naming the tier that carries it. This is
  the check the previous commit's error-detail work was diagnosing: such an
  id is accepted by this process, rejected by the provider, and 400s on
  EVERY run. "A model is wrong" does not say which of three lines to edit.

Defaults ship as a matched pair. defaultBaseURL and the three tier ids are
one decision, not four: an id is only meaningful against the service that
serves it, and a Groq id on an OpenAI base URL is the same failure from the
other side. The tiers also stop being one model — a tier whose cost does
not differ is a distinction that buys nothing.

Verified end to end against a stub of the wire, driving the real wiring
(config.Load in production mode, gateway.New, StreamComplete): streamed
deltas, tool-call decoding, the loopback credential exemption, and usage
totalling 150 rather than 190 — the cached-prefix subtraction still holds.

gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still
needs a reachable database.

NOT verified: the I7 planted-injection eval. Removing this path removed the
only model whose refusal behaviour had been measured against it, so the new
default is unproven there until `make eval-live` runs with a real key. The
Groq model ids should also be confirmed against Groq's current lineup.
Flagged in CLAUDE.md §12 and docs/handover.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-05 11:52:21 +05:30
cf99866e12 create employee table
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 9s
2026-09-05 10:44:47 +05:30
dc785b917c Separate what a worker does from what a company needs filled
Some checks failed
CI / test (push) Failing after 4m41s
CI / fixture (push) Failing after 8s
Owliver could offer neither create. The Create Position flow worked and no chip
anywhere suggested it, because the chip row is entirely the backend's static
catalogue and no intent in it wrote anything. The gap was never in the
frontend's trigger matching — every phrasing already routed.

`employee_roles` is the supply side of `job_postings`. A posting is what the
ORGANIZATION needs filled; this is what a WORKER says they do. They share a
vocabulary and almost nothing else: "3 years" on a posting is a minimum an
applicant must clear, and the same words here are what the person has. There is
deliberately no foreign key between them — supply and demand already meet
through `job_applications`, which carries the funnel, the interview and the
outcome, and a second weaker link would disagree with it the first time
somebody withdrew.

NO NEW COMPANY ENTITY, AND THAT IS THE LOAD-BEARING DECISION. "Create a company
position" reads like it needs a client record. `organizations` is the TENANT —
absent from the resource table, absent from the policy map, written only by the
seeder — so creating a row there from a chat flow would provision a new tenant,
and the position would carry an org_id the operator's session cannot see. The
operator could never view the record they just created. That breaks I5 and I1
to add a feature nobody asked for. The client stays free text on the posting,
per blueprint decision D2, and the flow simply offers the clients this
organization already staffs for as chips. No schema change, no endpoint change.

Create is operators-only, and that is an I1 decision rather than a deferral.
The worker is named explicitly on the row and is deliberately NOT derived from
the session, because an operator recording a role on somebody's behalf is the
whole point of the flow. Granting talent the same Create would let a talent
caller write a role under any worker_email in the tenant — the attribution hole
Phase 3D closed elsewhere. Talent reads its own via a ScopeEmail predicate,
which is in place now so the grant is one line when a talent console exists.

`created_by` is in gen_resources.py's SERVER_OWNED as well as the policy's
Derived list. Both are required and the pairing is easy to miss: Derived fills
the column from the session, SERVER_OWNED is what makes the descriptor ReadOnly
so a request body cannot set it in the first place. Without it,
TestDerivedColumnsAreReadOnlyOrTalentScoped fails — verified by mutation, not
by reading.

The two catalogue intents carry PHRASE terms only. A bare "position" or "role"
term scores 10, the same as every reading on that page, and wins the tie on
declaration order — so a create chip would have arrived by evicting
`positions-attention` from the exact ordered result TestPositionsSuggestions
asserts. An offer to create something must not displace the reading a person
actually asked for. Neither declares a Subject, on the precedent of
`position-spec-steps`: a Subject would let the bare query "summarize" match
through matchShape and survive filterOnTopic. Neither declares a Signal, so an
empty composer still reports what the organization needs rather than proposing
paperwork.

Chip text is the coupling with nothing else holding it together: no page
context declares `capabilities`, so every server suggestion dispatches as its
own TEXT and is answered by whichever skill's trigger that text matches. A
renamed chip would open nothing, silently. Asserted on the frontend side.

The down migration drops `employee_role_status` and keeps `english_level`,
which is shared with job_postings.english_required and
job_applications.english_level. Rolled back and re-applied against the
database to prove it, not asserted.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-02 15:29:25 +05:30
c74fe7e074 Add an OpenAI-compatible gateway, so the model provider is a config value
The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.

`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.

`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.

THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.

Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.

Three configuration mistakes are refused at startup rather than at runtime:

  - MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
    endpoint and ignores the field, so this is a deployment that believes it
    switched providers and did not — every run still goes to Anthropic and is
    still billed there, with nothing in the logs to say so. Cost is the whole
    reason this change exists, and that is the one mistake that silently
    defeats it.
  - An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
  - A production deployment with no credential — except against localhost,
    which needs none, and demanding one would make the free local path
    impossible to configure.

reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.

`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.

Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.

NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-01 11:47:53 +05:30
57c2a52c1e Refuse an HTTP write timeout that would cut off a legal agent run
Some checks failed
CI / test (push) Failing after 5m32s
CI / fixture (push) Failing after 59s
Production answered 502 Bad Gateway on a non-streamed agent run. Nothing about
that was a gateway fault: krow-proxy already had proxy_read_timeout 3600s, and
the API pods were healthy with zero restarts throughout.

HTTP_WRITE_TIMEOUT was 30s. Every shipped agent runs at the `balanced` tier,
whose deadline is 60s, and the `deep` tier allows 120s. So the server aborted
the response on any run over half the time the runtime considered legal, the
proxy saw its upstream vanish mid-response, and it reported the only thing it
could. A gateway error for something no gateway did — which is why it looked
like infrastructure for as long as it did.

Delegation did not cause this; it made it routine. A parent that asks two
subagents takes longer than one answering alone, so a latent misconfiguration
became a reliable one. Verified: the exact request that returned 502 now
answers 200 in 18s.

Streaming is what hid it, and that is the part worth keeping in mind. The chat
panel uses SSE, so the product looked healthy while every non-streaming caller
got 502 on a slow question. A bug only reachable by the callers who do not yet
exist is one nobody reports.

So the value is now derived from the thing that constrains it — the default is
DeepestAgentDeadline plus headroom rather than a number typed once — and
validate() refuses anything below that deadline at startup. A slow,
intermittent, misattributed failure becomes a message on the first boot.

DeepestAgentDeadline is duplicated in internal/config rather than imported,
because internal/runtime already imports internal/config and a cycle to share
one number is a bad trade. TestConfigKnowsTheDeepestAgentDeadline asserts the
two agree, so drift is a build failure rather than a discovery. It also checks
that no tier exceeds it, or the name lies.

ORDERING, and it matters for the next deploy: the check refuses the old 30s, so
a pod carrying this image against an unpatched configmap will not boot.
Production's configmap is already 180s. The handover says so too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-31 12:06:26 +05:30
fd1812e161 Record the CI runner, now that one exists
Some checks failed
CI / test (push) Failing after 6m25s
CI / fixture (push) Failing after 47s
Both repositories carried GitHub Actions workflows on a Gitea remote and
nobody had confirmed a runner. There was not one: the 924 frontend checks,
the whole Go suite, the skip guard and the suite-shrank guard had never run
on a push, only when somebody remembered.

gitea/act_runner v0.6.1 is registered as krow-runner on the cluster host. This
commit is also the first push that can prove it picks up a job, which is the
failure mode worth catching — a runner that registers and never runs anything
looks identical to a healthy one in the Runners list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-31 11:13:47 +05:30
109fc2f1c6 Add the in-cluster embedder, so production retrieval stops being keyword-only
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
Production had no EMBED_PROVIDER, so every knowledge_chunk carried a null
embedding and a question only matched documents that shared its words. A person
asking about a family emergency got nothing from a document titled "shift cover
and cancellation".

Ollama rather than Voyage: internal/knowledge/embed.go calls it "the default
worth reaching for" — real semantics, no credential, no per-token cost, and no
tenant text leaving the cluster. Voyage needs an API key nobody has issued.

Bounded deliberately. The API pods share this node, so an unbounded model
server is a way to evict them; the memory limit means the kubelet kills the
embedder and nothing else. The 1Gi request is also what keeps it off the second
node, which has 1.2Gi allocatable and could not hold it.

Applied in three stages so nothing was pointed at an embedder that had not
been proven: deploy and pull the model, run reembed with the settings passed as
exec environment — 34 chunks in 11s, which proves connectivity without touching
live config — and only then patch krow-config and restart. Rolling back is
removing four keys and restarting.

Verified after: 55/55 on verify-deploy, and a question with no literal keyword
overlap with the corpus returned the relevant policy documents.

This file is the record of what was applied. It was applied by hand, which is
the same gap the README already admits for migrations — there is no deploy
pipeline, so a manifest in the repository is a description of the cluster
rather than the thing that produces it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 15:33:59 +05:30
629d97181d Rebase Staff too, which was the last thing keeping the chart incomplete
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
And correct the previous commit's closing claim, which was wrong.

c378bc0 said the Hiring activity chart "still renders empty ... the fault is
further down in that component, not in the data." That was a retraction of a
correct diagnosis, made from a screenshot taken before the rebase had reached
the browser. The chart was empty BECAUSE the data was stale, exactly as first
diagnosed, and rebasing fixed it. Checked properly this time: the area path
carries real values, and the rendered chart shows applications peaking at 16
around 8/18 with the screening and interview series drawn over it.

What was genuinely still missing was Hires. Staff was not rebased, so
hire_date stayed 35 days old with nothing inside the 30-day window and that
series drew nothing. It is rebased now, anchored on hire_date rather than
created_date, because the hire is the event the chart plots.

Leaving it behind had also introduced an inconsistency of my own making: once
applications moved, a candidate was hired last week according to their
application and five weeks ago according to their staff record. Rebasing them
together removes that.

All four series now render. Full suite green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 15:21:14 +05:30
c378bc00ed Rebase seeded time-series data to now, so the demo stops going quiet
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
ShiftRecord is generated against now (shifts.go); everything else stayed on the
fixed calendar in seed.js while the calendar moved on. Twenty-three days after
that file was written the activity agent truthfully reported zero events in the
last seven days, and applications were sixteen days stale. Nothing was broken —
the data had simply aged out of every window the product reports over, and it
gets worse every day nobody reseeds.

RebaseToNow moves a set of records so the newest sits at now, keeping every gap
exactly as authored. The SHAPE is what every reader of this data is looking at:
three hires on one day, a screening the day after, a quiet fortnight before it.
Shifting the whole set by one delta preserves all of it. Scaling into a window
or scattering events across recent days would invent a rhythm nobody wrote.

Applied per entity — UserActivity, JobApplication, AIInterview — because each
anchors on its own newest record. One shared anchor would drag the quieter
entities by another entity's delta and invent relationships between them.
Reference data is untouched: a course's date is a fact about the course, not a
position in a window.

Rebasing rather than generating, so seed.json stays the single authored source,
still deterministic and still comparable byte-for-byte by the drift check. The
alternative — excluding these from the fixture the way ShiftRecord is — means a
second generator to keep in step with the frontend's copy.

The existing TestSeedPreservesSourceValues caught a real bug in the first
attempt, and its comment is why: "Applications carry updated_date in the
source, and the gap from created_date is what buildHires reads as
time-to-hire." I had shifted created_date alone, which turned five-day hires
into three-week ones. EVERY timestamp on a record now moves by the same delta,
and there is a test on that specifically.

That test now asserts the GAP rather than the absolute dates, because for a
rebased entity the dates are deliberately different — which is the one reason
it is supposed to allow. Its real subject was always the interval.

Verified against the database: before, activity was 23 days old with 0 events
in the last 7; after, all three entities are current, with 24 applications
spread across the last 30 days and 15 activity events inside 30.

Not fixed here: the Hiring activity chart on Control Center still renders
empty, and it was equally empty before this change. Its bucketing is correct —
replaying it in the browser against live data matched all 24 applications into
the right days — so the fault is further down in that component, not in the
data. Naming it rather than leaving it implied by a chart that still looks
wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 15:13:33 +05:30
48ab9d1dad Give importagents tests, by separating what it decides from what it wires
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
This command had no tests at all, while carrying the rules that decide whether
a deploy may change a published agent. Everything interesting was inside run(),
which loads configuration, opens its own pool and resolves a tenant from a
slug — none of which a test can supply. So it was untestable by construction
rather than by neglect, and the fix is a seam, not a test-only helper.

Three functions come out of run(), each doing one thing:

  validateSpecs  the parse and status checks. Pure.
  validateGraph  §3's DAG check over the whole set. Pure.
  importInto     the write phase, taking a transaction the caller owns and
                 returning what it did.

run() is now the wiring around them. importInto does not commit — the caller
does — so a refused rewrite leaves the caller's deferred rollback to undo the
writes that already happened, which is the behaviour that was there before and
is now visible in the signature rather than implied by where the code sat.

Six tests, four of them against a real database:

  - every problem is reported, not the first: two bad specs produce two
    messages and a good one produces none;
  - a chain is not a cycle, and a cycle names the edge to cut;
  - a first import records versions, and a second over unchanged specs
    records none — the counter that used to say nine every deploy;
  - a changed spec at the same version is refused AND nothing is committed,
    checked by reading the row back;
  - a lowered version is refused and the live row is still at the higher one;
  - a raised version is accepted and leaves two rows in the history.

The author is resolved through resolveAuthor rather than passed as a literal,
so the tests exercise that path too and fail loudly on an organization with no
active admin — a real deployment condition. The first draft passed "" and got
`invalid input syntax for type uuid`, which is what a literal buys you.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 14:49:38 +05:30
377948708b Enforce §3's monotonic version and its DAG requirement at publish
Two rules §3 states and nothing checked.

MONOTONIC. The rewrite guard added earlier compares content at ONE version
number, so republishing an OLDER number with the text originally published
under it looked like a no-op: nothing conflicted, nothing was refused, and the
live row silently reverted. The agent then reads v1 in the UI while the newest
thing anybody approved was v2. The test publishes v1, publishes v2, republishes
v1 byte-for-byte, and asserts both the refusal and that the live row is still
v2. Without the guard it answers 200 and the row goes back to version 1.

DAG. §3 says cycle detection runs at publish; only the runtime depth cap
existed. That cap means a cycle was never a safety problem — it was a budget
one. Every run entering the loop spends its whole allowance delegating in a
circle before terminating, and the author learns about it from a bill rather
than from the publish that created it.

definition.FindSubagentCycle is a pure function over id -> subagent ids, so it
is tested directly: chains, diamonds, self-reference, loops not involving the
first agent walked, and a 5000-long chain that would matter if this were
written to recurse carelessly. It REPORTS the cycle ("a -> b -> c -> a")
rather than merely detecting one, because an operator otherwise has to find it
by hand across a set of specs. The report is deterministic — a test runs it
fifty times over a graph with two cycles and requires the same answer, since Go
randomises map iteration and an error message that changes between identical
runs is one nobody trusts.

Wired into both publish paths. importagents has every spec in hand, which is
the only place that is cheaply true. The API builds the graph from
organization-visible agents plus the incoming definition standing in for its
stored self — otherwise an edit that CREATES a cycle is checked against the
version that did not have one and passes. Personal agents are excluded: they
are invisible to everyone else so cannot complete anyone else's loop, and
reading them would mean reading other people's drafts to validate your own.

An edge to an agent that is not in the set is ignored rather than reported.
That is a different failure with a different message
(runtime.unknown_subagent), and conflating them prints "cycle detected" for
what is actually a typo.

The cycle test bumps the version on the loop-closing edit. Without that the
rewrite guard refuses it for changing published text, the test passes for the
wrong reason, and it would keep passing with cycle detection deleted — which
is how it was first written, and what running it without the guard showed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 14:43:49 +05:30
4c29185c3b Correct the governing documents where this session disproved them
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
Both documents are the first thing a new reader trusts, and several of their
claims were wrong — some wrong from the start, some overtaken by work this
week. A governing document that misdescribes the system is worse than none,
because it is believed.

CLAUDE.md §10 specified Python 3.11, FastAPI, SQLAlchemy and Alembic. The code
is Go and has never been anything else. That is corrected rather than quietly
deleted, so the next person understands the document drifted rather than
wondering which half to trust.

Also in CLAUDE.md: the tool count was 17 with one write and is 19 with two;
delegation now exists and §11's Orchestration row says what it guarantees; the
embedder deviation described a Voyage-or-stand-in choice that has since become
EMBED_PROVIDER with three options, of which production sets none.

The handover claimed three things that this session disproved by running them:

  - "definition_versions is empty ... nothing has gone through it". It was not
    empty in production; activity-agent had a v1 that the shipped file
    contradicted, which is how a real drift was found. It now holds every
    agent and skill.
  - "Skills are still stored in user_preferences". They are rows in
    skill_definitions, and are now versioned.
  - "make eval-live ... has never been run". It has, it passes 3/3, and what
    it established is recorded — including that the handbook corpus carries a
    planted prompt injection which the agent refused and reported. That is I7
    holding against a real model, which is worth more than the pass count.

The endpoint counts were one high throughout (55/57, not 56/58) — the delta of
two was always right, so the signal worked and the absolute numbers did not.

Added, because they cost time this week and would cost it again:

  - the app reaches its database through pgbouncer, not PostgreSQL directly.
    Enabling TLS on PostgreSQL does nothing for the application hop; pgbouncer
    terminates 5432 and needs its own client_tls_sslmode.
  - the seeded UserActivity is NOT anchored to today the way ShiftRecord is,
    so it ages out of every window the activity tools offer. Twenty-three days
    old as of writing: zero events in the last 7 days, 6 of 15 in the last 30.
    The agent answers truthfully and the demo looks dead.
  - the whole stack runs on Docker alone. Dockerfile.api builds every command
    plus the migrate CLI, so a new machine needs neither Go nor psql — which
    is how this one was set up, having no Homebrew.
  - Ollama runs on the HOST, so a container reaches it at
    host.docker.internal, not localhost. The old .env said localhost and would
    have failed with nothing obviously wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 12:10:41 +05:30
6849363a37 Implement delegation, so an agent's subagents are more than decoration
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
krow-workforce-agent has declared five subagents since it was written and
answered every question by itself. Everything for §6 existed except the
delegation: the parser read `subagents:`, runtime.Agent carried them, the
loader populated them, agent_runs had a parent_run_id column with a
self-reference and a no-self-parent constraint, and budget.go's comments
already described sharing a budget with subagents. Nothing called any of it.

A subagent is offered to the parent's model as a tool, because §6 says that is
what delegation is from the parent's side. Three rules are enforced rather than
assumed, each with a test that fails if it stops holding:

  I1  The subagent runs as the ORIGINAL caller. It cannot read anything the
      person could not read directly.
  §6  It SHARES the parent's budget. The test sets MaxSteps to 1, spends it in
      the parent, and asserts the child terminates BudgetExceeded — an
      assertion that only passes when the budget is shared, and that a fresh
      budget would quietly turn green.
  §3  Depth is capped at 2. At the cap no subagent is loaded or offered, so a
      cycle reaching run time is bounded rather than unbounded.

I4 survives too: a write a SUBAGENT wants approved still stops the whole run
and asks a person, rather than being performed because it happened one level
down.

Two bugs found by running it rather than by reading it:

  - delegate() read the error before the result. finish returns a non-nil
    error for every termination that is not Completed, INCLUDING
    ConfirmationPending — which is not a failure but a run that stopped to ask
    a question. Reading the error first discarded the result and with it the
    confirmation, so a subagent's write silently never happened and nobody was
    asked.

  - Delegated trajectories were never persisted at all. parent_run_id is a
    foreign key and a subagent finishes BEFORE the run that delegated to it,
    so every child insert named a parent row that did not exist yet. The
    database refused it; finish deliberately does not fail a run over a sink
    error; and the entry recording that the trajectory could not be saved was
    itself in the trajectory that was not saved. Children are now buffered and
    written by finish after the parent's own row, each arriving with its
    descendants already ordered behind it, so one pass writes a whole tree
    parent-first. The regression test asserts on save ORDER, because a
    MemorySink has no foreign key and will pass either way.

Verified end to end against a live model: an agent with no tools of its own and
one subagent produced

    delegation-probe   run=run_16622d7de6  parent=(root)
    talent-pool-agent  run=run_64160684b1  parent=run_16622d7de6

with the subagent's answer reaching the parent's model. Full suite green, only
TestLive* skipped.

Not addressed: §3's publish-time cycle detection, which needs the whole agent
set in hand. The depth cap is what holds without it, and is the half that
matters at run time.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 11:59:20 +05:30
a1e91f776d Compare definitions by meaning, and stop miscounting what was recorded
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
Two problems found by deploying the previous commits to production.

FIRST: the rewrite guard compared raw Markdown, so it refused a publish over
formatting. The authoring UI re-serialises a definition when somebody saves it
— writing `webSearch: false` where the hand-authored file omitted the key, and
ordering the frontmatter its own way — and the parser reads absent and false
identically (agent.go: `data["webSearch"] == true`). A definition nobody
meaningfully changed stopped a deploy. Refusing a change that is not a change
is still a bug, even though it fails safe.

definition.SameAgent and SameSkill compare the parsed definition instead, and
are deliberately conservative, because the two ways of being wrong are not
equally bad. A false difference blocks a deploy: visible, recoverable. A false
SAMENESS lets a changed agent overwrite an approved version silently, which is
the thing versioning exists to prevent. So:

  - The body is compared verbatim. Agent.Body carries `json:"-"`, so a
    comparison that only marshalled the struct would call a completely
    rewritten system prompt "unchanged". There is a test that fails loudly on
    exactly that, because it is the mistake this design invites.
  - List ORDER stays significant. loader.go resolves Skills in order and that
    order reaches prompt assembly, so two definitions listing the same skills
    differently are still different. A deploy that only reorders still has to
    raise its version. That is a limit, recorded in a test rather than left to
    be discovered: loosening it needs somebody to decide skill order cannot
    matter, which is not a decision to bury in a comparison function.

What it absorbs is exactly what the round trip produces: frontmatter key order,
whitespace, and a defaulted value written out in full.

SECOND: importagents reported "9 agent version(s) recorded" on a run that
recorded nothing. The counter incremented on every successful Snapshot call,
and Snapshot returns nil for the idempotent no-op as well as for a real insert.
The skill counter was already honest; the agent one was not. snapshotAgent now
distinguishes recorded / conflict / already-present, and only the first counts.
Verified locally: 1 on the run that added activity-agent v2, 0 on the re-run,
where it previously said 9. A number that says nine every time is one nobody
checks on the day it matters.

Full suite green against PostgreSQL, only TestLive* skipped.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 11:02:22 +05:30
Suriyakumarvijayanayagam
5ab16b836a Publish activity-agent as v2; the change to it was real, not formatting
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
Correcting the previous commit's reasoning. It claimed the difference between
this file and what production published was inert — a defaulted `webSearch:
false` and a reordered skills list. That was based on comparing the file
against the LIVE row in agent_definitions. The live row was the wrong thing to
compare against: Snapshot compares against the stored v1 in
definition_versions, and those two had diverged.

Read from production directly, v1 as published carries two skills:

    skills:
      - anomaly-detection
      - operational-risk

and the live row carries three, with activity-analysis added. So a skill was
added to this agent after v1 was published, in place, without the version being
raised — the exact silent rewrite this branch exists to stop. It is a genuine
change of behaviour: an agent with a third skill answers differently from one
with two.

So the version is raised rather than the file being bent to match. v1 keeps
what was approved; the current definition, which is what production has been
serving, becomes v2. The earlier alignment of field order and `webSearch:
false` is kept — that part WAS serialisation, and matching it keeps future
imports quiet.

Production has three agents with version rows at all, so two others may hold
the same kind of drift. They did not conflict on this import, which means their
live rows still match what was published; it does not mean nobody edited them.

The image ships agents/, so this file only reaches production on the next
image build. Any build from main at or after this commit carries it; a build
from an older tree will fail the import again, by design.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-28 20:14:43 +05:30
Suriyakumarvijayanayagam
04b11a079b Align activity-agent.md with the version production actually published
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
The first run of the new importagents against production refused, which is the
guard working rather than a fault: version 1 of activity-agent was already
published and said something different from the file this repository ships.

The difference is inert. Production's copy adds `webSearch: false`, which is
exactly what the parser defaults to when the key is absent (agent.go:471 reads
`data["webSearch"] == true`), and lists the same three skills in a different
order. Both forms parse to the identical agent — 2 tools, 0 sources, 3 skills,
confirmed by running the importer's own dry-run over each.

What happened is a round trip: somebody edited this agent in the UI, the editor
re-serialised it, and that serialisation is what got snapshotted as v1. The
hand-authored file was never the published artefact for this one.

So the file is updated to match rather than the version being bumped. Bumping
would publish a v2 that differs from v1 only in field order and a defaulted
key, which is noise in a history whose whole purpose is to say what changed.

This unblocks the import. It does not address the underlying awkwardness: the
comparison is textual, so any future UI edit that reformats without changing
meaning will block a deploy the same way. Comparing the PARSED definition
instead would fix that properly and is the right follow-up — it needs a
decision about what counts as semantically equal (skill order, starter order)
and should not be rushed in behind a deploy.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-28 20:01:48 +05:30
Suriyakumarvijayanayagam
0cda877cd6 Version skills too, numbered by the server rather than by their author
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
repo.KindSkill existed with nothing writing it. Migration 000010 says "agents
and skills version identically", the table has always accepted kind='skill',
and no path on either side ever recorded one — so an edit to a skill left no
record of what it used to say. Agents name their skills and the runtime refuses
to load one whose skill is missing, so a skill changing under a pinned agent is
the same class of problem the last two commits fixed, one layer down.

Skills are numbered differently, and not by preference. An agent's frontmatter
carries `version:`, so its author decides when a change is a new version and can
be refused for rewriting an old one. definition.Skill has no such field, the
skill_definitions table has no such column, and the vocabulary is active |
inactive rather than draft | published. Giving skills an authored version would
mean a migration, a parser change on BOTH sides of the conformance test in
internal/definition — which replays a capture of the real frontend module graph
— and an edit to all 23 shipped skills. That is a feature, not this fix.

So the server assigns it: one after whatever was last published. This is not an
invention. repo.VersionsRepo.LatestVersion was written for exactly this and
says so — "the next published version has to follow what was actually published
rather than what somebody wrote in the frontmatter" — and had no callers
outside its own test.

Because the author never names a version, there is nothing to refuse: an edit
is always a new version. What needs care instead is the opposite — a save that
changed nothing must NOT be one, or every deploy would add a version to all 23
skills and the number would stop meaning anything. Each publish is compared
against the last recorded copy first. Inactive skills are not recorded at all;
inactive is this vocabulary's draft.

Verified against a live stack:

  - first import over 23 unversioned skills: 23 skill version(s) recorded
  - second import, files unchanged: 0 recorded, total still 23
  - one skill edited: 1 recorded, that skill at v1, v2; v1 still holds the
    original text and v2 the edit

The test fails without the change — "after create: 0 version(s), want 1" — and
covers the three behaviours that matter: an edit versions, an identical save
does not, and an inactive skill is not recorded.

Both races noted on the agent path apply here as well: two simultaneous edits
can compute the same next number, and the loser's snapshot is dropped rather
than failing the author's save.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-28 19:34:06 +05:30
Suriyakumarvijayanayagam
6b3dda8e5a Record versions when importagents publishes, and refuse a silent rewrite
The previous commit closed this hole on the authoring path. This is the other
half, and the larger one: every organization agent is published by this
command, so until now none of them were versioned at all. definition_versions
was empty on a fully deployed system, and each deploy rewrote v1 in place with
whatever the files happened to say.

Versions are now recorded through the same transaction as the definitions, so
the history and the row it describes cannot disagree — either both land or
neither does. repo.VersionsRepo.Snapshot is what refuses a spec whose content
changed without its `version:` being raised, and that refusal now stops the
import rather than being absent.

Every offending spec is collected instead of the first being returned, matching
how the parse errors above it already behave: an operator who forgot to bump
three files should see three. That is safe here because the refusal comes from
comparing a row this code read, not from a failed statement — the INSERT is ON
CONFLICT DO NOTHING, so the transaction stays healthy and the remaining specs
can still be checked.

The header comment claimed "it does not create versions" as a deliberate
omission, deferring immutability to Phase 3. Phase 3 shipped; the comment is
updated rather than left to describe a decision that has been reversed.

Verified against a live stack:

  - first run over nine unversioned agents: 9 version(s) recorded
  - second run, files unchanged: still 9, not 18 — republishing is a no-op
  - a spec edited without a bump: refused by name, exit 1, and the live row
    did NOT contain the edit; the whole transaction rolled back
  - the same spec with version: 2: exit 0, v1 and v2 both in history, live
    row at v2

Not addressed, and visible while testing this: the command does not enforce
monotonicity. A file whose version is LOWERED still overwrites the live row,
because the upsert writes whatever the frontmatter says. History is unharmed —
the older version is already recorded and matches — but the deployed
definition silently goes backwards. That wants its own change.

cmd/importagents still has no test files, which predates this. The refusal
itself is covered by repo/versions_test.go; what is untested here is the
collecting and rollback around it, and run() opens its own pool from config,
so making it testable is a refactor rather than an addition.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-28 19:24:56 +05:30
Suriyakumarvijayanayagam
80ba57ace3 Refuse an edit that would rewrite an already-published version
§3 says a published version is immutable and editing publishes a new one.
The machinery for that was all present — an append-only definition_versions
table, a trigger, and repo.VersionsRepo.Snapshot, which already refuses to
store a version number whose content differs from what is stored.

Nothing acted on that refusal. snapshotIfPublished's error was discarded at
both call sites (`_ = s.snapshotIfPublished(...)`), and deliberately so: the
comment there explains that losing an author's work to protect a record of it
is the wrong trade. That is right for a recording failure and wrong for
exactly one case. A conflict is not the history failing to record; it is the
invariant firing.

The effect was silent. Editing a published agent without raising the
frontmatter version answered 200: the live row took the new text, the history
kept the old, and two different definitions were both called v1. Because
runtime.LoadAgentVersion resolves a pin by returning the CURRENT definition
whenever the pinned number equals the current one, a conversation pinned to v1
then ran the rewritten instructions while the audit trail showed the
originals. Verified against a live stack before the fix: PATCH answered 200,
agent_definitions held "SILENTLY CHANGED" and definition_versions still held
the published text, both labelled v2.

So the conflict is now detected before anything is written, where refusing
costs the author nothing but a version bump. The post-write snapshot keeps its
original contract for every other kind of failure, and republishing a version
unchanged stays the no-op it was. Drafts are untouched: they carry no promise,
and are still rewritten in place.

Not addressed here, and each its own change:

  - cmd/importagents never creates versions at all (documented at main.go:10),
    so the nine file-published organization agents are outside this entirely
    and every deploy still mutates v1 in place.
  - skill definitions never snapshot, so KindSkill exists with nothing writing
    it. Fixing that changes skill authoring behaviour and wants its own pass.
  - a concurrent publish of one version number with differing content can still
    pass this check and be caught by the unique index afterwards, where it is
    swallowed as before. That is the pre-existing behaviour, narrowed rather
    than removed.

Tests: the new case fails without the fix — the live row takes the rewritten
text at version 1 — and passes with it. Full suite green against PostgreSQL,
with only TestLive* skipped, which is what CI allows.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-28 18:28:20 +05:30
Suriyakumarvijayanayagam
f48b5606df Make the local-db overlay actually start, and pass the model credential through
The overlay had never been run against a fresh volume. Two faults, the first
hiding the second:

  - postgres:16-alpine ships libssl but not the openssl CLI, so the first-boot
    certificate generation exited 127 in a restart loop. It failed invisibly:
    the 2>/dev/null on the openssl line swallowed sh's "not found" as well, so
    `docker logs` was completely empty. openssl is now installed on the boot
    that generates the certificate, inside the same guard, so a restart still
    needs no network.

  - the certificate was written into /var/lib/postgresql/data BEFORE initdb
    ran, and initdb refuses to initialise a directory that is not empty. That
    made a fresh volume unstartable regardless of the first fault. The
    certificate now lives in its own volume, which keeps it persistent — the
    reason it was put in the data directory — without touching the cluster's.

Separately, docker-compose.yml did not pass ANTHROPIC_API_KEY to the api
container, so a compose deployment could never register the agent run routes:
POST /agents/{id}/runs answered 404 and /version reported two endpoints fewer.
The model and embedder variables are now passed through, all defaulting to
empty so a deployment without them behaves exactly as it did.

Verified on a fresh volume: 56/56 verify-deploy checks against the resulting
stack, including a live agent run and 34 chunks embedded through Ollama.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-28 17:12:34 +05:30
d190fc8ee9 Preserve CLAUDE.md and add a handover document
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
Neither survived a machine change. CLAUDE.md sat in the directory ABOVE both
repositories, which is not a git repository at all, so the governing document
for the project existed on exactly one laptop. It is now in this repository;
place a copy at the parent level on a new machine, where it covers both.

docs/handover.md records what CLAUDE.md does not: what was decided and why,
what is deployed and how to verify it, and the conventions that produce
confident wrong numbers rather than errors — a score of 0 meaning "not rated",
screened_at being vestigial, shift data anchored to today.

Written because Claude Code's own memory is per-machine and keyed to the
absolute path of the checkout: it does not sync, and a different path on a new
machine reads a different folder. A file in the repository travels with the
code and is useful to a person besides.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
2026-08-28 13:56:50 +05:30
a222dcd3e4 Add evals for every shipped agent, a policy corpus, and CI
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
§9 says no agent ships without evals. Eight of the nine had none: the two
other suites in evals/ are harness fixtures rather than agents in the
registry, so the rule was being met by one agent in nine.

Evals — 40 new cases, five per agent, every one carrying mustNotLeak:

  - the agent is loaded from its real spec in agents/*.md rather than
    written out again in Go. A hand-copied agent tests the copy: it keeps
    passing after somebody edits the spec, which is the moment it most
    needed to fail.
  - callNamed calls the tool a case names. toolThenAnswer always called
    tools[0], so seven of positions-agent's eight tools were unreachable,
    and a boundary nothing calls is a boundary nothing tests.
  - seedWorkspace fills BOTH tenants. A leak test against an empty second
    tenant cannot fail.

Verified by breaking workersByScore's org predicate: six cases across four
agents fail with LEAKED "RIVAL".

Knowledge — six policy documents, taking the corpus from 2 to 8 (34
chunks). Three restricted to admin and employer, five tenant-wide. They
cover what the tools cannot: a tool reports how many shifts went unworked,
a policy says what cover costs inside 24 hours.

corpus_test.go treats those documents as product rather than fixtures. The
first version was tautological — it read audience: from a file and checked
that file's audience was enforced, so opening a restricted document passed.
mustNotBeTenantWide now holds that judgement apart from the files, with the
reason recorded for each.

CI — the checks this repository already had, made unskippable. testutil
calls t.Skipf on an unreachable database, so a dead service container would
produce a green build over a suite that ran almost nothing. Simulated: go
test exits 0 with 74 tests skipped, including every tenant-isolation test.
The guard exits 1 and names them, while still allowing TestLive* to skip
without a model key.

This CI tests; it does not deploy. The README's claim that migrations are
run by CI against the target database remains aspirational.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
2026-08-28 13:52:46 +05:30
f7df96c973 agent build 2026-08-28 12:21:44 +05:30
b6f8655909 aravind changes 2026-08-25 16:37:05 +05:30
cadea4bd92 Merge pull request 'Add CORS credentials, transactional endpoints, and container deployment' (#1) from feat/cors-transactions-docker into main
Reviewed-on: #1
2026-08-25 06:04:14 +00:00
Suriyakumarvijayanayagam
954ba9076f Add CORS credentials, transactional endpoints, and container deployment
CORS
  cors.go never set Access-Control-Allow-Credentials, so the
  cookie-authenticated API was unreadable from any cross-origin frontend:
  the server answered correctly and the browser blocked the page from
  reading it. Set for allowlisted origins on both the preflight and the
  actual response. Three tests added.

  HTTP_COOKIE_SAMESITE (lax|none|strict, default lax) is new. CORS is only
  half of what a cross-origin browser call needs; SameSite is judged on
  registrable domain, so a frontend on an unrelated domain gets perfect CORS
  headers and still no cookie. "none" is the only value that survives that,
  and validate() refuses it without the Secure flag.

  The "*" rejection now explains itself: browsers refuse Allow-Origin "*"
  together with credentials, so it would break every authenticated call
  rather than loosen anything.

Transactional endpoints (api-contract.md 12.1)
  POST /api/v1/job-applications/{id}/hire
  POST /api/v1/job-postings/{id}/assignments

  Replaces two client-side loops that wrote several records with no
  transaction and no rollback. Each is now one endpoint and one transaction,
  built over repo.Repo so org scoping, derived columns, type casts and error
  translation are not re-derived. Authorization reuses the existing policy
  table rather than adding a parallel one: a workflow is exactly as
  privileged as the writes it performs. 13 tests, including both rollback
  paths.

Bug fix in the repository layer
  repo.bindValue handled int64/int/float64/string but not int32, which is
  what pgx returns for a PostgreSQL `int` column. Nothing previously read a
  record and wrote one of its fields elsewhere, so it never surfaced; the
  hire flow does exactly that and failed with "ai_score must be a number".
  Both KindInt and KindFloat now accept the widths pgx actually produces.

Deployment
  infrastructure/Dockerfile.api  multi-stage, cross-compiling (BUILDPLATFORM
    + GOARCH) so linux/amd64 builds from arm64 are compiled rather than
    emulated. Alpine runtime, non-root uid 10001, 22.1 MB. Ships api, seed,
    setpassword and migrate, plus the migrations, so a Kubernetes
    initContainer can apply the schema from the same image and tag as the
    API. HEALTHCHECK keys on status code, not body, so a "degraded" instance
    is not pulled from rotation during a migration window.

  infrastructure/docker-compose.yml  migrations run to completion before the
    API starts. Assumes a managed PostgreSQL; the local-db overlay adds one
    with TLS enabled so APP_ENV=production is met rather than dodged.

  scripts/drop_public_tables.go  the one-off used to clear an unrelated
    schema from krowdb on 2026-08-24, kept for the record. Build-tagged
    ignore and gated on CONFIRM_DROP=yes.

Verified against PostgreSQL: 16/16 new tests pass, and the image was built,
run and exercised end to end (login, CORS preflight, authenticated reads,
transaction rollback).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CmQiGq73Uyfq7J4yR8Vxxw
2026-08-25 11:33:01 +05:30
7d12ebef3d first commit 2026-08-24 13:06:29 +05:30