Commit Graph

7 Commits

Author SHA1 Message Date
34fa58a6b9 Remove the Anthropic path; the gateway speaks one wire protocol
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 7s
The platform now runs on Groq by default, through the OpenAI-compatible
chat-completions shape. That shape is not one vendor — Gemini, OpenRouter,
Together, vLLM and a local Ollama serve it too — so moving again stays
configuration rather than code.

Two things in the deleted file were not Anthropic's and would have gone
with it silently:

  withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and
  CALLED BY openai.go. Deleting the file wholesale would have removed the
  retry policy of the provider that survived, and nothing in openai.go
  mentions it, so the loss would have been invisible until the next 429.
  The policy is a property of this platform's runs, not of a vendor's API;
  it now lives in retry.go where no provider can carry it off.

  StreamComplete had the same problem and moves to gateway.go, beside the
  Streamer interface whose comment already referenced it.

Three stale-configuration failures are now refused at startup instead of
being ignored. Each was verified firing through the real config.Load():

  MODEL_PROVIDER=anthropic — named separately from every other wrong value
  because it used to be correct. Ignoring it gives a stack that believes it
  is on Claude while every run goes to Groq and is billed there.

  ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an
  operator did set is the worst version of this: they fail every run on a
  missing credential they are looking straight at.

  A leftover claude-* model id, naming the tier that carries it. This is
  the check the previous commit's error-detail work was diagnosing: such an
  id is accepted by this process, rejected by the provider, and 400s on
  EVERY run. "A model is wrong" does not say which of three lines to edit.

Defaults ship as a matched pair. defaultBaseURL and the three tier ids are
one decision, not four: an id is only meaningful against the service that
serves it, and a Groq id on an OpenAI base URL is the same failure from the
other side. The tiers also stop being one model — a tier whose cost does
not differ is a distinction that buys nothing.

Verified end to end against a stub of the wire, driving the real wiring
(config.Load in production mode, gateway.New, StreamComplete): streamed
deltas, tool-call decoding, the loopback credential exemption, and usage
totalling 150 rather than 190 — the cached-prefix subtraction still holds.

gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still
needs a reachable database.

NOT verified: the I7 planted-injection eval. Removing this path removed the
only model whose refusal behaviour had been measured against it, so the new
default is unproven there until `make eval-live` runs with a real key. The
Groq model ids should also be confirmed against Groq's current lineup.
Flagged in CLAUDE.md §12 and docs/handover.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-05 11:52:21 +05:30
c74fe7e074 Add an OpenAI-compatible gateway, so the model provider is a config value
The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.

`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.

`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.

THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.

Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.

Three configuration mistakes are refused at startup rather than at runtime:

  - MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
    endpoint and ignores the field, so this is a deployment that believes it
    switched providers and did not — every run still goes to Anthropic and is
    still billed there, with nothing in the logs to say so. Cost is the whole
    reason this change exists, and that is the one mistake that silently
    defeats it.
  - An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
  - A production deployment with no credential — except against localhost,
    which needs none, and demanding one would make the free local path
    impossible to configure.

reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.

`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.

Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.

NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-01 11:47:53 +05:30
57c2a52c1e Refuse an HTTP write timeout that would cut off a legal agent run
Some checks failed
CI / test (push) Failing after 5m32s
CI / fixture (push) Failing after 59s
Production answered 502 Bad Gateway on a non-streamed agent run. Nothing about
that was a gateway fault: krow-proxy already had proxy_read_timeout 3600s, and
the API pods were healthy with zero restarts throughout.

HTTP_WRITE_TIMEOUT was 30s. Every shipped agent runs at the `balanced` tier,
whose deadline is 60s, and the `deep` tier allows 120s. So the server aborted
the response on any run over half the time the runtime considered legal, the
proxy saw its upstream vanish mid-response, and it reported the only thing it
could. A gateway error for something no gateway did — which is why it looked
like infrastructure for as long as it did.

Delegation did not cause this; it made it routine. A parent that asks two
subagents takes longer than one answering alone, so a latent misconfiguration
became a reliable one. Verified: the exact request that returned 502 now
answers 200 in 18s.

Streaming is what hid it, and that is the part worth keeping in mind. The chat
panel uses SSE, so the product looked healthy while every non-streaming caller
got 502 on a slow question. A bug only reachable by the callers who do not yet
exist is one nobody reports.

So the value is now derived from the thing that constrains it — the default is
DeepestAgentDeadline plus headroom rather than a number typed once — and
validate() refuses anything below that deadline at startup. A slow,
intermittent, misattributed failure becomes a message on the first boot.

DeepestAgentDeadline is duplicated in internal/config rather than imported,
because internal/runtime already imports internal/config and a cycle to share
one number is a bad trade. TestConfigKnowsTheDeepestAgentDeadline asserts the
two agree, so drift is a build failure rather than a discovery. It also checks
that no tier exceeds it, or the name lies.

ORDERING, and it matters for the next deploy: the check refuses the old 30s, so
a pod carrying this image against an unpatched configmap will not boot.
Production's configmap is already 180s. The handover says so too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-31 12:06:26 +05:30
fd1812e161 Record the CI runner, now that one exists
Some checks failed
CI / test (push) Failing after 6m25s
CI / fixture (push) Failing after 47s
Both repositories carried GitHub Actions workflows on a Gitea remote and
nobody had confirmed a runner. There was not one: the 924 frontend checks,
the whole Go suite, the skip guard and the suite-shrank guard had never run
on a push, only when somebody remembered.

gitea/act_runner v0.6.1 is registered as krow-runner on the cluster host. This
commit is also the first push that can prove it picks up a job, which is the
failure mode worth catching — a runner that registers and never runs anything
looks identical to a healthy one in the Runners list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-31 11:13:47 +05:30
629d97181d Rebase Staff too, which was the last thing keeping the chart incomplete
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
And correct the previous commit's closing claim, which was wrong.

c378bc0 said the Hiring activity chart "still renders empty ... the fault is
further down in that component, not in the data." That was a retraction of a
correct diagnosis, made from a screenshot taken before the rebase had reached
the browser. The chart was empty BECAUSE the data was stale, exactly as first
diagnosed, and rebasing fixed it. Checked properly this time: the area path
carries real values, and the rendered chart shows applications peaking at 16
around 8/18 with the screening and interview series drawn over it.

What was genuinely still missing was Hires. Staff was not rebased, so
hire_date stayed 35 days old with nothing inside the 30-day window and that
series drew nothing. It is rebased now, anchored on hire_date rather than
created_date, because the hire is the event the chart plots.

Leaving it behind had also introduced an inconsistency of my own making: once
applications moved, a candidate was hired last week according to their
application and five weeks ago according to their staff record. Rebasing them
together removes that.

All four series now render. Full suite green.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 15:21:14 +05:30
4c29185c3b Correct the governing documents where this session disproved them
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
Both documents are the first thing a new reader trusts, and several of their
claims were wrong — some wrong from the start, some overtaken by work this
week. A governing document that misdescribes the system is worse than none,
because it is believed.

CLAUDE.md §10 specified Python 3.11, FastAPI, SQLAlchemy and Alembic. The code
is Go and has never been anything else. That is corrected rather than quietly
deleted, so the next person understands the document drifted rather than
wondering which half to trust.

Also in CLAUDE.md: the tool count was 17 with one write and is 19 with two;
delegation now exists and §11's Orchestration row says what it guarantees; the
embedder deviation described a Voyage-or-stand-in choice that has since become
EMBED_PROVIDER with three options, of which production sets none.

The handover claimed three things that this session disproved by running them:

  - "definition_versions is empty ... nothing has gone through it". It was not
    empty in production; activity-agent had a v1 that the shipped file
    contradicted, which is how a real drift was found. It now holds every
    agent and skill.
  - "Skills are still stored in user_preferences". They are rows in
    skill_definitions, and are now versioned.
  - "make eval-live ... has never been run". It has, it passes 3/3, and what
    it established is recorded — including that the handbook corpus carries a
    planted prompt injection which the agent refused and reported. That is I7
    holding against a real model, which is worth more than the pass count.

The endpoint counts were one high throughout (55/57, not 56/58) — the delta of
two was always right, so the signal worked and the absolute numbers did not.

Added, because they cost time this week and would cost it again:

  - the app reaches its database through pgbouncer, not PostgreSQL directly.
    Enabling TLS on PostgreSQL does nothing for the application hop; pgbouncer
    terminates 5432 and needs its own client_tls_sslmode.
  - the seeded UserActivity is NOT anchored to today the way ShiftRecord is,
    so it ages out of every window the activity tools offer. Twenty-three days
    old as of writing: zero events in the last 7 days, 6 of 15 in the last 30.
    The agent answers truthfully and the demo looks dead.
  - the whole stack runs on Docker alone. Dockerfile.api builds every command
    plus the migrate CLI, so a new machine needs neither Go nor psql — which
    is how this one was set up, having no Homebrew.
  - Ollama runs on the HOST, so a container reaches it at
    host.docker.internal, not localhost. The old .env said localhost and would
    have failed with nothing obviously wrong.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-29 12:10:41 +05:30
d190fc8ee9 Preserve CLAUDE.md and add a handover document
Some checks failed
CI / test (push) Has been cancelled
CI / fixture (push) Has been cancelled
Neither survived a machine change. CLAUDE.md sat in the directory ABOVE both
repositories, which is not a git repository at all, so the governing document
for the project existed on exactly one laptop. It is now in this repository;
place a copy at the parent level on a new machine, where it covers both.

docs/handover.md records what CLAUDE.md does not: what was decided and why,
what is deployed and how to verify it, and the conventions that produce
confident wrong numbers rather than errors — a score of 0 meaning "not rated",
screened_at being vestigial, shift data anchored to today.

Written because Claude Code's own memory is per-machine and keyed to the
absolute path of the checkout: it does not sync, and a different path on a new
machine reads a different folder. A file in the repository travels with the
code and is useful to a person besides.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu
2026-08-28 13:56:50 +05:30