797ee5f2d275147e6f3438bcb4ce170b7e074324
7 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
| 5166fde764 |
Add GatewayFailure: the provider not answering is not a tool failing
terminationFor sent every gateway error that was not Refused or Timeout to ToolFailure, because the enum had nowhere else to put it. On 2026-09-22 that was 131 of 318 production runs, and not one of them was a tool failing: 56 were retired model ids, 25 an exhausted Anthropic balance, 45 Groq's free-tier rate limit -- the only one still happening. An operator reading the termination column saw "a tool is broken" for two weeks while the actual answer was "we are not paying for capacity". GatewayFailure is the seventh termination. Rate limited, request rejected, credential refused and unreachable land there; Refused and Deadline keep their own reasons; a non-gateway error is still the tool layer's. A delegation whose subagent died at the gateway now carries that reason up to the parent instead of reading as a tool call that failed. Migration 000016 widens the CHECK that 000006 chose precisely so this would be a migration rather than an ALTER TYPE. Its down folds any GatewayFailure rows back to ToolFailure BEFORE narrowing the constraint, which is the order that works; verified up, down and up again on a scratch database. Existing rows are left as they are -- the trajectory entries still carry the gateway.* code for anyone reclassifying history. The surface wording is the one termination where "try again" is honest advice, since the dominant cause clears within a minute. Full suite run against a real database, including the tests that skip without one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g |
|||
| 9d3192a9c4 |
Replace the model ids with ones Groq actually serves
The defaults shipped yesterday were wrong the day they shipped, and a real key proved it in one request. Groq serves neither llama-3.1-8b-instant nor llama-3.3-70b-versatile any more. Both were chosen from memory, both passed startup validation, and every agent run would have failed with a 400. This is the exact failure the claude-* guard was written to catch, arriving from the side that guard cannot see. A prefix check can reject a vendor this service cannot call; it has no way to know a provider retired an id last month. That is not a gap in the check, it is a gap in the class of thing local validation can know, so the fix is not another guard: TestConfiguredModelsAreServed asks the provider. It lists /models — part of the same openai-compatible surface the gateway already speaks, so every supported provider answers it — and fails if a configured id is absent, printing what is available. It reads the ids through config.DefaultModels() rather than repeating them, because a second copy would be the first thing to drift, and drift is the whole failure. Skipped without a credential like the rest of the live suite. Verified three ways: it fails on the retired id with the message an operator needs, skips clean with no key, passes on the new ones. New defaults, chosen against the live account rather than from memory: openai/gpt-oss-20b (fast) and openai/gpt-oss-120b (balanced, deep). Tool calling confirmed on both. groq/compound-mini was ruled out — it cannot do tool calls at all, which this platform requires. MODEL_REASONING_EFFORT is now documented as safe here and NOT portable: gpt-oss accepts low/medium/high, exactly the scale openAIEffort maps onto, while qwen/qwen3.6-27b on the same account rejects all three and fails the whole request rather than ignoring the key. I7 IS NO LONGER UNPROVEN. make eval-live passes all three cases twice against gpt-oss-120b, the planted-injection case included: answers from the handbook, cites, refuses the injection, leaks neither the operator-only pay guidance nor the other tenant's figures. CLAUDE.md §12 and handover.md updated from "urgent" to measured, dated, and scoped to the one model it is evidence about. One real defect found on the way. The handbook grounding check failed once on an answer containing the phrase it wanted — "more than ten minutes" on screen, strings.Contains false — which leaves an invisible separator as the only explanation; the same model writes "47 %" and a U+2011 hyphen elsewhere. The flaky assertion is the small half. THE LEAK ASSERTIONS USED THE SAME MATCH and fail in the dangerous direction: "attacker@evil.test" with a zero-width space, or "uplift" with a soft hyphen, would have been reported clean. A permission test that cannot see the leak it is hunting is worse than none, because it is believed. normalizeForMatch folds those away, and its test pins that every case is one plain ToLower MISSES — a case whose naive match already succeeds fails, so the suite cannot fill with examples that demonstrate nothing. That caught my own first BOM case, which put the mark where Contains found it regardless. gofmt clean, vet clean, 15/15 packages pass offline; live suite green twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g |
|||
| 34fa58a6b9 |
Remove the Anthropic path; the gateway speaks one wire protocol
The platform now runs on Groq by default, through the OpenAI-compatible chat-completions shape. That shape is not one vendor — Gemini, OpenRouter, Together, vLLM and a local Ollama serve it too — so moving again stays configuration rather than code. Two things in the deleted file were not Anthropic's and would have gone with it silently: withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and CALLED BY openai.go. Deleting the file wholesale would have removed the retry policy of the provider that survived, and nothing in openai.go mentions it, so the loss would have been invisible until the next 429. The policy is a property of this platform's runs, not of a vendor's API; it now lives in retry.go where no provider can carry it off. StreamComplete had the same problem and moves to gateway.go, beside the Streamer interface whose comment already referenced it. Three stale-configuration failures are now refused at startup instead of being ignored. Each was verified firing through the real config.Load(): MODEL_PROVIDER=anthropic — named separately from every other wrong value because it used to be correct. Ignoring it gives a stack that believes it is on Claude while every run goes to Groq and is billed there. ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an operator did set is the worst version of this: they fail every run on a missing credential they are looking straight at. A leftover claude-* model id, naming the tier that carries it. This is the check the previous commit's error-detail work was diagnosing: such an id is accepted by this process, rejected by the provider, and 400s on EVERY run. "A model is wrong" does not say which of three lines to edit. Defaults ship as a matched pair. defaultBaseURL and the three tier ids are one decision, not four: an id is only meaningful against the service that serves it, and a Groq id on an OpenAI base URL is the same failure from the other side. The tiers also stop being one model — a tier whose cost does not differ is a distinction that buys nothing. Verified end to end against a stub of the wire, driving the real wiring (config.Load in production mode, gateway.New, StreamComplete): streamed deltas, tool-call decoding, the loopback credential exemption, and usage totalling 150 rather than 190 — the cached-prefix subtraction still holds. gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still needs a reachable database. NOT verified: the I7 planted-injection eval. Removing this path removed the only model whose refusal behaviour had been measured against it, so the new default is unproven there until `make eval-live` runs with a real key. The Groq model ids should also be confirmed against Groq's current lineup. Flagged in CLAUDE.md §12 and docs/handover.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g |
|||
| dc785b917c |
Separate what a worker does from what a company needs filled
Owliver could offer neither create. The Create Position flow worked and no chip anywhere suggested it, because the chip row is entirely the backend's static catalogue and no intent in it wrote anything. The gap was never in the frontend's trigger matching — every phrasing already routed. `employee_roles` is the supply side of `job_postings`. A posting is what the ORGANIZATION needs filled; this is what a WORKER says they do. They share a vocabulary and almost nothing else: "3 years" on a posting is a minimum an applicant must clear, and the same words here are what the person has. There is deliberately no foreign key between them — supply and demand already meet through `job_applications`, which carries the funnel, the interview and the outcome, and a second weaker link would disagree with it the first time somebody withdrew. NO NEW COMPANY ENTITY, AND THAT IS THE LOAD-BEARING DECISION. "Create a company position" reads like it needs a client record. `organizations` is the TENANT — absent from the resource table, absent from the policy map, written only by the seeder — so creating a row there from a chat flow would provision a new tenant, and the position would carry an org_id the operator's session cannot see. The operator could never view the record they just created. That breaks I5 and I1 to add a feature nobody asked for. The client stays free text on the posting, per blueprint decision D2, and the flow simply offers the clients this organization already staffs for as chips. No schema change, no endpoint change. Create is operators-only, and that is an I1 decision rather than a deferral. The worker is named explicitly on the row and is deliberately NOT derived from the session, because an operator recording a role on somebody's behalf is the whole point of the flow. Granting talent the same Create would let a talent caller write a role under any worker_email in the tenant — the attribution hole Phase 3D closed elsewhere. Talent reads its own via a ScopeEmail predicate, which is in place now so the grant is one line when a talent console exists. `created_by` is in gen_resources.py's SERVER_OWNED as well as the policy's Derived list. Both are required and the pairing is easy to miss: Derived fills the column from the session, SERVER_OWNED is what makes the descriptor ReadOnly so a request body cannot set it in the first place. Without it, TestDerivedColumnsAreReadOnlyOrTalentScoped fails — verified by mutation, not by reading. The two catalogue intents carry PHRASE terms only. A bare "position" or "role" term scores 10, the same as every reading on that page, and wins the tie on declaration order — so a create chip would have arrived by evicting `positions-attention` from the exact ordered result TestPositionsSuggestions asserts. An offer to create something must not displace the reading a person actually asked for. Neither declares a Subject, on the precedent of `position-spec-steps`: a Subject would let the bare query "summarize" match through matchShape and survive filterOnTopic. Neither declares a Signal, so an empty composer still reports what the organization needs rather than proposing paperwork. Chip text is the coupling with nothing else holding it together: no page context declares `capabilities`, so every server suggestion dispatches as its own TEXT and is answered by whichever skill's trigger that text matches. A renamed chip would open nothing, silently. Asserted on the frontend side. The down migration drops `employee_role_status` and keeps `english_level`, which is shared with job_postings.english_required and job_applications.english_level. Rolled back and re-applied against the database to prove it, not asserted. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g |
|||
| c74fe7e074 |
Add an OpenAI-compatible gateway, so the model provider is a config value
The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.
`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.
`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.
THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.
Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.
Three configuration mistakes are refused at startup rather than at runtime:
- MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
endpoint and ignores the field, so this is a deployment that believes it
switched providers and did not — every run still goes to Anthropic and is
still billed there, with nothing in the logs to say so. Cost is the whole
reason this change exists, and that is the one mistake that silently
defeats it.
- An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
- A production deployment with no credential — except against localhost,
which needs none, and demanding one would make the free local path
impossible to configure.
reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.
`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.
Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.
NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
|
|||
| 4c29185c3b |
Correct the governing documents where this session disproved them
Both documents are the first thing a new reader trusts, and several of their
claims were wrong — some wrong from the start, some overtaken by work this
week. A governing document that misdescribes the system is worse than none,
because it is believed.
CLAUDE.md §10 specified Python 3.11, FastAPI, SQLAlchemy and Alembic. The code
is Go and has never been anything else. That is corrected rather than quietly
deleted, so the next person understands the document drifted rather than
wondering which half to trust.
Also in CLAUDE.md: the tool count was 17 with one write and is 19 with two;
delegation now exists and §11's Orchestration row says what it guarantees; the
embedder deviation described a Voyage-or-stand-in choice that has since become
EMBED_PROVIDER with three options, of which production sets none.
The handover claimed three things that this session disproved by running them:
- "definition_versions is empty ... nothing has gone through it". It was not
empty in production; activity-agent had a v1 that the shipped file
contradicted, which is how a real drift was found. It now holds every
agent and skill.
- "Skills are still stored in user_preferences". They are rows in
skill_definitions, and are now versioned.
- "make eval-live ... has never been run". It has, it passes 3/3, and what
it established is recorded — including that the handbook corpus carries a
planted prompt injection which the agent refused and reported. That is I7
holding against a real model, which is worth more than the pass count.
The endpoint counts were one high throughout (55/57, not 56/58) — the delta of
two was always right, so the signal worked and the absolute numbers did not.
Added, because they cost time this week and would cost it again:
- the app reaches its database through pgbouncer, not PostgreSQL directly.
Enabling TLS on PostgreSQL does nothing for the application hop; pgbouncer
terminates 5432 and needs its own client_tls_sslmode.
- the seeded UserActivity is NOT anchored to today the way ShiftRecord is,
so it ages out of every window the activity tools offer. Twenty-three days
old as of writing: zero events in the last 7 days, 6 of 15 in the last 30.
The agent answers truthfully and the demo looks dead.
- the whole stack runs on Docker alone. Dockerfile.api builds every command
plus the migrate CLI, so a new machine needs neither Go nor psql — which
is how this one was set up, having no Homebrew.
- Ollama runs on the HOST, so a container reaches it at
host.docker.internal, not localhost. The old .env said localhost and would
have failed with nothing obviously wrong.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
|
|||
| d190fc8ee9 |
Preserve CLAUDE.md and add a handover document
Neither survived a machine change. CLAUDE.md sat in the directory ABOVE both repositories, which is not a git repository at all, so the governing document for the project existed on exactly one laptop. It is now in this repository; place a copy at the parent level on a new machine, where it covers both. docs/handover.md records what CLAUDE.md does not: what was decided and why, what is deployed and how to verify it, and the conventions that produce confident wrong numbers rather than errors — a score of 0 meaning "not rated", screened_at being vestigial, shift data anchored to today. Written because Claude Code's own memory is per-machine and keyed to the absolute path of the checkout: it does not sync, and a different path on a new machine reads a different folder. A file in the repository travels with the code and is useful to a person besides. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0186JgqQUCDS8ZwGmyw3ymWu |