Files
krow_backend/.github/workflows/ci.yml
Suriyakumarvijayanayagam 34fa58a6b9
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 7s
Remove the Anthropic path; the gateway speaks one wire protocol
The platform now runs on Groq by default, through the OpenAI-compatible
chat-completions shape. That shape is not one vendor — Gemini, OpenRouter,
Together, vLLM and a local Ollama serve it too — so moving again stays
configuration rather than code.

Two things in the deleted file were not Anthropic's and would have gone
with it silently:

  withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and
  CALLED BY openai.go. Deleting the file wholesale would have removed the
  retry policy of the provider that survived, and nothing in openai.go
  mentions it, so the loss would have been invisible until the next 429.
  The policy is a property of this platform's runs, not of a vendor's API;
  it now lives in retry.go where no provider can carry it off.

  StreamComplete had the same problem and moves to gateway.go, beside the
  Streamer interface whose comment already referenced it.

Three stale-configuration failures are now refused at startup instead of
being ignored. Each was verified firing through the real config.Load():

  MODEL_PROVIDER=anthropic — named separately from every other wrong value
  because it used to be correct. Ignoring it gives a stack that believes it
  is on Claude while every run goes to Groq and is billed there.

  ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an
  operator did set is the worst version of this: they fail every run on a
  missing credential they are looking straight at.

  A leftover claude-* model id, naming the tier that carries it. This is
  the check the previous commit's error-detail work was diagnosing: such an
  id is accepted by this process, rejected by the provider, and 400s on
  EVERY run. "A model is wrong" does not say which of three lines to edit.

Defaults ship as a matched pair. defaultBaseURL and the three tier ids are
one decision, not four: an id is only meaningful against the service that
serves it, and a Groq id on an OpenAI base URL is the same failure from the
other side. The tiers also stop being one model — a tier whose cost does
not differ is a distinction that buys nothing.

Verified end to end against a stub of the wire, driving the real wiring
(config.Load in production mode, gateway.New, StreamComplete): streamed
deltas, tool-call decoding, the loopback credential exemption, and usage
totalling 150 rather than 190 — the cached-prefix subtraction still holds.

gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still
needs a reachable database.

NOT verified: the I7 planted-injection eval. Removing this path removed the
only model whose refusal behaviour had been measured against it, so the new
default is unproven there until `make eval-live` runs with a real key. The
Groq model ids should also be confirmed against Groq's current lineup.
Flagged in CLAUDE.md §12 and docs/handover.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-05 11:52:21 +05:30

150 lines
5.7 KiB
YAML

name: CI
# Every check this repository already had, run on every push instead of when
# somebody remembers. Nothing here is new verification — it is the verification
# that existed, made unskippable.
on:
push:
branches: [main]
pull_request:
workflow_dispatch:
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
jobs:
test:
runs-on: ubuntu-latest
# The tests SKIP when PostgreSQL is unreachable — see testutil.New, which
# calls t.Skipf rather than failing, so a developer without a database can
# still run the non-database tests. In CI that behaviour is a trap: a broken
# service container would produce a green build over a suite that tested
# almost nothing. The guard at the end of this job is what closes it.
services:
postgres:
image: postgres:16-alpine
env:
POSTGRES_PASSWORD: postgres
POSTGRES_USER: postgres
POSTGRES_DB: postgres
ports: ['5432:5432']
options: >-
--health-cmd "pg_isready -U postgres"
--health-interval 5s
--health-timeout 5s
--health-retries 20
env:
DATABASE_HOST: 127.0.0.1
DATABASE_PORT: '5432'
DATABASE_USER: postgres
DATABASE_PASSWORD: postgres
steps:
- uses: actions/checkout@v4
- uses: actions/setup-go@v5
with:
go-version-file: go-api/go.mod
cache-dependency-path: go-api/go.sum
- name: gofmt
working-directory: go-api
run: |
unformatted="$(gofmt -l ./cmd ./internal)"
if [ -n "$unformatted" ]; then
echo "not gofmt'd:"; echo "$unformatted"; exit 1
fi
- name: go vet
working-directory: go-api
run: go vet ./...
- name: Tests
working-directory: go-api
# -json so the guard below can count what actually ran, and `|| true` so
# a failure reaches that guard rather than ending the job here — the
# guard reports which tests failed, which the raw JSON does not.
run: go test ./... -count=1 -json > /tmp/test.json || true
- name: Fail if the database tests skipped
# The point of this job. testutil skips on an unreachable database, so
# "0 failures" is not the same as "the suite ran": a service container
# that never came up would otherwise look identical to a passing build.
run: |
python3 - <<'PY'
import json, sys
skipped, passed, failed = [], 0, []
for line in open('/tmp/test.json'):
line = line.strip()
if not line.startswith('{'):
continue
try:
e = json.loads(line)
except ValueError:
continue
if e.get('Action') == 'skip' and e.get('Test'):
skipped.append(e['Test'])
if e.get('Action') == 'pass' and e.get('Test'):
passed += 1
if e.get('Action') == 'fail' and e.get('Test'):
failed.append(e['Test'])
print(f"{passed} passed, {len(failed)} failed, {len(skipped)} skipped")
if failed:
print("FAILED:"); [print(" ", t) for t in failed[:40]]
sys.exit(1)
# A Test<Pkg>Live* / TestLive* test skips without MODEL_API_KEY, which
# is correct here: CI should not spend tokens on every push, and the key
# should not be present unless somebody put it there deliberately. Any
# OTHER skip means the database was unreachable, and that is the case
# this guard exists for — testutil calls t.Skipf rather than failing, so
# a dead service container would otherwise look exactly like a pass.
unexpected = [t for t in skipped if not t.startswith('TestLive')]
if skipped:
print("skipped:"); [print(" ", t) for t in skipped[:40]]
if unexpected:
print("\nThese are not live-model tests, so they skipped because the")
print("database was unreachable. A green build over a suite that did")
print("not run is worse than a red one.")
sys.exit(1)
if passed < 200:
sys.exit(f"only {passed} tests passed; the suite is far smaller than expected — did it run?")
PY
- name: Migrations are reversible
# §10: one migration per PR, reversible. Applying and rolling back the
# newest one is the cheapest way to find out that it is not.
working-directory: go-api
run: go test ./internal/repo/... -run 'Migration' -count=1
fixture:
# seed.json is generated from the frontend's seed module. The two used to be
# hand-maintained copies, and drift between them is silent: the demo and the
# API answer the same question with different numbers.
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Check out the frontend beside this repo
uses: actions/checkout@v4
with:
repository: ${{ github.repository_owner }}/krow-demo
path: ../krow-demo
# A private sibling needs a token with read access; without one this
# job reports that it could not check, rather than passing quietly.
token: ${{ secrets.FRONTEND_REPO_TOKEN }}
continue-on-error: true
- uses: actions/setup-node@v4
with:
node-version: '20'
- name: seed.json matches the frontend seed module
run: |
if [ ! -d ../krow-demo ]; then
echo "krow-demo is not available to this job, so the fixture could not be checked."
echo "Set FRONTEND_REPO_TOKEN to enable it. Not passing silently."
exit 1
fi
cd ../krow-demo && npm ci && npm run seed:check