The defaults shipped yesterday were wrong the day they shipped, and a real key proved it in one request. Groq serves neither llama-3.1-8b-instant nor llama-3.3-70b-versatile any more. Both were chosen from memory, both passed startup validation, and every agent run would have failed with a 400. This is the exact failure the claude-* guard was written to catch, arriving from the side that guard cannot see. A prefix check can reject a vendor this service cannot call; it has no way to know a provider retired an id last month. That is not a gap in the check, it is a gap in the class of thing local validation can know, so the fix is not another guard: TestConfiguredModelsAreServed asks the provider. It lists /models — part of the same openai-compatible surface the gateway already speaks, so every supported provider answers it — and fails if a configured id is absent, printing what is available. It reads the ids through config.DefaultModels() rather than repeating them, because a second copy would be the first thing to drift, and drift is the whole failure. Skipped without a credential like the rest of the live suite. Verified three ways: it fails on the retired id with the message an operator needs, skips clean with no key, passes on the new ones. New defaults, chosen against the live account rather than from memory: openai/gpt-oss-20b (fast) and openai/gpt-oss-120b (balanced, deep). Tool calling confirmed on both. groq/compound-mini was ruled out — it cannot do tool calls at all, which this platform requires. MODEL_REASONING_EFFORT is now documented as safe here and NOT portable: gpt-oss accepts low/medium/high, exactly the scale openAIEffort maps onto, while qwen/qwen3.6-27b on the same account rejects all three and fails the whole request rather than ignoring the key. I7 IS NO LONGER UNPROVEN. make eval-live passes all three cases twice against gpt-oss-120b, the planted-injection case included: answers from the handbook, cites, refuses the injection, leaks neither the operator-only pay guidance nor the other tenant's figures. CLAUDE.md §12 and handover.md updated from "urgent" to measured, dated, and scoped to the one model it is evidence about. One real defect found on the way. The handbook grounding check failed once on an answer containing the phrase it wanted — "more than ten minutes" on screen, strings.Contains false — which leaves an invisible separator as the only explanation; the same model writes "47 %" and a U+2011 hyphen elsewhere. The flaky assertion is the small half. THE LEAK ASSERTIONS USED THE SAME MATCH and fail in the dangerous direction: "attacker@evil.test" with a zero-width space, or "uplift" with a soft hyphen, would have been reported clean. A permission test that cannot see the leak it is hunting is worse than none, because it is believed. normalizeForMatch folds those away, and its test pins that every case is one plain ToLower MISSES — a case whose naive match already succeeds fails, so the suite cannot fill with examples that demonstrate nothing. That caught my own first BOM case, which put the mark where Contains found it regardless. gofmt clean, vet clean, 15/15 packages pass offline; live suite green twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
140 lines
7.3 KiB
Plaintext
140 lines
7.3 KiB
Plaintext
# ============================================================================
|
|
# Krow API — Docker environment
|
|
#
|
|
# cd infrastructure && cp .env.docker.example .env
|
|
#
|
|
# Compose reads `.env` from the directory the compose file lives in, so this
|
|
# copy belongs in infrastructure/, NOT the repository root .env used by
|
|
# `make run`. Both are gitignored.
|
|
#
|
|
# Every value here is a placeholder. No real password or host belongs in a
|
|
# file that is committed.
|
|
# ============================================================================
|
|
|
|
# ── Application ─────────────────────────────────────────────────────────────
|
|
APP_ENV=production # development | staging | production
|
|
LOG_LEVEL=info # debug | info | warn | error
|
|
VERSION=1.0.0 # build arg only
|
|
IMAGE=doormile/krowbackend:latest # what `docker compose build` tags and `up` runs
|
|
|
|
# ── Where the API is published ──────────────────────────────────────────────
|
|
# Loopback by default: put a reverse proxy in front to terminate TLS. The API
|
|
# speaks plain HTTP and its session cookie is Secure outside development, so a
|
|
# browser will not send that cookie back over a plain connection anyway.
|
|
API_BIND=127.0.0.1
|
|
API_PORT=8080
|
|
|
|
# Browser origins allowed to call this API, comma-separated. Exact strings:
|
|
# scheme included, no trailing slash. Cookies are the auth mechanism, so
|
|
# anything listed here can hold a live session.
|
|
#
|
|
# "*" is REJECTED at startup — browsers refuse Allow-Origin: "*" together with
|
|
# credentials, so it would break every authenticated call rather than loosen
|
|
# anything.
|
|
HTTP_CORS_ORIGINS=https://platform.krowforce.com,https://mcp.krowforce.com
|
|
|
|
# SameSite on the session cookie: lax | none | strict.
|
|
#
|
|
# "lax" is right while the API is on *.krowforce.com — the two origins above
|
|
# are then cross-ORIGIN (CORS applies) but same-SITE (the cookie is still
|
|
# sent). Move the API to any other registrable domain and this must become
|
|
# "none", or login will succeed and every following request will arrive
|
|
# anonymous.
|
|
HTTP_COOKIE_SAMESITE=lax
|
|
|
|
# ── Database ────────────────────────────────────────────────────────────────
|
|
# Point at a managed PostgreSQL. With docker-compose.local-db.yml layered on
|
|
# top, set DATABASE_HOST=postgres instead.
|
|
DATABASE_HOST=your-database-host.example.com
|
|
DATABASE_PORT=5432
|
|
DATABASE_NAME=krow
|
|
DATABASE_USER=krow
|
|
DATABASE_PASSWORD=change-me
|
|
DATABASE_SCHEMA=public
|
|
|
|
# require encrypts but does not verify the server; verify-full also checks the
|
|
# certificate against a CA and is what you want against a managed database.
|
|
#
|
|
# `disable` is REJECTED at startup when APP_ENV=production. That is deliberate.
|
|
DATABASE_SSLMODE=require
|
|
|
|
# ── Migration URL ───────────────────────────────────────────────────────────
|
|
# golang-migrate takes a single URL, and Compose cannot assemble or encode one.
|
|
#
|
|
# ⚠️ PERCENT-ENCODE THE PASSWORD. A password containing @ : / ? # or % will
|
|
# otherwise be parsed as part of the host or the path, and the failure looks
|
|
# like a wrong host rather than a wrong password:
|
|
#
|
|
# @ → %40 : → %3A / → %2F ? → %3F # → %23 % → %25
|
|
#
|
|
# python3 -c 'import urllib.parse,sys; print(urllib.parse.quote(sys.argv[1], safe=""))' 'your-password'
|
|
#
|
|
# search_path must match DATABASE_SCHEMA above.
|
|
DATABASE_URL=postgres://krow:change-me@your-database-host.example.com:5432/krow?sslmode=require&search_path=public
|
|
|
|
# ── Pool and timeouts ───────────────────────────────────────────────────────
|
|
DATABASE_MAX_OPEN_CONNS=25
|
|
DATABASE_MIN_IDLE_CONNS=2
|
|
DATABASE_CONN_MAX_LIFETIME=30m
|
|
DATABASE_CONNECT_TIMEOUT=10s
|
|
DATABASE_STATEMENT_TIMEOUT=10s
|
|
|
|
# ── Model gateway ───────────────────────────────────────────────────────────
|
|
# THIS BLOCK WAS MISSING and a deployment copying this file could not start:
|
|
# APP_ENV=production with no MODEL_API_KEY is refused, because the alternative
|
|
# is an API that accepts agent runs and fails every one of them at the gateway.
|
|
#
|
|
# One wire protocol: the openai chat-completions shape. Groq, Gemini,
|
|
# OpenRouter, Together, vLLM and a local Ollama all serve it, so switching
|
|
# vendors is a base URL and a model id, not a code change.
|
|
#
|
|
# Empty MODEL_PROVIDER means openai — the only implementation. Empty
|
|
# MODEL_BASE_URL means Groq.
|
|
MODEL_PROVIDER=openai
|
|
MODEL_BASE_URL=https://api.groq.com/openai/v1
|
|
|
|
# REQUIRED in production. There is no ANTHROPIC_API_KEY fallback: that variable
|
|
# is now REFUSED at startup if it is set while this one is empty, because
|
|
# silently authenticating to Groq with a key named for a vendor this service
|
|
# cannot call is a lie the next operator has to unpick. Rename it here, and
|
|
# replace the value — an Anthropic key boots fine and then fails every run with
|
|
# 401, which the startup check cannot catch and only the model call can.
|
|
MODEL_API_KEY=
|
|
|
|
# Model ids must be ones MODEL_BASE_URL actually serves. A leftover claude-*
|
|
# id is refused at startup by name and tier: nothing configured serves one, so
|
|
# every run on that tier would 400 at the gateway.
|
|
MODEL_FAST=openai/gpt-oss-20b
|
|
MODEL_BALANCED=openai/gpt-oss-120b
|
|
MODEL_DEEP=openai/gpt-oss-120b
|
|
|
|
# Safe to turn on with the gpt-oss ids above: both accept reasoning_effort at
|
|
# low, medium and high, which is exactly the scale the gateway maps its three
|
|
# efforts onto. VERIFIED against the live Groq API, not assumed.
|
|
#
|
|
# It is NOT portable. qwen/qwen3.6-27b on the same account rejects all three
|
|
# ("must be one of `none` or `default`") and fails the whole request rather than
|
|
# ignoring the key, so switching model id and leaving this on breaks every run.
|
|
# Re-check it whenever MODEL_* changes.
|
|
MODEL_REASONING_EFFORT=
|
|
|
|
# ── HTTP timeouts ───────────────────────────────────────────────────────────
|
|
HTTP_READ_TIMEOUT=15s
|
|
# 180s, not 30s. internal/config REFUSES TO START when this is below the deep
|
|
# tier's 2m agent deadline: the server would abort the response mid-run and the
|
|
# caller would see 502 from the proxy in front, a gateway error for something no
|
|
# gateway did. 30s shipped here for a long time and was the cause of exactly
|
|
# that incident. Anything at or under 2m0s is a container that will not boot.
|
|
HTTP_WRITE_TIMEOUT=180s
|
|
HTTP_IDLE_TIMEOUT=60s
|
|
HTTP_SHUTDOWN_TIMEOUT=10s
|
|
|
|
# ── Container resources ─────────────────────────────────────────────────────
|
|
API_CPU_LIMIT=2
|
|
API_MEMORY_LIMIT=512M
|
|
|
|
# ── Local database only (docker-compose.local-db.yml) ───────────────────────
|
|
POSTGRES_BIND=127.0.0.1
|
|
POSTGRES_PORT=5432
|
|
POSTGRES_MEMORY_LIMIT=1G
|