The defaults shipped yesterday were wrong the day they shipped, and a real key proved it in one request. Groq serves neither llama-3.1-8b-instant nor llama-3.3-70b-versatile any more. Both were chosen from memory, both passed startup validation, and every agent run would have failed with a 400. This is the exact failure the claude-* guard was written to catch, arriving from the side that guard cannot see. A prefix check can reject a vendor this service cannot call; it has no way to know a provider retired an id last month. That is not a gap in the check, it is a gap in the class of thing local validation can know, so the fix is not another guard: TestConfiguredModelsAreServed asks the provider. It lists /models — part of the same openai-compatible surface the gateway already speaks, so every supported provider answers it — and fails if a configured id is absent, printing what is available. It reads the ids through config.DefaultModels() rather than repeating them, because a second copy would be the first thing to drift, and drift is the whole failure. Skipped without a credential like the rest of the live suite. Verified three ways: it fails on the retired id with the message an operator needs, skips clean with no key, passes on the new ones. New defaults, chosen against the live account rather than from memory: openai/gpt-oss-20b (fast) and openai/gpt-oss-120b (balanced, deep). Tool calling confirmed on both. groq/compound-mini was ruled out — it cannot do tool calls at all, which this platform requires. MODEL_REASONING_EFFORT is now documented as safe here and NOT portable: gpt-oss accepts low/medium/high, exactly the scale openAIEffort maps onto, while qwen/qwen3.6-27b on the same account rejects all three and fails the whole request rather than ignoring the key. I7 IS NO LONGER UNPROVEN. make eval-live passes all three cases twice against gpt-oss-120b, the planted-injection case included: answers from the handbook, cites, refuses the injection, leaks neither the operator-only pay guidance nor the other tenant's figures. CLAUDE.md §12 and handover.md updated from "urgent" to measured, dated, and scoped to the one model it is evidence about. One real defect found on the way. The handbook grounding check failed once on an answer containing the phrase it wanted — "more than ten minutes" on screen, strings.Contains false — which leaves an invisible separator as the only explanation; the same model writes "47 %" and a U+2011 hyphen elsewhere. The flaky assertion is the small half. THE LEAK ASSERTIONS USED THE SAME MATCH and fail in the dangerous direction: "attacker@evil.test" with a zero-width space, or "uplift" with a soft hyphen, would have been reported clean. A permission test that cannot see the leak it is hunting is worse than none, because it is believed. normalizeForMatch folds those away, and its test pins that every case is one plain ToLower MISSES — a case whose naive match already succeeds fails, so the suite cannot fill with examples that demonstrate nothing. That caught my own first BOM case, which put the mark where Contains found it regardless. gofmt clean, vet clean, 15/15 packages pass offline; live suite green twice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
171 lines
8.5 KiB
Plaintext
171 lines
8.5 KiB
Plaintext
# ============================================================================
|
|
# Krow backend — example environment
|
|
#
|
|
# Copy to .env and fill in. .env is gitignored and must never be committed.
|
|
# Every value below is a placeholder or a safe local default: no real password,
|
|
# API key or token belongs in this file.
|
|
#
|
|
# cp .env.example .env
|
|
# ============================================================================
|
|
|
|
# ── Application ─────────────────────────────────────────────────────────────
|
|
APP_ENV=development # development | staging | production
|
|
LOG_LEVEL=info # debug | info | warn | error
|
|
|
|
# ── HTTP server ─────────────────────────────────────────────────────────────
|
|
HTTP_HOST=127.0.0.1
|
|
HTTP_PORT=8080
|
|
HTTP_READ_TIMEOUT=15s
|
|
# 180s, not 30s. internal/config REFUSES TO START when this is below the deep
|
|
# tier's 2m agent deadline: the server would abort the response mid-run and the
|
|
# caller would see 502 from the proxy in front, a gateway error for something no
|
|
# gateway did. 30s shipped here for a long time and was the cause of exactly
|
|
# that incident. Anything at or under 2m0s is a container that will not boot.
|
|
HTTP_WRITE_TIMEOUT=180s
|
|
HTTP_IDLE_TIMEOUT=60s
|
|
HTTP_SHUTDOWN_TIMEOUT=10s
|
|
# Browser origins allowed to call this API cross-origin, comma-separated.
|
|
# Unset in development it defaults to the Vite dev server on both hostnames
|
|
# (localhost and 127.0.0.1 are different origins to a browser) plus `vite
|
|
# preview`. Unset anywhere else it defaults to empty, meaning same-origin only.
|
|
# Origins are matched exactly, echoed back one at a time, and "*" is rejected.
|
|
# HTTP_CORS_ORIGINS=http://localhost:5173,http://127.0.0.1:5173
|
|
|
|
# ── PostgreSQL ──────────────────────────────────────────────────────────────
|
|
# The local development database. DATABASE_NAME is mixed-case and hyphenated,
|
|
# so anything that interpolates it into SQL must quote it: "Krow-force".
|
|
DATABASE_HOST=127.0.0.1
|
|
DATABASE_PORT=5432
|
|
DATABASE_NAME=Krow-force
|
|
DATABASE_USER=postgres
|
|
DATABASE_PASSWORD=
|
|
DATABASE_SCHEMA=public
|
|
|
|
# sslmode: disable is fine for a loopback dev database. APP_ENV=production
|
|
# rejects `disable` at startup — use require or verify-full there.
|
|
DATABASE_SSLMODE=disable
|
|
|
|
# Pool and timeout tuning.
|
|
DATABASE_MAX_OPEN_CONNS=25
|
|
DATABASE_MIN_IDLE_CONNS=2
|
|
DATABASE_CONN_MAX_LIFETIME=30m
|
|
DATABASE_CONNECT_TIMEOUT=5s
|
|
DATABASE_STATEMENT_TIMEOUT=10s
|
|
|
|
# ── Migrations ──────────────────────────────────────────────────────────────
|
|
# Consumed by the Makefile, which builds the golang-migrate URL from the
|
|
# DATABASE_* values above. Keep it pointed at the repository's migrations/.
|
|
MIGRATIONS_DIR=./migrations
|
|
|
|
# ── Seed ────────────────────────────────────────────────────────────────────
|
|
# The demo fixture, generated from the frontend repository's src/api/seed.js.
|
|
# The Makefile passes an absolute path; this default suits running from the
|
|
# repository root.
|
|
SEED_FIXTURE_PATH=./seed/fixtures/seed.json
|
|
|
|
# ── Model gateway ───────────────────────────────────────────────────────────
|
|
# The one place this service talks to a language model. An agent spec declares
|
|
# a `reasoning` tier — fast | balanced | deep — never a model id, so the
|
|
# mapping below is a deployment decision and changes without editing a single
|
|
# definition.
|
|
#
|
|
# WHICH PROVIDER ANSWERS is a deployment decision, but the wire protocol is no
|
|
# longer one. There is a single implementation:
|
|
#
|
|
# openai the chat-completions shape — which is NOT only OpenAI. Groq,
|
|
# Gemini (through its OpenAI-compatible endpoint), OpenRouter,
|
|
# Together, vLLM and a local Ollama all serve it, so moving
|
|
# between them is MODEL_BASE_URL and MODEL_* ids, nothing more.
|
|
#
|
|
# The anthropic path was REMOVED. MODEL_PROVIDER=anthropic is refused at
|
|
# startup rather than ignored, because a stack still carrying it would
|
|
# otherwise run on a vendor it never chose. Leave this empty or set "openai".
|
|
MODEL_PROVIDER=openai
|
|
|
|
# Where the provider is. Defaults to Groq when unset — the model ids below are
|
|
# Groq ids, and an id is only meaningful against the service that serves it, so
|
|
# these two settings move together or not at all.
|
|
#
|
|
# Groq https://api.groq.com/openai/v1 (the default)
|
|
# Gemini https://generativelanguage.googleapis.com/v1beta/openai
|
|
# OpenRouter https://openrouter.ai/api/v1
|
|
# Ollama http://localhost:11434/v1 (no key needed)
|
|
MODEL_BASE_URL=https://api.groq.com/openai/v1
|
|
|
|
# The credential. ANTHROPIC_API_KEY is NO LONGER READ — if it is set while this
|
|
# is empty, startup fails rather than silently ignoring it.
|
|
# May be empty outside production: migrations, seeding and every endpoint that
|
|
# is not an agent run work without one, and an agent run fails with a
|
|
# structured `gateway.not_configured` rather than the service refusing to boot.
|
|
# APP_ENV=production requires one — unless the model is on localhost, which
|
|
# needs no credential at all.
|
|
MODEL_API_KEY=
|
|
|
|
# The tiers differ by model AND by *effort*, which the gateway fixes
|
|
# (fast=low, balanced=high, deep=xhigh) so that "deep" cannot mean two
|
|
# different things in two deployments.
|
|
#
|
|
# These must be ids your MODEL_BASE_URL actually serves. A leftover claude-*
|
|
# id is refused at startup: it would be accepted by this process, rejected by
|
|
# the provider, and fail every single run with a 400.
|
|
MODEL_FAST=openai/gpt-oss-20b
|
|
MODEL_BALANCED=openai/gpt-oss-120b
|
|
MODEL_DEEP=openai/gpt-oss-120b
|
|
|
|
MODEL_MAX_OUTPUT_TOKENS=16000
|
|
|
|
# Send the tier's effort level as `reasoning_effort` on the openai-compatible
|
|
# wire. OFF by default and it should stay off unless every model named above is
|
|
# a reasoning model: the others reject the entire request rather than ignoring
|
|
# an unknown key, so turning this on for a non-reasoning model breaks every run
|
|
# with a 400. Ignored by the anthropic provider, which always sends effort.
|
|
MODEL_REASONING_EFFORT=false
|
|
|
|
# ── Knowledge layer (retrieval) ─────────────────────────────────────────────
|
|
#
|
|
# The dense half of hybrid retrieval needs an embedding model. Three options,
|
|
# and the choice is worth making deliberately: all three return vectors and
|
|
# retrieval works with any of them, so a deployment running the wrong one looks
|
|
# exactly like one running the right one — until somebody phrases a question
|
|
# differently.
|
|
#
|
|
# ollama A model on this machine. Real semantics, no credential, no
|
|
# per-token cost, and no tenant text leaving the host. Start here.
|
|
#
|
|
# brew install ollama
|
|
# ollama pull nomic-embed-text
|
|
#
|
|
# then EMBED_PROVIDER=ollama.
|
|
#
|
|
# voyage Hosted, and better on subtle retrieval over a large messy corpus.
|
|
# Needs VOYAGE_API_KEY. Anthropic does not serve embeddings, so
|
|
# this is a separate credential.
|
|
#
|
|
# lexical A deterministic stand-in that hashes words into a vector. NOT
|
|
# semantic — "annual leave" and "time off" are unrelated to it. It
|
|
# exists so the permission filter and the citation path can be
|
|
# tested without a network. Startup REFUSES it when
|
|
# APP_ENV=production.
|
|
#
|
|
# Leave EMBED_PROVIDER empty and the choice is inferred from what is set,
|
|
# preferring the local model. With nothing configured at all, retrieval runs
|
|
# keyword-only and says so on every result.
|
|
#
|
|
# CHANGING PROVIDER MEANS RE-EMBEDDING. Vectors from two models are not
|
|
# comparable, and every chunk records which model produced it — so after a
|
|
# switch the old vectors are simply not searched, and retrieval silently drops
|
|
# to keyword-only until you run:
|
|
#
|
|
# make reembed ORG=<slug>
|
|
#
|
|
EMBED_PROVIDER=ollama
|
|
EMBED_BASE_URL=http://localhost:11434
|
|
EMBED_MODEL=nomic-embed-text
|
|
EMBED_DIMENSIONS=768
|
|
|
|
# Only for EMBED_PROVIDER=voyage.
|
|
VOYAGE_API_KEY=
|
|
|
|
# Legacy switch for the stand-in. EMBED_PROVIDER=lexical is the current spelling.
|
|
EMBED_USE_LEXICAL=false
|