Files
krow_backend/.env.example
Suriyakumarvijayanayagam c74fe7e074 Add an OpenAI-compatible gateway, so the model provider is a config value
The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.

`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.

`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.

THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.

Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.

Three configuration mistakes are refused at startup rather than at runtime:

  - MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
    endpoint and ignores the field, so this is a deployment that believes it
    switched providers and did not — every run still goes to Anthropic and is
    still billed there, with nothing in the logs to say so. Cost is the whole
    reason this change exists, and that is the one mistake that silently
    defeats it.
  - An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
  - A production deployment with no credential — except against localhost,
    which needs none, and demanding one would make the free local path
    impossible to configure.

reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.

`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.

Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.

NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-01 11:47:53 +05:30

163 lines
8.1 KiB
Plaintext

# ============================================================================
# Krow backend — example environment
#
# Copy to .env and fill in. .env is gitignored and must never be committed.
# Every value below is a placeholder or a safe local default: no real password,
# API key or token belongs in this file.
#
# cp .env.example .env
# ============================================================================
# ── Application ─────────────────────────────────────────────────────────────
APP_ENV=development # development | staging | production
LOG_LEVEL=info # debug | info | warn | error
# ── HTTP server ─────────────────────────────────────────────────────────────
HTTP_HOST=127.0.0.1
HTTP_PORT=8080
HTTP_READ_TIMEOUT=15s
HTTP_WRITE_TIMEOUT=30s
HTTP_IDLE_TIMEOUT=60s
HTTP_SHUTDOWN_TIMEOUT=10s
# Browser origins allowed to call this API cross-origin, comma-separated.
# Unset in development it defaults to the Vite dev server on both hostnames
# (localhost and 127.0.0.1 are different origins to a browser) plus `vite
# preview`. Unset anywhere else it defaults to empty, meaning same-origin only.
# Origins are matched exactly, echoed back one at a time, and "*" is rejected.
# HTTP_CORS_ORIGINS=http://localhost:5173,http://127.0.0.1:5173
# ── PostgreSQL ──────────────────────────────────────────────────────────────
# The local development database. DATABASE_NAME is mixed-case and hyphenated,
# so anything that interpolates it into SQL must quote it: "Krow-force".
DATABASE_HOST=127.0.0.1
DATABASE_PORT=5432
DATABASE_NAME=Krow-force
DATABASE_USER=postgres
DATABASE_PASSWORD=
DATABASE_SCHEMA=public
# sslmode: disable is fine for a loopback dev database. APP_ENV=production
# rejects `disable` at startup — use require or verify-full there.
DATABASE_SSLMODE=disable
# Pool and timeout tuning.
DATABASE_MAX_OPEN_CONNS=25
DATABASE_MIN_IDLE_CONNS=2
DATABASE_CONN_MAX_LIFETIME=30m
DATABASE_CONNECT_TIMEOUT=5s
DATABASE_STATEMENT_TIMEOUT=10s
# ── Migrations ──────────────────────────────────────────────────────────────
# Consumed by the Makefile, which builds the golang-migrate URL from the
# DATABASE_* values above. Keep it pointed at the repository's migrations/.
MIGRATIONS_DIR=./migrations
# ── Seed ────────────────────────────────────────────────────────────────────
# The demo fixture, generated from the frontend repository's src/api/seed.js.
# The Makefile passes an absolute path; this default suits running from the
# repository root.
SEED_FIXTURE_PATH=./seed/fixtures/seed.json
# ── Model gateway ───────────────────────────────────────────────────────────
# The one place this service talks to a language model. An agent spec declares
# a `reasoning` tier — fast | balanced | deep — never a model id, so the
# mapping below is a deployment decision and changes without editing a single
# definition.
#
# WHICH PROVIDER ANSWERS is a deployment decision. Two wire protocols:
#
# anthropic the Claude API. The default, and what an unset value means.
# openai the chat-completions shape — which is NOT only OpenAI. Groq,
# Gemini (through its OpenAI-compatible endpoint), OpenRouter,
# Together, vLLM and a local Ollama all serve it, so moving
# between them is MODEL_BASE_URL and MODEL_* ids, nothing more.
MODEL_PROVIDER=anthropic
# Where the openai-compatible provider points. IGNORED — and refused at
# startup — unless MODEL_PROVIDER=openai, because a base URL set against the
# anthropic provider is a deployment that believes it has switched and has not:
# every run would still go to Anthropic, and still be billed there.
#
# Groq https://api.groq.com/openai/v1
# Gemini https://generativelanguage.googleapis.com/v1beta/openai
# OpenRouter https://openrouter.ai/api/v1
# Ollama http://localhost:11434/v1 (no key needed)
MODEL_BASE_URL=
# The credential. MODEL_API_KEY is the provider-neutral name and wins;
# ANTHROPIC_API_KEY still works so no existing deployment needs an edit.
# Either may be empty outside production: migrations, seeding and every
# endpoint that is not an agent run work without one, and an agent run fails
# with a structured `gateway.not_configured` rather than the service refusing
# to boot. APP_ENV=production requires one — unless the model is on localhost,
# which needs no credential at all.
MODEL_API_KEY=
ANTHROPIC_API_KEY=
# All three tiers default to the same model. They differ by *effort*, which the
# gateway fixes (fast=low, balanced=high, deep=xhigh) so that "deep" cannot
# mean two different things in two deployments. Point a tier at a different
# model only as a deliberate choice — never as a silent cost downgrade.
MODEL_FAST=claude-opus-5
MODEL_BALANCED=claude-opus-5
MODEL_DEEP=claude-opus-5
# Hard ceiling on a single unstreamed response. Not the run's token budget —
# that spans every call in a run and belongs to the runtime.
MODEL_MAX_OUTPUT_TOKENS=16000
# Send the tier's effort level as `reasoning_effort` on the openai-compatible
# wire. OFF by default and it should stay off unless every model named above is
# a reasoning model: the others reject the entire request rather than ignoring
# an unknown key, so turning this on for a non-reasoning model breaks every run
# with a 400. Ignored by the anthropic provider, which always sends effort.
MODEL_REASONING_EFFORT=false
# ── Knowledge layer (retrieval) ─────────────────────────────────────────────
#
# The dense half of hybrid retrieval needs an embedding model. Three options,
# and the choice is worth making deliberately: all three return vectors and
# retrieval works with any of them, so a deployment running the wrong one looks
# exactly like one running the right one — until somebody phrases a question
# differently.
#
# ollama A model on this machine. Real semantics, no credential, no
# per-token cost, and no tenant text leaving the host. Start here.
#
# brew install ollama
# ollama pull nomic-embed-text
#
# then EMBED_PROVIDER=ollama.
#
# voyage Hosted, and better on subtle retrieval over a large messy corpus.
# Needs VOYAGE_API_KEY. Anthropic does not serve embeddings, so
# this is a separate credential.
#
# lexical A deterministic stand-in that hashes words into a vector. NOT
# semantic — "annual leave" and "time off" are unrelated to it. It
# exists so the permission filter and the citation path can be
# tested without a network. Startup REFUSES it when
# APP_ENV=production.
#
# Leave EMBED_PROVIDER empty and the choice is inferred from what is set,
# preferring the local model. With nothing configured at all, retrieval runs
# keyword-only and says so on every result.
#
# CHANGING PROVIDER MEANS RE-EMBEDDING. Vectors from two models are not
# comparable, and every chunk records which model produced it — so after a
# switch the old vectors are simply not searched, and retrieval silently drops
# to keyword-only until you run:
#
# make reembed ORG=<slug>
#
EMBED_PROVIDER=ollama
EMBED_BASE_URL=http://localhost:11434
EMBED_MODEL=nomic-embed-text
EMBED_DIMENSIONS=768
# Only for EMBED_PROVIDER=voyage.
VOYAGE_API_KEY=
# Legacy switch for the stand-in. EMBED_PROVIDER=lexical is the current spelling.
EMBED_USE_LEXICAL=false