Files
krow_backend/.env.example
Suriyakumarvijayanayagam 34fa58a6b9
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 7s
Remove the Anthropic path; the gateway speaks one wire protocol
The platform now runs on Groq by default, through the OpenAI-compatible
chat-completions shape. That shape is not one vendor — Gemini, OpenRouter,
Together, vLLM and a local Ollama serve it too — so moving again stays
configuration rather than code.

Two things in the deleted file were not Anthropic's and would have gone
with it silently:

  withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and
  CALLED BY openai.go. Deleting the file wholesale would have removed the
  retry policy of the provider that survived, and nothing in openai.go
  mentions it, so the loss would have been invisible until the next 429.
  The policy is a property of this platform's runs, not of a vendor's API;
  it now lives in retry.go where no provider can carry it off.

  StreamComplete had the same problem and moves to gateway.go, beside the
  Streamer interface whose comment already referenced it.

Three stale-configuration failures are now refused at startup instead of
being ignored. Each was verified firing through the real config.Load():

  MODEL_PROVIDER=anthropic — named separately from every other wrong value
  because it used to be correct. Ignoring it gives a stack that believes it
  is on Claude while every run goes to Groq and is billed there.

  ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an
  operator did set is the worst version of this: they fail every run on a
  missing credential they are looking straight at.

  A leftover claude-* model id, naming the tier that carries it. This is
  the check the previous commit's error-detail work was diagnosing: such an
  id is accepted by this process, rejected by the provider, and 400s on
  EVERY run. "A model is wrong" does not say which of three lines to edit.

Defaults ship as a matched pair. defaultBaseURL and the three tier ids are
one decision, not four: an id is only meaningful against the service that
serves it, and a Groq id on an OpenAI base URL is the same failure from the
other side. The tiers also stop being one model — a tier whose cost does
not differ is a distinction that buys nothing.

Verified end to end against a stub of the wire, driving the real wiring
(config.Load in production mode, gateway.New, StreamComplete): streamed
deltas, tool-call decoding, the loopback credential exemption, and usage
totalling 150 rather than 190 — the cached-prefix subtraction still holds.

gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still
needs a reachable database.

NOT verified: the I7 planted-injection eval. Removing this path removed the
only model whose refusal behaviour had been measured against it, so the new
default is unproven there until `make eval-live` runs with a real key. The
Groq model ids should also be confirmed against Groq's current lineup.
Flagged in CLAUDE.md §12 and docs/handover.md.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-05 11:52:21 +05:30

166 lines
8.2 KiB
Plaintext

# ============================================================================
# Krow backend — example environment
#
# Copy to .env and fill in. .env is gitignored and must never be committed.
# Every value below is a placeholder or a safe local default: no real password,
# API key or token belongs in this file.
#
# cp .env.example .env
# ============================================================================
# ── Application ─────────────────────────────────────────────────────────────
APP_ENV=development # development | staging | production
LOG_LEVEL=info # debug | info | warn | error
# ── HTTP server ─────────────────────────────────────────────────────────────
HTTP_HOST=127.0.0.1
HTTP_PORT=8080
HTTP_READ_TIMEOUT=15s
HTTP_WRITE_TIMEOUT=30s
HTTP_IDLE_TIMEOUT=60s
HTTP_SHUTDOWN_TIMEOUT=10s
# Browser origins allowed to call this API cross-origin, comma-separated.
# Unset in development it defaults to the Vite dev server on both hostnames
# (localhost and 127.0.0.1 are different origins to a browser) plus `vite
# preview`. Unset anywhere else it defaults to empty, meaning same-origin only.
# Origins are matched exactly, echoed back one at a time, and "*" is rejected.
# HTTP_CORS_ORIGINS=http://localhost:5173,http://127.0.0.1:5173
# ── PostgreSQL ──────────────────────────────────────────────────────────────
# The local development database. DATABASE_NAME is mixed-case and hyphenated,
# so anything that interpolates it into SQL must quote it: "Krow-force".
DATABASE_HOST=127.0.0.1
DATABASE_PORT=5432
DATABASE_NAME=Krow-force
DATABASE_USER=postgres
DATABASE_PASSWORD=
DATABASE_SCHEMA=public
# sslmode: disable is fine for a loopback dev database. APP_ENV=production
# rejects `disable` at startup — use require or verify-full there.
DATABASE_SSLMODE=disable
# Pool and timeout tuning.
DATABASE_MAX_OPEN_CONNS=25
DATABASE_MIN_IDLE_CONNS=2
DATABASE_CONN_MAX_LIFETIME=30m
DATABASE_CONNECT_TIMEOUT=5s
DATABASE_STATEMENT_TIMEOUT=10s
# ── Migrations ──────────────────────────────────────────────────────────────
# Consumed by the Makefile, which builds the golang-migrate URL from the
# DATABASE_* values above. Keep it pointed at the repository's migrations/.
MIGRATIONS_DIR=./migrations
# ── Seed ────────────────────────────────────────────────────────────────────
# The demo fixture, generated from the frontend repository's src/api/seed.js.
# The Makefile passes an absolute path; this default suits running from the
# repository root.
SEED_FIXTURE_PATH=./seed/fixtures/seed.json
# ── Model gateway ───────────────────────────────────────────────────────────
# The one place this service talks to a language model. An agent spec declares
# a `reasoning` tier — fast | balanced | deep — never a model id, so the
# mapping below is a deployment decision and changes without editing a single
# definition.
#
# WHICH PROVIDER ANSWERS is a deployment decision, but the wire protocol is no
# longer one. There is a single implementation:
#
# openai the chat-completions shape — which is NOT only OpenAI. Groq,
# Gemini (through its OpenAI-compatible endpoint), OpenRouter,
# Together, vLLM and a local Ollama all serve it, so moving
# between them is MODEL_BASE_URL and MODEL_* ids, nothing more.
#
# The anthropic path was REMOVED. MODEL_PROVIDER=anthropic is refused at
# startup rather than ignored, because a stack still carrying it would
# otherwise run on a vendor it never chose. Leave this empty or set "openai".
MODEL_PROVIDER=openai
# Where the provider is. Defaults to Groq when unset — the model ids below are
# Groq ids, and an id is only meaningful against the service that serves it, so
# these two settings move together or not at all.
#
# Groq https://api.groq.com/openai/v1 (the default)
# Gemini https://generativelanguage.googleapis.com/v1beta/openai
# OpenRouter https://openrouter.ai/api/v1
# Ollama http://localhost:11434/v1 (no key needed)
MODEL_BASE_URL=https://api.groq.com/openai/v1
# The credential. ANTHROPIC_API_KEY is NO LONGER READ — if it is set while this
# is empty, startup fails rather than silently ignoring it.
# May be empty outside production: migrations, seeding and every endpoint that
# is not an agent run work without one, and an agent run fails with a
# structured `gateway.not_configured` rather than the service refusing to boot.
# APP_ENV=production requires one — unless the model is on localhost, which
# needs no credential at all.
MODEL_API_KEY=
# The tiers differ by model AND by *effort*, which the gateway fixes
# (fast=low, balanced=high, deep=xhigh) so that "deep" cannot mean two
# different things in two deployments.
#
# These must be ids your MODEL_BASE_URL actually serves. A leftover claude-*
# id is refused at startup: it would be accepted by this process, rejected by
# the provider, and fail every single run with a 400.
MODEL_FAST=llama-3.1-8b-instant
MODEL_BALANCED=llama-3.3-70b-versatile
MODEL_DEEP=llama-3.3-70b-versatile
MODEL_MAX_OUTPUT_TOKENS=16000
# Send the tier's effort level as `reasoning_effort` on the openai-compatible
# wire. OFF by default and it should stay off unless every model named above is
# a reasoning model: the others reject the entire request rather than ignoring
# an unknown key, so turning this on for a non-reasoning model breaks every run
# with a 400. Ignored by the anthropic provider, which always sends effort.
MODEL_REASONING_EFFORT=false
# ── Knowledge layer (retrieval) ─────────────────────────────────────────────
#
# The dense half of hybrid retrieval needs an embedding model. Three options,
# and the choice is worth making deliberately: all three return vectors and
# retrieval works with any of them, so a deployment running the wrong one looks
# exactly like one running the right one — until somebody phrases a question
# differently.
#
# ollama A model on this machine. Real semantics, no credential, no
# per-token cost, and no tenant text leaving the host. Start here.
#
# brew install ollama
# ollama pull nomic-embed-text
#
# then EMBED_PROVIDER=ollama.
#
# voyage Hosted, and better on subtle retrieval over a large messy corpus.
# Needs VOYAGE_API_KEY. Anthropic does not serve embeddings, so
# this is a separate credential.
#
# lexical A deterministic stand-in that hashes words into a vector. NOT
# semantic — "annual leave" and "time off" are unrelated to it. It
# exists so the permission filter and the citation path can be
# tested without a network. Startup REFUSES it when
# APP_ENV=production.
#
# Leave EMBED_PROVIDER empty and the choice is inferred from what is set,
# preferring the local model. With nothing configured at all, retrieval runs
# keyword-only and says so on every result.
#
# CHANGING PROVIDER MEANS RE-EMBEDDING. Vectors from two models are not
# comparable, and every chunk records which model produced it — so after a
# switch the old vectors are simply not searched, and retrieval silently drops
# to keyword-only until you run:
#
# make reembed ORG=<slug>
#
EMBED_PROVIDER=ollama
EMBED_BASE_URL=http://localhost:11434
EMBED_MODEL=nomic-embed-text
EMBED_DIMENSIONS=768
# Only for EMBED_PROVIDER=voyage.
VOYAGE_API_KEY=
# Legacy switch for the stand-in. EMBED_PROVIDER=lexical is the current spelling.
EMBED_USE_LEXICAL=false