agent build

This commit is contained in:
2026-08-28 12:21:44 +05:30
parent b6f8655909
commit f7df96c973
138 changed files with 24164 additions and 207 deletions

View File

@@ -57,3 +57,75 @@ MIGRATIONS_DIR=./migrations
# The Makefile passes an absolute path; this default suits running from the
# repository root.
SEED_FIXTURE_PATH=./seed/fixtures/seed.json
# ── Model gateway ───────────────────────────────────────────────────────────
# The one place this service talks to a language model. An agent spec declares
# a `reasoning` tier — fast | balanced | deep — never a model id, so the
# mapping below is a deployment decision and changes without editing a single
# definition.
#
# The key may be left empty outside production: migrations, seeding and every
# endpoint that is not an agent run work without one, and an agent run fails
# with a structured `gateway.not_configured` rather than the service refusing
# to boot. APP_ENV=production requires it.
ANTHROPIC_API_KEY=
# All three tiers default to the same model. They differ by *effort*, which the
# gateway fixes (fast=low, balanced=high, deep=xhigh) so that "deep" cannot
# mean two different things in two deployments. Point a tier at a different
# model only as a deliberate choice — never as a silent cost downgrade.
MODEL_FAST=claude-opus-5
MODEL_BALANCED=claude-opus-5
MODEL_DEEP=claude-opus-5
# Hard ceiling on a single unstreamed response. Not the run's token budget —
# that spans every call in a run and belongs to the runtime.
MODEL_MAX_OUTPUT_TOKENS=16000
# ── Knowledge layer (retrieval) ─────────────────────────────────────────────
#
# The dense half of hybrid retrieval needs an embedding model. Three options,
# and the choice is worth making deliberately: all three return vectors and
# retrieval works with any of them, so a deployment running the wrong one looks
# exactly like one running the right one — until somebody phrases a question
# differently.
#
# ollama A model on this machine. Real semantics, no credential, no
# per-token cost, and no tenant text leaving the host. Start here.
#
# brew install ollama
# ollama pull nomic-embed-text
#
# then EMBED_PROVIDER=ollama.
#
# voyage Hosted, and better on subtle retrieval over a large messy corpus.
# Needs VOYAGE_API_KEY. Anthropic does not serve embeddings, so
# this is a separate credential.
#
# lexical A deterministic stand-in that hashes words into a vector. NOT
# semantic — "annual leave" and "time off" are unrelated to it. It
# exists so the permission filter and the citation path can be
# tested without a network. Startup REFUSES it when
# APP_ENV=production.
#
# Leave EMBED_PROVIDER empty and the choice is inferred from what is set,
# preferring the local model. With nothing configured at all, retrieval runs
# keyword-only and says so on every result.
#
# CHANGING PROVIDER MEANS RE-EMBEDDING. Vectors from two models are not
# comparable, and every chunk records which model produced it — so after a
# switch the old vectors are simply not searched, and retrieval silently drops
# to keyword-only until you run:
#
# make reembed ORG=<slug>
#
EMBED_PROVIDER=ollama
EMBED_BASE_URL=http://localhost:11434
EMBED_MODEL=nomic-embed-text
EMBED_DIMENSIONS=768
# Only for EMBED_PROVIDER=voyage.
VOYAGE_API_KEY=
# Legacy switch for the stand-in. EMBED_PROVIDER=lexical is the current spelling.
EMBED_USE_LEXICAL=false