Files
krow_backend/go-api/internal/runtime/wire.go
Suriyakumarvijayanayagam c74fe7e074 Add an OpenAI-compatible gateway, so the model provider is a config value
The platform could only talk to one vendor. Moving off Claude — for cost, or
because a client asks for Gemini — meant a rewrite behind an interface that
already had exactly the right shape and one implementation.

`openai` is not only OpenAI. Groq, Gemini's compatibility endpoint, OpenRouter,
Together, vLLM and a local Ollama all serve the chat-completions shape, so one
implementation reaches all of them and the difference between them is a base
URL and three model ids. That is why this is one file and not a package per
vendor.

`routing.go` had the vendor baked into the routing table every provider has to
read: effort was `anthropic.OutputConfigEffort`. Nothing was wrong with that
while there was one implementation; it became wrong the moment there were two,
because the OpenAI path would have had to import the Anthropic SDK to learn how
hard to think. Effort is now the platform's own three-value vocabulary and each
implementation maps it onto whatever its API calls the same idea.

THE ACCOUNTING DIFFERS BETWEEN THE TWO WIRES, and getting it wrong would have
been invisible. OpenAI reports prompt_tokens INCLUSIVE of the cached prefix;
Anthropic reports input tokens EXCLUSIVE of it and carries the cache
separately. Usage.Total() adds all four fields, so copying both numbers across
verbatim bills the cached prefix twice — worst on long conversations, which is
exactly where I3's budget matters most. The run would still answer; it would
just hit BudgetExceeded early, for no visible reason. normalise() subtracts,
and there is a test named after it.

Streamed tool calls are keyed by their wire index, not appended in arrival
order. Providers interleave the fragments of parallel calls, so appending
splices one call's arguments onto another's — and the result is usually two
calls that are each valid JSON and both wrong, which means the tools run with
inputs the model never chose and nothing errors. Mutation-checked: ignoring the
index produces `{"day"{"week":"friday"}:"next"}` and the test catches it.

Three configuration mistakes are refused at startup rather than at runtime:

  - MODEL_BASE_URL without MODEL_PROVIDER=openai. The anthropic path has one
    endpoint and ignores the field, so this is a deployment that believes it
    switched providers and did not — every run still goes to Anthropic and is
    still billed there, with nothing in the logs to say so. Cost is the whole
    reason this change exists, and that is the one mistake that silently
    defeats it.
  - An unrecognised MODEL_PROVIDER, once at boot instead of once per run.
  - A production deployment with no credential — except against localhost,
    which needs none, and demanding one would make the free local path
    impossible to configure.

reasoning_effort is opt-in via MODEL_REASONING_EFFORT. Reasoning models accept
it; most others reject the entire request with a 400 rather than ignoring an
unknown key, so every deployment would have had to opt out instead.

`make eval-live` now reads the same environment the service does and logs which
provider answered, because a suite that cannot say which model produced a
result is a suite whose result cannot be compared with another run's. That is
the point of this change: §12 leaves model hosting open, and this makes the
decision cheap to reverse and possible to settle on evidence. Weigh the I7 case
heaviest — a cheaper model that follows the planted injection is a security
regression, not a saving.

Default behaviour is unchanged: MODEL_PROVIDER unset means anthropic, and
ANTHROPIC_API_KEY still works, so no existing deployment needs an edit.

NOT verified against a live provider — no credential was available on this
machine. Tested against a fake endpoint covering both paths, and the three
guarantees above are mutation-checked.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-09-01 11:47:53 +05:30

167 lines
6.6 KiB
Go

package runtime
import (
"github.com/krow/krow-backend/go-api/internal/config"
"github.com/krow/krow-backend/go-api/internal/gateway"
"github.com/krow/krow-backend/go-api/internal/knowledge"
"github.com/krow/krow-backend/go-api/internal/repo"
"github.com/krow/krow-backend/go-api/internal/tools"
)
// NewModelEngine builds the production runtime: the loader, the agent loop, a
// live model gateway and a Postgres trajectory sink.
//
// One call, because the alternative is four, and four assembled at a call site
// is how a deployment ends up running with a DiscardSink nobody chose. A test
// that wants a fake model still reaches for NewEngine with WithAgentExecutor —
// this function is the wiring, not a second way to configure the runtime.
//
// Skills keep the refusing stub. A skill has no executor of its own: the loop
// runs agents, and a skill reaches a model only as a capability an agent
// carries. Handing SkillExec a model would create a second, unbounded path to
// one — which is exactly the shape I3 exists to prevent.
func NewModelEngine(db repo.Querier, cfg config.Config) *Engine {
// gateway.New, not NewAnthropic: which provider answers is a deployment
// decision now, and hardcoding the constructor here would have meant every
// alternative provider needed an edit to this file to be reachable.
gw := gateway.New(gateway.FromConfig(cfg.Model))
retriever := knowledge.NewRetriever(db, NewEmbedder(cfg))
exec := NewModelExecutor(gw, NewPostgresSink(db), DefaultTools(db, retriever)).
WithRetriever(retriever).
// Without this, a spec's `subagents:` parses, loads, and is then
// dropped — which is how krow-workforce-agent came to declare five
// subagents and answer every question by itself. The resolver is the
// same Loader the engine uses, so a subagent is loaded exactly the way
// a directly-invoked agent is: same skills, same dependency checks.
WithSubagents(NewLoader(db))
return NewEngine(db, WithAgentExecutor(exec))
}
// NewEmbedder picks the embedding provider from configuration.
//
// Explicit first, then what is configured, then nothing. The order is the whole
// design: three providers all return vectors and retrieval works with any of
// them, so a deployment running the wrong one looks identical to one running
// the right one until somebody phrases a question differently. Naming the
// provider is how that stops being a silent condition.
//
// ollama A model on this machine. Real semantics, no credential, no
// per-token cost, no tenant text leaving the host. The default
// worth reaching for.
// voyage Hosted. Better on subtle retrieval over a large messy corpus,
// and the only one that needs a credential.
// lexical The deterministic stand-in. NOT semantic — it matches shared
// vocabulary and nothing else. Development only; config.validate
// refuses it in production.
//
// Returns nil when nothing is configured, and retrieval then runs keyword-only,
// saying so on every result. Nil rather than a hosted client with an empty key:
// both end up keyword-only, but nil says "no embedder is configured" once, at
// wiring time, instead of failing an HTTP call per query to learn the same
// thing.
func NewEmbedder(cfg config.Config) knowledge.Embedder {
k := cfg.Knowledge
provider := k.EmbedProvider
if provider == "" {
// Nothing named. Infer from what is actually present, preferring the
// one that costs nothing and keeps text local.
switch {
case k.UseLexicalEmbedder:
provider = "lexical"
case k.EmbedBaseURL != "":
provider = "ollama"
case k.EmbedAPIKey != "":
provider = "voyage"
default:
return nil
}
}
switch provider {
case "ollama":
return knowledge.NewOllama(k.EmbedBaseURL, k.EmbedModel, k.EmbedDims)
case "voyage":
if k.EmbedAPIKey == "" {
// Named but unusable. Nil, so retrieval degrades honestly rather
// than failing a request per query on a credential nobody set.
return nil
}
model, dims := k.EmbedModel, k.EmbedDims
if model == "" {
model = knowledge.DefaultVoyageModel
}
if dims == 0 {
dims = knowledge.DefaultVoyageDims
}
return knowledge.NewVoyage(k.EmbedAPIKey, model, dims)
case "lexical":
dims := k.EmbedDims
if dims == 0 {
dims = 256
}
e := knowledge.NewLexical(dims)
// Told what environment it is in, so its own refusal is the backstop
// behind config.validate's.
e.Production = cfg.AppEnv == "production"
return e
}
return nil
}
// DefaultTools is the tool registry this service ships with.
//
// One function, so "which tools exist" has a single answer that a test and the
// server reach the same way. Registration panics on a malformed tool: a
// service that booted without a capability its specs name would fail one run
// at a time instead of once, loudly, at startup.
func DefaultTools(db repo.Querier, retriever *knowledge.Retriever) *tools.Registry {
// The confirmation store is Postgres-backed, not in-process. A pending
// write is asked about in one request and approved in another, and nothing
// guarantees those two reach the same replica — an in-memory store would
// refuse a large share of perfectly good approvals, for a reason invisible
// to the person clicking. See tools.MemoryStore's own warning.
reg := tools.NewRegistryWithStore(tools.NewPostgresStore(db))
for _, t := range []tools.Tool{
// Activity
tools.ActivityBreakdown(db),
tools.ActivitySignals(db),
// Workforce
tools.WorkforceAttendance(db),
tools.WorkforceOvertime(db),
tools.WorkforceCoverage(db),
tools.WorkforceTraining(db),
// Hiring
tools.CandidatesQuality(db),
tools.HiresRecent(db),
tools.HiresPerformance(db),
tools.PositionsRisk(db),
tools.TalentPool(db),
// Cross-domain
tools.WorkspaceSummary(db),
tools.OperationsRisk(db),
// Assignments: the two lookups that yield ids, and the one write that
// consumes them. assign_worker is the only tool here with an effect,
// and it cannot run without an approval — see tools/confirm.go.
tools.OpenPositions(db),
tools.AvailableWorkers(db),
tools.AssignWorker(db),
// The hiring funnel: the lookup that yields application ids, and the
// write that moves somebody through it. Replaces the browser panel's
// interview matcher, which was the one capability the old templates had
// that the tool layer did not.
tools.CandidatesAwaiting(db),
tools.MoveApplication(db),
// Knowledge. Registered once; which corpora it may read comes from the
// running agent's spec by way of the tool Context, so this single
// registration serves every agent without any of them being able to
// name another's documents.
tools.KnowledgeSearch(retriever),
} {
reg.MustRegister(t)
}
return reg
}