Files
krow_backend/go-api/internal/runtime/wire.go
Aravind 43dabb5f72
Some checks failed
CI / test (push) Failing after 4m38s
CI / fixture (push) Failing after 8s
Switch memory on: the agent may keep one, and carries them into the next run
The store and the write trigger existed and nothing used either. This wires
both ends, so "we have long-term memory" stops being a statement about code
that exists and becomes one about behaviour.

READ. Memories are recalled before retrieval and placed before it in the
prompt: it is the smaller block and the more general one — a standing
preference frames how the documents should be read, where a document does not
frame a preference. The question stays last, because a model reads the last
thing and answers it, and evidence after the question becomes the prompt.

Skipped for smalltalk on the same terms as retrieval. Nobody needs remembering
to say good morning, and paying for it is how "hi" came to cost six thousand
tokens.

FAILS QUIET, RECORDED LOUDLY. A memory store that is unreachable must not take
the run with it: an answer without memory is worse, not wrong, and the
alternative is an outage in the knowledge layer becoming an outage in the
product. The trajectory records the failure, and records separately when the
store returned recency instead of relevance — a reader comparing two answers
needs to know which one got which.

WRITE. tools.Remember is registered only where there is somewhere to put it,
through DefaultToolsWithMemory rather than a nil check inside the old
constructor: a deployment that has not migrated 000017 must not offer a tool
whose every call fails against a table that is not there. Passing nil
registers exactly the catalogue that was there before.

Memory shares retrieval's embedder, and memory.Embedder is knowledge.Embedder
by structure so it cannot be given a different one. Two embedding models in
one deployment produce vectors that cannot be compared, and the failure is
silent: a recall that returns nothing rather than an error.

Five new tests on the loop, including the two that matter — a failing store
still answers, and a greeting carries nothing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-10-07 20:09:59 +05:30

193 lines
7.8 KiB
Go

package runtime
import (
"github.com/krow/krow-backend/go-api/internal/config"
"github.com/krow/krow-backend/go-api/internal/gateway"
"github.com/krow/krow-backend/go-api/internal/knowledge"
"github.com/krow/krow-backend/go-api/internal/memory"
"github.com/krow/krow-backend/go-api/internal/repo"
"github.com/krow/krow-backend/go-api/internal/tools"
)
// NewModelEngine builds the production runtime: the loader, the agent loop, a
// live model gateway and a Postgres trajectory sink.
//
// One call, because the alternative is four, and four assembled at a call site
// is how a deployment ends up running with a DiscardSink nobody chose. A test
// that wants a fake model still reaches for NewEngine with WithAgentExecutor —
// this function is the wiring, not a second way to configure the runtime.
//
// Skills keep the refusing stub. A skill has no executor of its own: the loop
// runs agents, and a skill reaches a model only as a capability an agent
// carries. Handing SkillExec a model would create a second, unbounded path to
// one — which is exactly the shape I3 exists to prevent.
func NewModelEngine(db repo.Querier, cfg config.Config) *Engine {
// gateway.New, not NewAnthropic: which provider answers is a deployment
// decision now, and hardcoding the constructor here would have meant every
// alternative provider needed an edit to this file to be reachable.
gw := gateway.New(gateway.FromConfig(cfg.Model))
retriever := knowledge.NewRetriever(db, NewEmbedder(cfg))
// Long-term memory shares the embedder with retrieval, deliberately: two
// embedding models in one deployment produce vectors that cannot be
// compared, and the failure is silent — a recall that quietly returns
// nothing rather than an error.
memories := memory.New(db, NewEmbedder(cfg))
exec := NewModelExecutor(gw, NewPostgresSink(db), DefaultToolsWithMemory(db, retriever, memories)).
WithRetriever(retriever).
WithMemory(memories).
// Without this, a spec's `subagents:` parses, loads, and is then
// dropped — which is how krow-workforce-agent came to declare five
// subagents and answer every question by itself. The resolver is the
// same Loader the engine uses, so a subagent is loaded exactly the way
// a directly-invoked agent is: same skills, same dependency checks.
WithSubagents(NewLoader(db))
return NewEngine(db, WithAgentExecutor(exec))
}
// NewEmbedder picks the embedding provider from configuration.
//
// Explicit first, then what is configured, then nothing. The order is the whole
// design: three providers all return vectors and retrieval works with any of
// them, so a deployment running the wrong one looks identical to one running
// the right one until somebody phrases a question differently. Naming the
// provider is how that stops being a silent condition.
//
// ollama A model on this machine. Real semantics, no credential, no
// per-token cost, no tenant text leaving the host. The default
// worth reaching for.
// voyage Hosted. Better on subtle retrieval over a large messy corpus,
// and the only one that needs a credential.
// lexical The deterministic stand-in. NOT semantic — it matches shared
// vocabulary and nothing else. Development only; config.validate
// refuses it in production.
//
// Returns nil when nothing is configured, and retrieval then runs keyword-only,
// saying so on every result. Nil rather than a hosted client with an empty key:
// both end up keyword-only, but nil says "no embedder is configured" once, at
// wiring time, instead of failing an HTTP call per query to learn the same
// thing.
func NewEmbedder(cfg config.Config) knowledge.Embedder {
k := cfg.Knowledge
provider := k.EmbedProvider
if provider == "" {
// Nothing named. Infer from what is actually present, preferring the
// one that costs nothing and keeps text local.
switch {
case k.UseLexicalEmbedder:
provider = "lexical"
case k.EmbedBaseURL != "":
provider = "ollama"
case k.EmbedAPIKey != "":
provider = "voyage"
default:
return nil
}
}
switch provider {
case "ollama":
return knowledge.NewOllama(k.EmbedBaseURL, k.EmbedModel, k.EmbedDims)
case "voyage":
if k.EmbedAPIKey == "" {
// Named but unusable. Nil, so retrieval degrades honestly rather
// than failing a request per query on a credential nobody set.
return nil
}
model, dims := k.EmbedModel, k.EmbedDims
if model == "" {
model = knowledge.DefaultVoyageModel
}
if dims == 0 {
dims = knowledge.DefaultVoyageDims
}
return knowledge.NewVoyage(k.EmbedAPIKey, model, dims)
case "lexical":
dims := k.EmbedDims
if dims == 0 {
dims = 256
}
e := knowledge.NewLexical(dims)
// Told what environment it is in, so its own refusal is the backstop
// behind config.validate's.
e.Production = cfg.AppEnv == "production"
return e
}
return nil
}
// DefaultTools is the tool registry this service ships with.
//
// One function, so "which tools exist" has a single answer that a test and the
// server reach the same way. Registration panics on a malformed tool: a
// service that booted without a capability its specs name would fail one run
// at a time instead of once, loudly, at startup.
func DefaultTools(db repo.Querier, retriever *knowledge.Retriever) *tools.Registry {
return DefaultToolsWithMemory(db, retriever, nil)
}
// DefaultToolsWithMemory is DefaultTools with long-term memory available.
//
// A separate constructor rather than a nil check inside the old one, because a
// deployment that has not migrated 000017 must not offer a tool whose every
// call would fail against a table that is not there. Passing nil registers the
// catalogue exactly as it was.
func DefaultToolsWithMemory(db repo.Querier, retriever *knowledge.Retriever, memories tools.MemoryWriter) *tools.Registry {
// The confirmation store is Postgres-backed, not in-process. A pending
// write is asked about in one request and approved in another, and nothing
// guarantees those two reach the same replica — an in-memory store would
// refuse a large share of perfectly good approvals, for a reason invisible
// to the person clicking. See tools.MemoryStore's own warning.
reg := tools.NewRegistryWithStore(tools.NewPostgresStore(db))
for _, t := range []tools.Tool{
// Activity
tools.ActivityBreakdown(db),
tools.ActivitySignals(db),
// Workforce
tools.WorkforceAttendance(db),
tools.WorkforceOvertime(db),
tools.WorkforceCoverage(db),
tools.WorkforceTraining(db),
// Hiring
tools.CandidatesQuality(db),
tools.HiresRecent(db),
tools.HiresPerformance(db),
tools.PositionsRisk(db),
tools.TalentPool(db),
// Cross-domain
tools.WorkspaceSummary(db),
tools.OperationsRisk(db),
// Assignments: the two lookups that yield ids, and the one write that
// consumes them. assign_worker is the only tool here with an effect,
// and it cannot run without an approval — see tools/confirm.go.
tools.OpenPositions(db),
tools.AvailableWorkers(db),
tools.AssignWorker(db),
// The hiring funnel: the lookup that yields application ids, and the
// write that moves somebody through it. Replaces the browser panel's
// interview matcher, which was the one capability the old templates had
// that the tool layer did not.
tools.CandidatesAwaiting(db),
tools.MoveApplication(db),
// Knowledge. Registered once; which corpora it may read comes from the
// running agent's spec by way of the tool Context, so this single
// registration serves every agent without any of them being able to
// name another's documents.
tools.KnowledgeSearch(retriever),
} {
reg.MustRegister(t)
}
// Memory, only where there is somewhere to put it. It is a confirmed
// write like assign_worker: it stores personal data that shapes later
// hiring answers, so a person sees the sentence before it is kept.
if memories != nil {
reg.MustRegister(tools.Remember(memories))
}
return reg
}