krow-workforce-agent has declared five subagents since it was written and
answered every question by itself. Everything for §6 existed except the
delegation: the parser read `subagents:`, runtime.Agent carried them, the
loader populated them, agent_runs had a parent_run_id column with a
self-reference and a no-self-parent constraint, and budget.go's comments
already described sharing a budget with subagents. Nothing called any of it.
A subagent is offered to the parent's model as a tool, because §6 says that is
what delegation is from the parent's side. Three rules are enforced rather than
assumed, each with a test that fails if it stops holding:
I1 The subagent runs as the ORIGINAL caller. It cannot read anything the
person could not read directly.
§6 It SHARES the parent's budget. The test sets MaxSteps to 1, spends it in
the parent, and asserts the child terminates BudgetExceeded — an
assertion that only passes when the budget is shared, and that a fresh
budget would quietly turn green.
§3 Depth is capped at 2. At the cap no subagent is loaded or offered, so a
cycle reaching run time is bounded rather than unbounded.
I4 survives too: a write a SUBAGENT wants approved still stops the whole run
and asks a person, rather than being performed because it happened one level
down.
Two bugs found by running it rather than by reading it:
- delegate() read the error before the result. finish returns a non-nil
error for every termination that is not Completed, INCLUDING
ConfirmationPending — which is not a failure but a run that stopped to ask
a question. Reading the error first discarded the result and with it the
confirmation, so a subagent's write silently never happened and nobody was
asked.
- Delegated trajectories were never persisted at all. parent_run_id is a
foreign key and a subagent finishes BEFORE the run that delegated to it,
so every child insert named a parent row that did not exist yet. The
database refused it; finish deliberately does not fail a run over a sink
error; and the entry recording that the trajectory could not be saved was
itself in the trajectory that was not saved. Children are now buffered and
written by finish after the parent's own row, each arriving with its
descendants already ordered behind it, so one pass writes a whole tree
parent-first. The regression test asserts on save ORDER, because a
MemorySink has no foreign key and will pass either way.
Verified end to end against a live model: an agent with no tools of its own and
one subagent produced
delegation-probe run=run_16622d7de6 parent=(root)
talent-pool-agent run=run_64160684b1 parent=run_16622d7de6
with the subagent's answer reaching the parent's model. Full suite green, only
TestLive* skipped.
Not addressed: §3's publish-time cycle detection, which needs the whole agent
set in hand. The depth cap is what holds without it, and is the half that
matters at run time.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
164 lines
6.4 KiB
Go
164 lines
6.4 KiB
Go
package runtime
|
|
|
|
import (
|
|
"github.com/krow/krow-backend/go-api/internal/config"
|
|
"github.com/krow/krow-backend/go-api/internal/gateway"
|
|
"github.com/krow/krow-backend/go-api/internal/knowledge"
|
|
"github.com/krow/krow-backend/go-api/internal/repo"
|
|
"github.com/krow/krow-backend/go-api/internal/tools"
|
|
)
|
|
|
|
// NewModelEngine builds the production runtime: the loader, the agent loop, a
|
|
// live model gateway and a Postgres trajectory sink.
|
|
//
|
|
// One call, because the alternative is four, and four assembled at a call site
|
|
// is how a deployment ends up running with a DiscardSink nobody chose. A test
|
|
// that wants a fake model still reaches for NewEngine with WithAgentExecutor —
|
|
// this function is the wiring, not a second way to configure the runtime.
|
|
//
|
|
// Skills keep the refusing stub. A skill has no executor of its own: the loop
|
|
// runs agents, and a skill reaches a model only as a capability an agent
|
|
// carries. Handing SkillExec a model would create a second, unbounded path to
|
|
// one — which is exactly the shape I3 exists to prevent.
|
|
func NewModelEngine(db repo.Querier, cfg config.Config) *Engine {
|
|
gw := gateway.NewAnthropic(gateway.FromConfig(cfg.Model))
|
|
retriever := knowledge.NewRetriever(db, NewEmbedder(cfg))
|
|
exec := NewModelExecutor(gw, NewPostgresSink(db), DefaultTools(db, retriever)).
|
|
WithRetriever(retriever).
|
|
// Without this, a spec's `subagents:` parses, loads, and is then
|
|
// dropped — which is how krow-workforce-agent came to declare five
|
|
// subagents and answer every question by itself. The resolver is the
|
|
// same Loader the engine uses, so a subagent is loaded exactly the way
|
|
// a directly-invoked agent is: same skills, same dependency checks.
|
|
WithSubagents(NewLoader(db))
|
|
return NewEngine(db, WithAgentExecutor(exec))
|
|
}
|
|
|
|
// NewEmbedder picks the embedding provider from configuration.
|
|
//
|
|
// Explicit first, then what is configured, then nothing. The order is the whole
|
|
// design: three providers all return vectors and retrieval works with any of
|
|
// them, so a deployment running the wrong one looks identical to one running
|
|
// the right one until somebody phrases a question differently. Naming the
|
|
// provider is how that stops being a silent condition.
|
|
//
|
|
// ollama A model on this machine. Real semantics, no credential, no
|
|
// per-token cost, no tenant text leaving the host. The default
|
|
// worth reaching for.
|
|
// voyage Hosted. Better on subtle retrieval over a large messy corpus,
|
|
// and the only one that needs a credential.
|
|
// lexical The deterministic stand-in. NOT semantic — it matches shared
|
|
// vocabulary and nothing else. Development only; config.validate
|
|
// refuses it in production.
|
|
//
|
|
// Returns nil when nothing is configured, and retrieval then runs keyword-only,
|
|
// saying so on every result. Nil rather than a hosted client with an empty key:
|
|
// both end up keyword-only, but nil says "no embedder is configured" once, at
|
|
// wiring time, instead of failing an HTTP call per query to learn the same
|
|
// thing.
|
|
func NewEmbedder(cfg config.Config) knowledge.Embedder {
|
|
k := cfg.Knowledge
|
|
|
|
provider := k.EmbedProvider
|
|
if provider == "" {
|
|
// Nothing named. Infer from what is actually present, preferring the
|
|
// one that costs nothing and keeps text local.
|
|
switch {
|
|
case k.UseLexicalEmbedder:
|
|
provider = "lexical"
|
|
case k.EmbedBaseURL != "":
|
|
provider = "ollama"
|
|
case k.EmbedAPIKey != "":
|
|
provider = "voyage"
|
|
default:
|
|
return nil
|
|
}
|
|
}
|
|
|
|
switch provider {
|
|
case "ollama":
|
|
return knowledge.NewOllama(k.EmbedBaseURL, k.EmbedModel, k.EmbedDims)
|
|
|
|
case "voyage":
|
|
if k.EmbedAPIKey == "" {
|
|
// Named but unusable. Nil, so retrieval degrades honestly rather
|
|
// than failing a request per query on a credential nobody set.
|
|
return nil
|
|
}
|
|
model, dims := k.EmbedModel, k.EmbedDims
|
|
if model == "" {
|
|
model = knowledge.DefaultVoyageModel
|
|
}
|
|
if dims == 0 {
|
|
dims = knowledge.DefaultVoyageDims
|
|
}
|
|
return knowledge.NewVoyage(k.EmbedAPIKey, model, dims)
|
|
|
|
case "lexical":
|
|
dims := k.EmbedDims
|
|
if dims == 0 {
|
|
dims = 256
|
|
}
|
|
e := knowledge.NewLexical(dims)
|
|
// Told what environment it is in, so its own refusal is the backstop
|
|
// behind config.validate's.
|
|
e.Production = cfg.AppEnv == "production"
|
|
return e
|
|
}
|
|
return nil
|
|
}
|
|
|
|
// DefaultTools is the tool registry this service ships with.
|
|
//
|
|
// One function, so "which tools exist" has a single answer that a test and the
|
|
// server reach the same way. Registration panics on a malformed tool: a
|
|
// service that booted without a capability its specs name would fail one run
|
|
// at a time instead of once, loudly, at startup.
|
|
func DefaultTools(db repo.Querier, retriever *knowledge.Retriever) *tools.Registry {
|
|
// The confirmation store is Postgres-backed, not in-process. A pending
|
|
// write is asked about in one request and approved in another, and nothing
|
|
// guarantees those two reach the same replica — an in-memory store would
|
|
// refuse a large share of perfectly good approvals, for a reason invisible
|
|
// to the person clicking. See tools.MemoryStore's own warning.
|
|
reg := tools.NewRegistryWithStore(tools.NewPostgresStore(db))
|
|
for _, t := range []tools.Tool{
|
|
// Activity
|
|
tools.ActivityBreakdown(db),
|
|
tools.ActivitySignals(db),
|
|
// Workforce
|
|
tools.WorkforceAttendance(db),
|
|
tools.WorkforceOvertime(db),
|
|
tools.WorkforceCoverage(db),
|
|
tools.WorkforceTraining(db),
|
|
// Hiring
|
|
tools.CandidatesQuality(db),
|
|
tools.HiresRecent(db),
|
|
tools.HiresPerformance(db),
|
|
tools.PositionsRisk(db),
|
|
tools.TalentPool(db),
|
|
// Cross-domain
|
|
tools.WorkspaceSummary(db),
|
|
tools.OperationsRisk(db),
|
|
// Assignments: the two lookups that yield ids, and the one write that
|
|
// consumes them. assign_worker is the only tool here with an effect,
|
|
// and it cannot run without an approval — see tools/confirm.go.
|
|
tools.OpenPositions(db),
|
|
tools.AvailableWorkers(db),
|
|
tools.AssignWorker(db),
|
|
// The hiring funnel: the lookup that yields application ids, and the
|
|
// write that moves somebody through it. Replaces the browser panel's
|
|
// interview matcher, which was the one capability the old templates had
|
|
// that the tool layer did not.
|
|
tools.CandidatesAwaiting(db),
|
|
tools.MoveApplication(db),
|
|
// Knowledge. Registered once; which corpora it may read comes from the
|
|
// running agent's spec by way of the tool Context, so this single
|
|
// registration serves every agent without any of them being able to
|
|
// name another's documents.
|
|
tools.KnowledgeSearch(retriever),
|
|
} {
|
|
reg.MustRegister(t)
|
|
}
|
|
return reg
|
|
}
|