Nearle Buddy answers a typed question
Phase 2: the loop and the model gateway. The composer in the console has
said "Not connected yet" since it was built, because there was no
assistant endpoint anywhere. There is one now.
- utils/chat.go the gateway, a sibling of embedding.go: one small
interface, a provider switch, the shared postJSON, no
framework. Agents name a TIER (fast/balanced/deep) and
config maps tier to model, so changing provider does not
touch an agent.
- services/assistantService.go one loop for every agent. An agent is a
name, a tier, a prompt and an allow-list — data, not a
class — so a sixth is config rather than a subclass.
- the endpoint under /v1/web, inheriting middleware.WebAuth along with
every other console route. The assistant reads the same data the console
does and must read it as the same person.
What the model does not get to decide:
whose data the caller is built from the verified session in the
controller, never from the request body — there is no
tenant field to fill in. A test scripts the model calling
a tool with {"tenantid": 916} and asserts it ran for 1147.
which tools the registry enforces the agent's allow-list; a test
scripts a call to a tool the agent lacks and asserts the
handler never ran.
when to stop steps and tool calls are counted here. A model that keeps
calling tools is stopped by arithmetic, not by being
asked nicely.
Two quiet failures have tests of their own. A finish_reason of "length"
means the provider cut the reply off mid-sentence, which reads exactly
like a complete answer unless it is flagged. And a truncated tool result
reaches the model in words it will repeat — otherwise it describes a
capped list and an empty one identically.
A refused tool goes back as a message, not an error: a model told "that
tool needs a tenant" can explain it, where a model handed nothing says
"something went wrong".
Optional, like the embedder. Without ASSISTANT_PROVIDER the endpoint
answers "not switched on here", the composer stays disabled, and the tools
still work — they are ordinary Go functions, and only turning a sentence
into a tool call needs a model.
14 tests, against a scripted model rather than a live provider: these are
about what the loop refuses to let a model do, and that has to hold for
any model, including one behaving badly.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -71,6 +71,9 @@ type Config struct {
|
||||
S3 S3Config
|
||||
MQTT MQTTConfig
|
||||
Embedding EmbeddingConfig
|
||||
// Assistant is the model behind Nearle Buddy. Empty provider = no typed
|
||||
// questions; the tools still work.
|
||||
Assistant AssistantConfig
|
||||
|
||||
// POSTokenSecret signs terminal sessions. Falls back to JWTSecret when
|
||||
// unset, matching utils/postoken.go.
|
||||
@@ -141,6 +144,54 @@ type EmbeddingConfig struct {
|
||||
|
||||
func (e EmbeddingConfig) Enabled() bool { return e.Provider != "" }
|
||||
|
||||
// AssistantConfig is the model behind Nearle Buddy.
|
||||
//
|
||||
// Optional, like the embedder. With no provider the assistant refuses typed
|
||||
// questions and says so — the tools still work and still answer correctly,
|
||||
// because they are ordinary Go functions; only the part that turns a sentence
|
||||
// into a tool call is missing.
|
||||
//
|
||||
// ── Why three models and not one ────────────────────────────────────────────
|
||||
//
|
||||
// An agent names a TIER, never a model. "Which branch is underperforming?"
|
||||
// and "why is the cancel rate high?" want different amounts of thinking, and
|
||||
// wiring a model name into an agent means changing every agent to change
|
||||
// provider. The tiers are the stable vocabulary; this map is the only place a
|
||||
// model name appears.
|
||||
//
|
||||
// `ASSISTANT_MODEL` alone sets all three, which is the sane default for a
|
||||
// deployment that has not thought about it yet.
|
||||
type AssistantConfig struct {
|
||||
Provider string // "openai" — any OpenAI-compatible endpoint
|
||||
BaseURL string // default https://api.openai.com/v1
|
||||
APIKey string
|
||||
// Tier → model name. Empty falls back to Balanced, which falls back to
|
||||
// ASSISTANT_MODEL.
|
||||
Fast string
|
||||
Balanced string
|
||||
Deep string
|
||||
}
|
||||
|
||||
func (a AssistantConfig) Enabled() bool { return a.Provider != "" && a.Balanced != "" }
|
||||
|
||||
// ModelFor resolves a tier to a model name, falling back rather than failing.
|
||||
//
|
||||
// A missing `fast` model should answer a cheap question with the balanced one,
|
||||
// not refuse it. A deployment that sets one model gets one model everywhere.
|
||||
func (a AssistantConfig) ModelFor(tier string) string {
|
||||
switch tier {
|
||||
case "fast":
|
||||
if a.Fast != "" {
|
||||
return a.Fast
|
||||
}
|
||||
case "deep":
|
||||
if a.Deep != "" {
|
||||
return a.Deep
|
||||
}
|
||||
}
|
||||
return a.Balanced
|
||||
}
|
||||
|
||||
// IsProduction is true under APP_ENV=production.
|
||||
func (c *Config) IsProduction() bool { return c.AppEnv == EnvProduction }
|
||||
|
||||
@@ -197,6 +248,17 @@ func Load() (*Config, error) {
|
||||
BaseURL: env("EMBEDDING_BASE_URL", ""),
|
||||
},
|
||||
|
||||
Assistant: AssistantConfig{
|
||||
Provider: strings.ToLower(env("ASSISTANT_PROVIDER", "")),
|
||||
BaseURL: env("ASSISTANT_BASE_URL", ""),
|
||||
APIKey: env("ASSISTANT_API_KEY", ""),
|
||||
Fast: env("ASSISTANT_MODEL_FAST", ""),
|
||||
// ASSISTANT_MODEL alone sets every tier, for a deployment that has
|
||||
// not thought about tiers yet.
|
||||
Balanced: env("ASSISTANT_MODEL_BALANCED", env("ASSISTANT_MODEL", "")),
|
||||
Deep: env("ASSISTANT_MODEL_DEEP", ""),
|
||||
},
|
||||
|
||||
POSTokenSecret: env("POS_TOKEN_SECRET", ""),
|
||||
JWTSecret: env("JWT_SECRET_KEY", ""),
|
||||
UserContextKey: env("USER_CONTEXT_KEY", "nearle"),
|
||||
|
||||
Reference in New Issue
Block a user