Nearle Buddy answers a typed question

Phase 2: the loop and the model gateway. The composer in the console has
said "Not connected yet" since it was built, because there was no
assistant endpoint anywhere. There is one now.

- utils/chat.go   the gateway, a sibling of embedding.go: one small
                  interface, a provider switch, the shared postJSON, no
                  framework. Agents name a TIER (fast/balanced/deep) and
                  config maps tier to model, so changing provider does not
                  touch an agent.
- services/assistantService.go  one loop for every agent. An agent is a
                  name, a tier, a prompt and an allow-list — data, not a
                  class — so a sixth is config rather than a subclass.
- the endpoint under /v1/web, inheriting middleware.WebAuth along with
  every other console route. The assistant reads the same data the console
  does and must read it as the same person.

What the model does not get to decide:

  whose data      the caller is built from the verified session in the
                  controller, never from the request body — there is no
                  tenant field to fill in. A test scripts the model calling
                  a tool with {"tenantid": 916} and asserts it ran for 1147.
  which tools     the registry enforces the agent's allow-list; a test
                  scripts a call to a tool the agent lacks and asserts the
                  handler never ran.
  when to stop    steps and tool calls are counted here. A model that keeps
                  calling tools is stopped by arithmetic, not by being
                  asked nicely.

Two quiet failures have tests of their own. A finish_reason of "length"
means the provider cut the reply off mid-sentence, which reads exactly
like a complete answer unless it is flagged. And a truncated tool result
reaches the model in words it will repeat — otherwise it describes a
capped list and an empty one identically.

A refused tool goes back as a message, not an error: a model told "that
tool needs a tenant" can explain it, where a model handed nothing says
"something went wrong".

Optional, like the embedder. Without ASSISTANT_PROVIDER the endpoint
answers "not switched on here", the composer stays disabled, and the tools
still work — they are ordinary Go functions, and only turning a sentence
into a tool call needs a model.

14 tests, against a scripted model rather than a live provider: these are
about what the loop refuses to let a model do, and that has to hold for
any model, including one behaving badly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-09-23 13:13:20 +05:30
parent bb5f40926f
commit 8e1549764b
9 changed files with 1193 additions and 2 deletions

View File

@@ -71,6 +71,9 @@ type Config struct {
S3 S3Config
MQTT MQTTConfig
Embedding EmbeddingConfig
// Assistant is the model behind Nearle Buddy. Empty provider = no typed
// questions; the tools still work.
Assistant AssistantConfig
// POSTokenSecret signs terminal sessions. Falls back to JWTSecret when
// unset, matching utils/postoken.go.
@@ -141,6 +144,54 @@ type EmbeddingConfig struct {
func (e EmbeddingConfig) Enabled() bool { return e.Provider != "" }
// AssistantConfig is the model behind Nearle Buddy.
//
// Optional, like the embedder. With no provider the assistant refuses typed
// questions and says so — the tools still work and still answer correctly,
// because they are ordinary Go functions; only the part that turns a sentence
// into a tool call is missing.
//
// ── Why three models and not one ────────────────────────────────────────────
//
// An agent names a TIER, never a model. "Which branch is underperforming?"
// and "why is the cancel rate high?" want different amounts of thinking, and
// wiring a model name into an agent means changing every agent to change
// provider. The tiers are the stable vocabulary; this map is the only place a
// model name appears.
//
// `ASSISTANT_MODEL` alone sets all three, which is the sane default for a
// deployment that has not thought about it yet.
type AssistantConfig struct {
Provider string // "openai" — any OpenAI-compatible endpoint
BaseURL string // default https://api.openai.com/v1
APIKey string
// Tier → model name. Empty falls back to Balanced, which falls back to
// ASSISTANT_MODEL.
Fast string
Balanced string
Deep string
}
func (a AssistantConfig) Enabled() bool { return a.Provider != "" && a.Balanced != "" }
// ModelFor resolves a tier to a model name, falling back rather than failing.
//
// A missing `fast` model should answer a cheap question with the balanced one,
// not refuse it. A deployment that sets one model gets one model everywhere.
func (a AssistantConfig) ModelFor(tier string) string {
switch tier {
case "fast":
if a.Fast != "" {
return a.Fast
}
case "deep":
if a.Deep != "" {
return a.Deep
}
}
return a.Balanced
}
// IsProduction is true under APP_ENV=production.
func (c *Config) IsProduction() bool { return c.AppEnv == EnvProduction }
@@ -197,6 +248,17 @@ func Load() (*Config, error) {
BaseURL: env("EMBEDDING_BASE_URL", ""),
},
Assistant: AssistantConfig{
Provider: strings.ToLower(env("ASSISTANT_PROVIDER", "")),
BaseURL: env("ASSISTANT_BASE_URL", ""),
APIKey: env("ASSISTANT_API_KEY", ""),
Fast: env("ASSISTANT_MODEL_FAST", ""),
// ASSISTANT_MODEL alone sets every tier, for a deployment that has
// not thought about tiers yet.
Balanced: env("ASSISTANT_MODEL_BALANCED", env("ASSISTANT_MODEL", "")),
Deep: env("ASSISTANT_MODEL_DEEP", ""),
},
POSTokenSecret: env("POS_TOKEN_SECRET", ""),
JWTSecret: env("JWT_SECRET_KEY", ""),
UserContextKey: env("USER_CONTEXT_KEY", "nearle"),