Phase 1 of Nearle Buddy: an agent names a tool, and the registry decides whether that is allowed, whether the arguments make sense, who is asking, and what gets recorded — then runs a handler a person wrote and tested. No agent gets raw table access. The usual argument for tools over generated SQL is safety; here there is a harder one. The fields on this backend do not mean what their names say, and it is measured: orders.deliverystatus is an empty string on all 181 rows of tenant 1147, orders.orderstatus never carries the six middle delivery stages, deliveries.ridername holds statuses as often as names, deliverytype is empty on every row in production. A model writing SQL gets each of those wrong with no error — it reports a cancel rate from a column of empty strings and nobody can tell. A model calling a tool cannot, because the correction lives in the handler beside the measurement that justified it. Call does five things in order: find the tool, check the agent's allow-list, validate arguments, confirm the caller is scoped to something, run the handler — writing exactly one audit row whatever happens, refusals included. A trail of successes answers "did anything try to read another tenant?" with silence, which reads the same as no. The model has no say in whose data is read. stuck_orders has no tenantid field on its schema — absent, not rejected — and the tenant comes from the session claims added in the previous commit. Arguments the tool did not declare are dropped rather than passed on, so a model sending a `where` clause gets it discarded. stuck_orders: deliveries a rider was given and has not accepted, ten minutes for a look, twenty-five for somebody now. Derived from assigntime and orderstatus, so it does not depend on anyone having been watching. Carries the wait in minutes, what to do, where to check it, and what it covered. A capped answer says so — an empty result and a truncated one look identical to a model and it will call both "none". The audit sink writes to the log for now; a database sink is phase 8. Nothing calls the registry yet: the loop and the model gateway are phase 2. 37 tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
163 lines
4.8 KiB
Go
163 lines
4.8 KiB
Go
package tools
|
|
|
|
import (
|
|
"context"
|
|
"encoding/json"
|
|
"log"
|
|
"sort"
|
|
"strings"
|
|
"time"
|
|
)
|
|
|
|
// The audit trail.
|
|
//
|
|
// One row per call, including every refusal — the refusals are the interesting
|
|
// ones. A registry that recorded only successes would answer "did anything try
|
|
// to read another tenant?" with silence, which reads the same as "no".
|
|
//
|
|
// ── Why the sink is an interface ────────────────────────────────────────────
|
|
//
|
|
// Phase 1 writes to the log, because a table is a migration and this needs to
|
|
// work before that lands. Nothing else in the package knows that: the registry
|
|
// holds an `AuditSink`, so the database sink arrives later as a second
|
|
// implementation and no call site changes.
|
|
//
|
|
// ── Writes record intent, not outcome ───────────────────────────────────────
|
|
//
|
|
// Reads record what happened, which is all a read can be asked for. When write
|
|
// tools arrive they must record the ATTEMPT before the call leaves, not the
|
|
// result after it returns: a crash mid-write has to leave a trace that it was
|
|
// tried, and a row written only on success is a row that is missing exactly
|
|
// when it is needed.
|
|
|
|
const (
|
|
OutcomeOK = "ok"
|
|
OutcomeRefused = "refused"
|
|
OutcomeFailed = "failed"
|
|
)
|
|
|
|
// AuditEntry is one attempt to use a tool.
|
|
type AuditEntry struct {
|
|
At time.Time
|
|
Agent string
|
|
Tool string
|
|
Scope string
|
|
Userid int
|
|
Tenantid int
|
|
// The arguments as the handler received them — validated and defaulted, not
|
|
// as the model sent them. What actually ran is what is worth keeping.
|
|
Args map[string]any
|
|
// ok | refused | failed. `refused` is the guard saying no; `failed` is the
|
|
// handler breaking. Collapsing the two would hide a broken tool inside a
|
|
// count of things working as designed.
|
|
Outcome string
|
|
Detail string
|
|
Rows int
|
|
Took time.Duration
|
|
}
|
|
|
|
// AuditSink is where entries go.
|
|
//
|
|
// No error returned, deliberately. An audit sink that can fail a call gives a
|
|
// full disk the power to take the assistant down; one that cannot means a lost
|
|
// row, which is worse in theory and better in practice. A sink that cares
|
|
// should retry or buffer internally.
|
|
type AuditSink interface {
|
|
Write(ctx context.Context, entry AuditEntry)
|
|
}
|
|
|
|
// DiscardAudit keeps nothing. For tests that are not about the audit trail.
|
|
type DiscardAudit struct{}
|
|
|
|
func (DiscardAudit) Write(context.Context, AuditEntry) {}
|
|
|
|
// LogAudit writes one line per call to the standard logger.
|
|
//
|
|
// A line rather than JSON per field, because this is read by a person tailing
|
|
// logs during the rollout. The database sink can be structured.
|
|
type LogAudit struct{}
|
|
|
|
func (LogAudit) Write(_ context.Context, entry AuditEntry) {
|
|
log.Printf("assistant: %s", entry.Line())
|
|
}
|
|
|
|
// Line renders an entry for a log.
|
|
//
|
|
// Arguments are rendered sorted so two identical calls produce identical lines
|
|
// and a grep for one of them finds both. Go's map iteration is randomised, so
|
|
// without the sort the same call logs differently every time.
|
|
func (e AuditEntry) Line() string {
|
|
var b strings.Builder
|
|
b.WriteString(e.Outcome)
|
|
b.WriteString(" ")
|
|
b.WriteString(e.Agent)
|
|
b.WriteString("/")
|
|
b.WriteString(e.Tool)
|
|
|
|
if e.Tenantid > 0 {
|
|
b.WriteString(" tenant=")
|
|
b.WriteString(itoa(e.Tenantid))
|
|
} else {
|
|
// Explicitly, rather than by omission: "no tenant" on an assistant call
|
|
// is either staff or a bug, and both are worth being able to search for.
|
|
b.WriteString(" tenant=none")
|
|
}
|
|
b.WriteString(" user=")
|
|
b.WriteString(itoa(e.Userid))
|
|
|
|
if len(e.Args) > 0 {
|
|
keys := make([]string, 0, len(e.Args))
|
|
for key := range e.Args {
|
|
keys = append(keys, key)
|
|
}
|
|
sort.Strings(keys)
|
|
parts := make([]string, 0, len(keys))
|
|
for _, key := range keys {
|
|
value, err := json.Marshal(e.Args[key])
|
|
if err != nil {
|
|
value = []byte("?")
|
|
}
|
|
parts = append(parts, key+"="+string(value))
|
|
}
|
|
b.WriteString(" args{")
|
|
b.WriteString(strings.Join(parts, " "))
|
|
b.WriteString("}")
|
|
}
|
|
|
|
if e.Outcome == OutcomeOK {
|
|
b.WriteString(" rows=")
|
|
b.WriteString(itoa(e.Rows))
|
|
}
|
|
if e.Detail != "" {
|
|
b.WriteString(" detail=")
|
|
value, err := json.Marshal(e.Detail)
|
|
if err != nil {
|
|
b.WriteString("?")
|
|
} else {
|
|
b.Write(value)
|
|
}
|
|
}
|
|
b.WriteString(" took=")
|
|
b.WriteString(e.Took.Round(time.Millisecond).String())
|
|
return b.String()
|
|
}
|
|
|
|
func itoa(n int) string {
|
|
value, _ := json.Marshal(n)
|
|
return string(value)
|
|
}
|
|
|
|
// CollectAudit keeps entries in memory, for tests that ARE about the trail.
|
|
type CollectAudit struct{ Entries []AuditEntry }
|
|
|
|
func (c *CollectAudit) Write(_ context.Context, entry AuditEntry) {
|
|
c.Entries = append(c.Entries, entry)
|
|
}
|
|
|
|
func (c *CollectAudit) Last() (AuditEntry, bool) {
|
|
if len(c.Entries) == 0 {
|
|
return AuditEntry{}, false
|
|
}
|
|
return c.Entries[len(c.Entries)-1], true
|
|
}
|