Phase 1 of Nearle Buddy: an agent names a tool, and the registry decides whether that is allowed, whether the arguments make sense, who is asking, and what gets recorded — then runs a handler a person wrote and tested. No agent gets raw table access. The usual argument for tools over generated SQL is safety; here there is a harder one. The fields on this backend do not mean what their names say, and it is measured: orders.deliverystatus is an empty string on all 181 rows of tenant 1147, orders.orderstatus never carries the six middle delivery stages, deliveries.ridername holds statuses as often as names, deliverytype is empty on every row in production. A model writing SQL gets each of those wrong with no error — it reports a cancel rate from a column of empty strings and nobody can tell. A model calling a tool cannot, because the correction lives in the handler beside the measurement that justified it. Call does five things in order: find the tool, check the agent's allow-list, validate arguments, confirm the caller is scoped to something, run the handler — writing exactly one audit row whatever happens, refusals included. A trail of successes answers "did anything try to read another tenant?" with silence, which reads the same as no. The model has no say in whose data is read. stuck_orders has no tenantid field on its schema — absent, not rejected — and the tenant comes from the session claims added in the previous commit. Arguments the tool did not declare are dropped rather than passed on, so a model sending a `where` clause gets it discarded. stuck_orders: deliveries a rider was given and has not accepted, ten minutes for a look, twenty-five for somebody now. Derived from assigntime and orderstatus, so it does not depend on anyone having been watching. Carries the wait in minutes, what to do, where to check it, and what it covered. A capped answer says so — an empty result and a truncated one look identical to a model and it will call both "none". The audit sink writes to the log for now; a database sink is phase 8. Nothing calls the registry yet: the loop and the model gateway are phase 2. 37 tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
4.8 KiB
4.8 KiB