725 lines
42 KiB
Markdown
725 lines
42 KiB
Markdown
# Doormile Agent Platform — Plan & Agent Registry
|
||
|
||
Status: **Phases 0–4 done (uncommitted, not deployed), Phase 5 next** · Written 2026-09-29 · Scope: `krow_talent_app` (the
|
||
Doormile console), `doormile_backend` (Go API), `AI_engine` (Python agent swarm).
|
||
|
||
This is the document `src/pages/doormile/settings/Settings.jsx` has been pointing
|
||
at ("Phase 6 of docs/agent-platform-plan.md"). That file never existed until now,
|
||
so the phase numbers below replace the ones the comment assumed.
|
||
|
||
---
|
||
|
||
## 1. Where things actually stand (verified 2026-09-29)
|
||
|
||
### Console — `krow_talent_app`
|
||
- **Settings → Skills & Tools (Agent Studio) is a mock.** `agentRegistryData.js`
|
||
holds 2 agents, 3 surfaces, 4 skills and 6 tools. Everything is saved to
|
||
`localStorage` and nowhere else. The Test tab (`AgentPlayground.jsx:37`) is a
|
||
`setTimeout` that always replies "succeeded with 100% confidence". Configure and
|
||
Insights use hard-coded figures ("99.4%", "540 calls") and a model
|
||
`claude-3-5-sonnet` that nothing reads.
|
||
- **Three of the six mock tools are Krow leftovers:** `open_add_skill_training`,
|
||
`open_add_training` and `query_learning_analytics` (workforce training, not
|
||
logistics). Delete them; don't migrate them.
|
||
- **The Agents page** (`/doormile/agents`) runs on a static snapshot in
|
||
`src/lib/agentNetwork.js`, dated 16–20 Sep 2026. It is honest about being a
|
||
snapshot, but its test (`tests/integration/agentsPage.test.jsx`) is stale: that
|
||
is the 1 failing suite and 9 failing tests out of 45 suites.
|
||
- **Unmerged branch `feat/agentic-ops-layer`** (dharaneesh, 2026-09-01) is the only
|
||
real client-side agent work:
|
||
- `SkillRegistry` with 8 rule-based skills: SlaGuardian, DoorstepStall,
|
||
FleetBalancer, HighValueCod, RiderBatterySafety, HubCongestion, LateDispatch,
|
||
CashExposure.
|
||
- `tools.js` with 5 tools, following the rule "read-only tools execute, mutating
|
||
tools return a Proposal".
|
||
- AgentFactory, the ops briefing, proposal executors with a verify pass, and 20+
|
||
test files.
|
||
- It is **29 commits behind main and conflicts in 10 files**. Its skill config is
|
||
also stored in `localStorage`.
|
||
- **The Home assistant** runs on the regex catalogue `src/lib/assistant/intents.js`
|
||
(109 KB) and has no concept of skills.
|
||
- **Other state:** lint shows 35 auto-fixable errors, and `Deliveries.jsx` has an
|
||
uncommitted cosmetic change (a tidy-up of the batch filter).
|
||
|
||
### Backend — `doormile_backend`
|
||
- Build, vet and test all pass (Go 1.26.4). There are 233 routes; CLAUDE.md is
|
||
stale on counts, retry, routing, PIN auth and SMS (see its drift list).
|
||
- **Nothing for an agent registry exists yet.** `internal/ai/` is an empty,
|
||
untracked directory.
|
||
- **What does exist:**
|
||
- `models.AgentDecision` (table `agent_decisions`, pgvector 1536-dim — the
|
||
384-vs-1536 question is still open).
|
||
- Internal routes, all behind `X-Internal-Key`:
|
||
- `POST /internal/agent-decisions`
|
||
- `GET /internal/agent-decisions/similar`
|
||
- `PATCH /internal/agent-decisions/:id/outcome`
|
||
- `/internal/express/{riders,bookings,assign}`
|
||
- **Admin authorization is flat.** Roles 1, 3 and 4 get identical access, and no
|
||
handler checks the role further. Many admin write handlers have no tenant guard.
|
||
So any registry write endpoint must add its own role check (see §5).
|
||
|
||
### Engine — `AI_engine`
|
||
- **It has no registry to read from.** Agents are hard-coded classes in
|
||
`main.py:55-70`. `core/tool_registry.py` is an in-memory dict holding 3 tools,
|
||
and none of them load in production. The two "skills" in `core/skills/` are
|
||
stubs that nothing imports.
|
||
- **The agent process has no HTTP API.** The Command Center (FastAPI, :8600) is a
|
||
separate, unauthenticated NATS tap.
|
||
- **Runs, escalations and decisions live in memory only.** They are capped at 500
|
||
and lost on restart.
|
||
- **Tests:** 72 pass, but only with pytest installed by hand; `requirements.txt`
|
||
lacks pytest.
|
||
|
||
**Conclusion:** the registry must be **built**, not wired up. It belongs in
|
||
`doormile_backend` (Postgres). The console reads it, and the engine reads it.
|
||
Neither of them owns it.
|
||
|
||
---
|
||
|
||
## 2. The Agent Registry — what gets seeded
|
||
|
||
This is the real inventory. Every row here becomes a seed row in Phase 1.
|
||
"Status" is what the code does today, not what the docs claim.
|
||
|
||
### 2.1 Agents
|
||
|
||
| id | Class / file | Purpose | Trigger | Writes to system | LLM | Status |
|
||
|---|---|---|---|---|---|---|
|
||
| `JARVIS` | `MasterAgent` · `core/agent.py:196` | Orchestrator, escalation inbox | `logistics.direct.JARVIS` | none | — | Partial. The inbox works; `orchestrate_order` sends the wrong payload key |
|
||
| `DISPATCH_AGENT` | `agents/dispatch_agent.py:72` | Watches assignment outcomes, flags coverage gaps | JetStream `booking.assigned`, `booking.assignment_failed` | Alerts; customer notify only when autonomous | `decide_assignment_failure` | **Production-grade.** Idle until `assignment_failed` is confirmed as published (see §7) |
|
||
| `EXCEPTION_AGENT` | `agents/exception_agent.py:106` | Stalled-rider detection and response | TRACKING `miler.location.updated`, `miler.stalled`, plus a 60 s DB sweep | `POST /internal/bookings/:id/reassign` (needs autonomy flag and confidence ≥ 0.7), `/internal/notify` | `decide_stall_response` | **Production-grade** (the stall path only) |
|
||
| `EXPRESS_DISPATCH_AGENT` | `agents/express_dispatch_agent.py:80` | Tenant batch assign plus road sequencing | JetStream `express.dispatch_requested` | `POST /internal/express/assign` | — (greedy) | Implemented. Autonomy defaults to `true` in code and `false` in compose |
|
||
| `CUSTOMER_AGENT` | `agents/customer_agent.py:50` | Customer notifications | Direct tasks, only from an autonomous Dispatch | `/internal/notify` | — | Implemented, rarely reached |
|
||
| `ORDER_AGENT` | `agents/order_agent.py:15` | Order intake and validation | Direct tasks (no sender in production) | `/admin/crmbooking` — a **dead path, no auth header** | — | Broken. `update_status` crashes (`order_agent.py:182`) |
|
||
| `HUB_AGENT` | `agents/hub_agent.py:35` | Hub capacity | Direct tasks | none | — | **Simulation** (8 fictional hubs) |
|
||
| `FLEET_AGENT` | `agents/fleet_agent.py:34` | Vehicles | Direct tasks | none | — | **Simulation** (19 fake vehicles) |
|
||
| `ROUTE_OPTIMIZER` | `agents/route_optimizer_agent.py:41` | Routing | Direct tasks | none | — | **Simulation** (haversine only) |
|
||
| `CONSOLE_ASSISTANT` | `src/lib/assistant/*` (console) | Home chat: orders, bulk, assign, repeat | Operator prompt | Via proposals the operator confirms | RAG sidecar (off) | Live, regex-based |
|
||
| `CONSOLE_OPS_AGENT` | `feat/agentic-ops-layer` | Ops briefing plus 8 monitoring skills | Page load / poll | Via proposals the operator confirms | — (rules) | **Unmerged** |
|
||
|
||
The registry shows `status` as one of `live | partial | simulation | broken |
|
||
unmerged | retired`. **The console must display this badge.** A simulation agent
|
||
that looks live is the exact failure the Agents page comment warns about.
|
||
|
||
### 2.2 Tools (the real capabilities, not the stubs)
|
||
|
||
`kind` is `read`, `write` or `notify`. Every `write` or `notify` tool has
|
||
`requires_confirmation = true` unless its agent is explicitly autonomous.
|
||
|
||
| name | Owner agent(s) | Kind | Target | Today lives at |
|
||
|---|---|---|---|---|
|
||
| `reassign_booking` | EXCEPTION | write | Go `/internal/bookings/:id/reassign` | `exception_agent.py:466` |
|
||
| `notify_customer` | EXCEPTION, CUSTOMER | notify | Go `/internal/notify` | `exception_agent.py:476`, `customer_agent.py:349` |
|
||
| `list_express_bookings` | EXPRESS | read | Go `/internal/express/bookings` | `express_dispatch_agent.py:351` |
|
||
| `list_express_riders` | EXPRESS | read | Go `/internal/express/riders` | `express_dispatch_agent.py:342` |
|
||
| `assign_express_batch` | EXPRESS | write | Go `/internal/express/assign` | `express_dispatch_agent.py:222` |
|
||
| `sequence_stops` | EXPRESS | read (compute) | `routes.workolik.com /optimization/doormile/sequence` | `express_dispatch_agent.py:308` |
|
||
| `get_booking_cache` | CUSTOMER | read | Go `/bookings/cache/:id` | `customer_agent.py:243` |
|
||
| `nearby_milers` | DISPATCH, EXCEPTION | read | Redis GEO `milers:locations` | `dispatch_agent.py:170-240` |
|
||
| `publish_miler_stalled` | EXCEPTION | write (event) | NATS `miler.stalled` | `exception_agent.py:377` |
|
||
| `decide_stall_response` | EXCEPTION | read (LLM) | Claude | `core/llm.py:157` |
|
||
| `decide_assignment_failure` | DISPATCH | read (LLM) | Claude | `core/llm.py:227` |
|
||
| `record_agent_decision` | all | write | Go `/internal/agent-decisions` | Go route exists |
|
||
| `diagnose_operations` | CONSOLE_OPS | read | Console queries | branch `tools.js` |
|
||
| `lookup_order` | CONSOLE_OPS | read | `/admin/bookings` | branch `tools.js` |
|
||
| `lookup_miler` | CONSOLE_OPS | read | `/admin/milers` | branch `tools.js` |
|
||
| `propose_miler_reassignment` | CONSOLE_OPS | write → proposal | `/admin/bookings/:id/assign-miler` | branch `tools.js` |
|
||
| `simulate_pricing_quote` | CONSOLE_OPS, CONSOLE_ASSISTANT | read | `/admin/pricing/simulate` | branch `tools.js` |
|
||
| `create_single_order` | CONSOLE_ASSISTANT | write → proposal | `/admin/expressbooking` | current mock |
|
||
| `rebalance_riders` | CONSOLE_OPS | write → proposal | `/hub/bookings/batch-assign` | current mock (no backing code) |
|
||
|
||
**Not seeded:**
|
||
- `ask_question`, `order_intake_skill`, `repeat_run_skill` (engine stubs that are
|
||
never loaded).
|
||
- The 3 Krow training tools.
|
||
- ORDER_AGENT's `crmbooking` calls: that route was renamed to `expressbooking`, so
|
||
those calls hit nothing.
|
||
|
||
### 2.3 Skills
|
||
|
||
A skill is a named behaviour that belongs to one agent and uses one or more tools.
|
||
It has a switch (`enabled`) and tunable `thresholds` (JSON, validated against a
|
||
per-skill schema).
|
||
|
||
| id | Agent | Tools | Source |
|
||
|---|---|---|---|
|
||
| `stall_response` | EXCEPTION | nearby_milers, decide_stall_response, reassign_booking, notify_customer | engine |
|
||
| `assignment_failure_triage` | DISPATCH | nearby_milers, decide_assignment_failure | engine |
|
||
| `express_batch_dispatch` | EXPRESS | list_express_*, assign_express_batch, sequence_stops | engine |
|
||
| `sla_guardian` · `doorstep_stall` · `fleet_balancer` · `high_value_cod` · `rider_battery_safety` · `hub_congestion` · `late_dispatch` · `cash_exposure` | CONSOLE_OPS | branch `tools.js` set | branch (thresholds move from localStorage to the registry) |
|
||
| `order_intake_auto_schedule` | CONSOLE_ASSISTANT | create_single_order, simulate_pricing_quote | current mock, real behaviour in `orderFlow.js` |
|
||
| `dispatch_rebalance` | CONSOLE_OPS | rebalance_riders | current mock. **Keep disabled** until a backing endpoint exists |
|
||
|
||
### 2.4 Surfaces
|
||
`/doormile/home` (assistant), `/doormile/control-x` (dispatch board),
|
||
`/doormile/agents` (status board), and the ops banner (branch
|
||
`AgentOperationsBanner`). `/doormile/dispatch` is only a redirect to Control X, so
|
||
it is not a separate surface.
|
||
|
||
---
|
||
|
||
## 3. Schema (Phase 1) — ⚠ additive schema change, review before deploy
|
||
|
||
New tables in `doormile_backend`, added through GORM AutoMigrate. They are
|
||
additive only and touch no existing table.
|
||
|
||
```
|
||
ai_agents id text PK, name, class_ref, runtime ('engine'|'console'),
|
||
purpose, trigger jsonb, status, autonomous bool,
|
||
model text NULL, created_at, updated_at
|
||
ai_tools name text PK, description, kind ('read'|'write'|'notify'),
|
||
target text, input_schema jsonb, requires_confirmation bool,
|
||
enabled bool, created_at, updated_at
|
||
ai_skills id text PK, agent_id FK→ai_agents, title, category,
|
||
description, sample_prompt, enabled bool,
|
||
thresholds jsonb, thresholds_schema jsonb,
|
||
version int, updated_by int NULL, updated_at
|
||
ai_skill_tools skill_id FK, tool_name FK, PRIMARY KEY (skill_id, tool_name)
|
||
ai_registry_audit id bigserial, entity, entity_id, field, old jsonb, new jsonb,
|
||
changed_by int, changed_at
|
||
```
|
||
|
||
Rules:
|
||
- **Tools and agents are code-defined.** The console can toggle `enabled` and
|
||
`autonomous`, but it cannot invent a tool. A tool with no implementation is a lie
|
||
on screen. "New skill" in the console therefore picks from existing tools only.
|
||
- **Every write goes into `ai_registry_audit` in the same transaction.**
|
||
- **Seeding is idempotent** (upsert by id) and lives in `migrations/`, not
|
||
`scratch/`.
|
||
- **Never store a secret.** Env var *names* only; the engine's `.env` stays out of
|
||
the registry.
|
||
|
||
---
|
||
|
||
## 4. API contract (Phase 1)
|
||
|
||
All handlers go in `controllers/aiRegistryController.go` and the logic in
|
||
`internal/ai/registry`. Responses use `utils.OK` / `utils.List`, following the
|
||
console conventions.
|
||
|
||
| Method | Path | Auth | Notes |
|
||
|---|---|---|---|
|
||
| GET | `/admin/ai/agents` | admin (1,3,4), Doormile staff only | Includes a `status` badge and skill/tool counts |
|
||
| GET | `/admin/ai/agents/:id` | same | Agent plus its skills and tools |
|
||
| GET | `/admin/ai/skills` | same | `?agent=` filter |
|
||
| GET | `/admin/ai/tools` | same | `?kind=` filter |
|
||
| PATCH | `/admin/ai/skills/:id` | **roleid 1 only** | Allowed fields: `enabled`, `thresholds` (validated against schema). Bumps `version` and writes an audit row |
|
||
| POST | `/admin/ai/skills` | roleid 1 only | New skill from existing tools only |
|
||
| PATCH | `/admin/ai/agents/:id` | roleid 1 only | Allowed fields: `autonomous`, `model`. **Extra confirm in UI** — this changes what an agent does without a human |
|
||
| GET | `/admin/ai/audit` | admin | Registry change history |
|
||
| GET | `/admin/ai/decisions` | admin | Paged `agent_decisions` (Phase 4) |
|
||
| GET | `/internal/ai/registry` | `X-Internal-Key` | The engine pulls its config from here (Phase 5). It sends an ETag so the engine can poll cheaply |
|
||
|
||
"Doormile staff only" means `consoleTenantID == 0`. A tenant or client login gets
|
||
403. It does **not** get an empty list, because an empty list would hide a
|
||
misconfiguration.
|
||
|
||
---
|
||
|
||
## 5. Phases
|
||
|
||
Each phase ships on its own and leaves the system working. Nothing is committed or
|
||
pushed without an explicit ask.
|
||
|
||
### Phase 0 — Prerequisites and hygiene (small, do first)
|
||
1. **Rotate secrets and untrack them.**
|
||
- `doormile_backend`: `.env` and `doormile-abee7-*.json` are tracked; the
|
||
`.gitignore` line `#.env` is commented out.
|
||
- `AI_engine`: `.env` is tracked.
|
||
- `config/config.go:60-71` has hard-coded fallback secrets. Fail at boot
|
||
instead.
|
||
- This is a **user action** (rotation needs the providers' consoles). I can do
|
||
the untracking and ignore rules.
|
||
2. **Decide `feat/agentic-ops-layer`** (see §6, decision A).
|
||
3. **Console:** fix the stale `agentsPage.test.jsx`, run `lint:fix` for the 35
|
||
unused imports, and delete the 3 Krow training tools from
|
||
`agentRegistryData.js`.
|
||
4. **Add a PREVIEW banner** on the Agent Studio tab itself. Today only a code
|
||
comment says so; Agents shows a snapshot label, Agent Studio shows nothing.
|
||
5. **Backend:** remove the broken `.claude/skills/*` symlink stubs, `.agents/`,
|
||
`skills.md` and `skills-lock.json`. The global plugin already provides these
|
||
skills.
|
||
|
||
**Done when:** tests are green, lint is clean, and no secret is in `git ls-files`.
|
||
|
||
**Phase 0 result (2026-09-29, uncommitted):**
|
||
- Console: lint clean; **46/46 suites, 1192 tests pass**.
|
||
- Agent Studio: 4 Krow tools and 2 Krow skills removed (the plan said 3 tools;
|
||
`open_miletruth_ai` was a fourth). `dispatch_rebalance` ships disabled.
|
||
Storage keys moved to `_v3`. An on-screen Preview note was added.
|
||
- Agents page: this was not just a stale test. The 24–25 Sep rebuild presented
|
||
a simulation as live. Per decision, the design was kept and labelled:
|
||
- a `Snapshot · 16–20 Sep 2026` stamp;
|
||
- "Sample Activity — Simulation · not live data";
|
||
- no pulsing dot;
|
||
- status counts taken from `networkStats()`;
|
||
- "Autonomy gates on: 0 / 3", where the old "0 / 8" implied 8 gates.
|
||
The tests were rewritten, keeping the honesty checks.
|
||
- Backend:
|
||
- **Reverted 2026-09-29 at Suriya's request:** these are back in git
|
||
exactly as at HEAD:
|
||
- the untracking of `.env` and the service-account key (the key is
|
||
still committed, so rotation still stands);
|
||
- the `.gitignore` edit;
|
||
- the removal of `skills.md`, `skills-lock.json`, `.agents/` and
|
||
`.claude/skills/`.
|
||
AI_engine's `.env` is tracked again too. Do not redo any of this without
|
||
asking.
|
||
- `main.go` now refuses to boot when `ENV=production` and `JWT_SECRET_KEY`,
|
||
`DB_PASSWORD` or `NATS_PASSWORD` is unset.
|
||
- build, vet and test are green.
|
||
- AI_engine: `.env` was untracked (it was already in `.gitignore`).
|
||
- **Still yours:**
|
||
- Rotate every secret that was committed. Git history still holds the
|
||
values.
|
||
- **Before the next backend deploy,** confirm that production sets all
|
||
three secrets. If it has been running on a fallback, the new check stops
|
||
it from starting.
|
||
|
||
### Phase 1 — Registry in the backend
|
||
- Add the §3 tables, the idempotent seed from §2, and the §4 read endpoints plus
|
||
PATCH/POST with the role check and audit trail.
|
||
- **Tests:** a seed-idempotency test, a role test (roles 3 and 4 get 403 on PATCH;
|
||
a tenant login gets 403 on GET), a threshold-schema validation test, and an audit
|
||
test.
|
||
- **Done when:** `go build/vet/test` is green and `curl` against a staging DB
|
||
returns the §2 inventory.
|
||
|
||
**Phase 1 result (2026-09-29, uncommitted, NOT deployed, no real DB touched):**
|
||
- **Tables.** They follow the codebase's naming, not the names in §3:
|
||
`aiagents`, `aitools`, `aiskills`, `aiskilltools`, `airegistryaudit`.
|
||
The columns are as in §3, with a few changes:
|
||
- The agent's trigger column is named `wakeon`.
|
||
- `hasautonomygate` is new. Autonomy can only be set on Dispatch,
|
||
Exception and Express.
|
||
- `source` (engine/console/custom) is on skills.
|
||
- There is no `enabled` flag on tools.
|
||
- **Code.**
|
||
- `internal/ai/registry`: seed, rules, store.
|
||
- `controllers/aiRegistryController.go`
|
||
- `middlewares/staff_only.go` (`DoormileStaffOnly`)
|
||
- Routes are under `/admin/ai/*` and `/internal/ai/registry`, as in §4.
|
||
- The seed runs in `migrations.Migrate`. It logs a failure and does not
|
||
stop the boot.
|
||
- **Seed.** 11 agents, 26 tools and 15 skills.
|
||
- Four skills were added to §2.3 so that every tool belongs to a skill:
|
||
`customer_notifications`, `ops_briefing`, and the 8 branch tools
|
||
folded into their skills.
|
||
- The console ops skills keep the branch ids and threshold keys, so
|
||
Phase 3 is a 1:1 mapping.
|
||
- `record_agent_decision` was dropped. It is a log the Go side writes, not
|
||
a capability.
|
||
- **Rules enforced server-side.**
|
||
- Switching autonomy ON needs `confirm` set to the agent id.
|
||
- Model ids must match `claude-*`, and only engine agents have one.
|
||
- Thresholds are validated for range and step, and a patch is all-or-nothing.
|
||
- Custom skills can be added to console agents only, and only from existing
|
||
tools.
|
||
- A patch that changes nothing does not bump the version or write an audit
|
||
row.
|
||
- **Tests.**
|
||
- 26 unit tests and 6 HTTP gate tests always run.
|
||
- 10 Postgres integration tests and 1 HTTP end-to-end test run only when
|
||
`REGISTRY_TEST_DSN` is set. They need a throwaway database; each package
|
||
uses its own schema.
|
||
- All of them passed against a disposable `postgres:16-alpine` container.
|
||
- **Bug the Postgres run caught.** With `enabled` tagged `default:true`, gorm
|
||
dropped `false` from the INSERT, so `dispatch_rebalance` came up
|
||
**enabled**. Fixed by removing the column default.
|
||
- **Not yet proven.** The migration has not run against the real database;
|
||
that happens on your next deploy. It is 5 new tables and touches nothing
|
||
existing.
|
||
|
||
### Phase 2 — Console Agent Studio reads the registry
|
||
- Replace `agentRegistryData.js` and its `localStorage` with React Query hooks
|
||
(`useAiAgents`, `useAiSkills`, `useAiTools`) in `src/lib/doormileHooks.js`, and
|
||
add the endpoints to `src/api/doormile/endpoints.js`.
|
||
- Keep the existing components; only the data source changes. Also show the status
|
||
badge, the tool `kind`, and a confirmation marker.
|
||
- The skill toggle and "New skill" become real PATCH and POST calls. Add a
|
||
`loading`/`error` state; remove the optimistic toast that claims success before
|
||
the server answers.
|
||
- Configure tab: model and autonomy from the registry. The temperature slider is
|
||
**dropped**, because nothing reads it.
|
||
- Insights and Test stay behind the PREVIEW banner until Phases 4 and 6.
|
||
- Rewrite `tests/integration/agentStudio.test.jsx` against mocked hooks.
|
||
- **Done when:** a toggle made in one browser shows in another, and survives a
|
||
reload.
|
||
|
||
**Phase 2 result (2026-09-29, uncommitted, NOT deployed):**
|
||
- **Console.**
|
||
- Endpoints were added to `api/doormile/endpoints.js`: `getAiAgents`,
|
||
`getAiSkills`, `getAiTools`, `updateAiSkill`, `createAiSkill`,
|
||
`updateAiAgent` and `getAiRegistryAudit`.
|
||
- Hooks were added in `lib/doormileHooks.js`: `useAiAgents`, `useAiSkills`,
|
||
`useAiTools` and three mutations. They share one `['doormile','ai']` key.
|
||
- The adapter layer is `agentStudio/registryAdapters.js`, a set of pure
|
||
functions.
|
||
- `agentRegistryData.js` now keeps only the surfaces list and the
|
||
selected-agent preference.
|
||
- **Changes on screen.**
|
||
- Skills are filtered to the selected agent; before, every agent's skills
|
||
showed.
|
||
- Agent status badges appear in the switcher.
|
||
- The skill drawer has a real enable toggle, a threshold editor (range and
|
||
step checked in the browser, then on the server), and shows source and
|
||
version.
|
||
- The tool table shows the kind, "Used by", the system each tool touches
|
||
and where it is implemented. It no longer calls a read-only tool
|
||
"Autonomous".
|
||
- Configure shows the agent's record, a model picker (engine agents only)
|
||
and an autonomy switch (gated agents only) with a typed confirmation.
|
||
The fake temperature slider and GPT/DeepSeek list are gone.
|
||
- Insights shows "No run data yet" instead of invented figures.
|
||
- "New skill" is disabled on AI_engine agents and for anyone who is not an
|
||
admin.
|
||
- With no saved choice the page opens on the first agent that has skills,
|
||
not on JARVIS, which has none.
|
||
- **Backend addition.** `airegistryaudit.changedbyemail` records the token's
|
||
email. The end-to-end run showed every audit row with `changedby = 0`: an
|
||
admin login without an appusers row carries user id 0.
|
||
- **Tests.**
|
||
- Console: 18 Agent Studio tests (adapters plus the page with only HTTP
|
||
mocked). The full suite is 46/46 suites and 1193 tests, and lint is clean.
|
||
- Backend: everything is green with the database attached, including the
|
||
parallel run that clashed before per-package schemas.
|
||
- **End-to-end, in a real browser, fully local.**
|
||
- Setup: a throwaway Postgres; the backend running with no `.env` and every
|
||
host pinned to localhost; a second console on :5174; throwaway admin and
|
||
manager logins.
|
||
- Checked in the database: seed counts, the toggle, the threshold change,
|
||
autonomy on (with confirmation) and off, and custom skill creation, each
|
||
with its audit row.
|
||
- Checked in the browser: the manager view is read-only.
|
||
- Checked by direct API call: a manager's write gets 403.
|
||
- Everything was removed afterwards: container, image, scripts, test
|
||
credentials and temporary launch entries.
|
||
|
||
### Phase 3 — Land the ops-layer skills on main
|
||
- Port the 8 rule-based skills, `tools.js`, the proposal executors and the banner
|
||
from the branch onto current main. Resolve the 10 conflicts; the branch's
|
||
`Deliveries.jsx` edits collide with today's uncommitted change.
|
||
- `SkillRegistry` then reads `enabled`/`thresholds` from `/admin/ai/skills` and
|
||
falls back to code defaults when offline. It stops reading `localStorage`.
|
||
- **Keep the branch invariant:** write tools return Proposals, and a human
|
||
confirms.
|
||
- **Done when:** the branch's 20+ test files pass on main, and a threshold changed
|
||
in Agent Studio changes the banner's output.
|
||
|
||
**Phase 3 result (2026-09-29, uncommitted, NOT deployed):**
|
||
- **Ported, as unstaged file copies (no merge).**
|
||
- The 8 skill definitions.
|
||
- `agent/{AgentFactory,signals,normalise,briefing,actions}.js`.
|
||
- `SlaRemediationCard`.
|
||
- `AgentOperationsBanner`, mounted on the Exceptions page.
|
||
- The `opsBriefing` chat intent in `lib/assistant/intents.js`.
|
||
- The "Needs attention" chip on the Exceptions context.
|
||
- 11 branch test suites.
|
||
- **Settings.** `SkillRegistry` now takes enabled and thresholds from
|
||
`/admin/ai/skills` through `useSkillRegistrySync`, mounted once in
|
||
`AdminLayout`. Nothing is kept in localStorage. It falls back to code
|
||
defaults, and says so on the banner, when the registry cannot be read.
|
||
- **Not ported.**
|
||
- `tools.js`: nothing imported it.
|
||
- `AgentStudioModal`: a second, localStorage-only settings UI. "Configure
|
||
skills" goes to Settings → Skills & Tools instead.
|
||
- `AgentDecisionDrawer` and the `/internal/agent-decisions` endpoints: they
|
||
return 403 from the console.
|
||
- The AI-panel "Autonomous Fleet Agent" card, and the branch's cosmetic
|
||
edits.
|
||
- **Defects found and fixed while porting.**
|
||
1. `OpenToast('success', msg)` has its arguments swapped; the signature is
|
||
`(message, variant)`. Every successful action would have shown a red
|
||
error toast reading "success". The branch's tests asserted the same
|
||
wrong order.
|
||
2. The `assignMiler` executor posted to `/hub/bookings/batch-assign`, which
|
||
is behind `HubStaffAuth` (role 6 only). Every console click would 403.
|
||
It is now review-only, with the reason in `actions.js`.
|
||
3. **Three skills could never fire.** High-Value COD, Cash Exposure and
|
||
Battery Safety read payment and battery fields that `/admin/bookings`
|
||
rows do not carry. They would report a false all-clear. They ship **off**
|
||
(`dataGap` in code, `Enabled: false` plus the reason in the seed).
|
||
4. The chat trigger was narrowed. It no longer claims "late/delayed orders"
|
||
or "operations summary", which would have replaced existing answers.
|
||
Routing tests pin both directions.
|
||
5. `SlaRemediationCard.test.jsx` could not have run: no `lucide-react` stub.
|
||
- **Registry seed updated to match main.**
|
||
- `CONSOLE_OPS_AGENT` is `live`.
|
||
- Paths point at main.
|
||
- The console tools are now the proposal verbs: `scan_bookings`,
|
||
`notify_riders` (the only executor), `assign_riders` (review-only), and
|
||
five review-only actions, each labelled REVIEW ONLY.
|
||
- New backend tests guard the no-data skills and the review-only labels.
|
||
- **Tests.**
|
||
- Console: 64/64 suites, 1308 tests; lint is clean and the build is green.
|
||
- Backend: build, vet and all tests are green. The Postgres-gated tests were
|
||
not re-run; the seed change is data-only and unit-tested.
|
||
- `Deliveries.jsx` was untouched; the uncommitted change there is still only
|
||
Suriya's.
|
||
- **Open items.**
|
||
- Feeding the three off skills needs `/admin/bookings` to include payment
|
||
amounts and mode, and the rider's battery. That is a backend
|
||
response-shape change.
|
||
- "Assign riders" needs an admin batch-assign route. `useBatchAssignBookings`
|
||
has the same 403 problem, but nothing calls it.
|
||
- `lib/assistant/CLAUDE.md` is now stale: it says proactive alerts were
|
||
"not started" and that the `components/assistant` copies are live, but
|
||
`AIPanel` imports `lib/assistant`.
|
||
|
||
### Phase 4 — Persist decisions and runs (makes Insights real)
|
||
- A Go NATS consumer on `telemetry.task` plus the engine's LLM decisions, written to
|
||
`agent_decisions`, plus a new `ai_agent_runs` table (⚠ schema change).
|
||
- `/admin/ai/decisions` and run stats feed the Insights tab and the Agents page,
|
||
replacing the `agentNetwork.js` snapshot.
|
||
- **Resolve before relying on vector search:** the `context_embedding`
|
||
1536-vs-384 dimension question.
|
||
|
||
**Phase 4 result (2026-09-29, uncommitted, NOT deployed):**
|
||
- **Finding.** Nothing in doormile_backend or AI_engine writes
|
||
`agent_decisions`. routemate (external) returns an `agent_decision_id` from
|
||
`/decide-assignment`, so it presumably writes through
|
||
`POST /internal/agent-decisions`. Whether production has rows is unverified.
|
||
AI_engine's two LLM decisions (stall, assignment-failure) are **not**
|
||
persisted anywhere; they appear only in logs.
|
||
- **Backend.**
|
||
- One new table, `aiagentruns`: append-only, unique on (agentid, taskid),
|
||
pruned after 30 days.
|
||
- `internal/ai/telemetry` queue-subscribes (`doormile-backend-telemetry`) to
|
||
`telemetry.task`, and writes runs in batches (2 s / 200). The NATS callback
|
||
never blocks: a full buffer drops the event, counts it and logs it.
|
||
- `telemetry.agent` heartbeats go to Redis (`ai:agent:state:<id>`, 5-minute
|
||
TTL), not Postgres.
|
||
- New endpoints, Doormile staff only: `GET /admin/ai/insights?days=1..30`
|
||
(runs, failures and average time per agent; decisions by type and outcome;
|
||
live state; a `receiving` flag) and `GET /admin/ai/decisions`
|
||
(keyset-paged; the `context` column is excluded because it holds rider
|
||
data).
|
||
- Windows use the backend clock (`utils.DBNow`). The engine's naive
|
||
timestamp is stored for display only.
|
||
- **Console.** Insights shows those figures with a 24 h / 7 d / 30 d window.
|
||
Silent agents appear as "silent" with zero runs rather than being left out.
|
||
When `receiving` is false, the page says telemetry is not received rather
|
||
than showing "0 runs" as if nothing happened.
|
||
- **Tests.**
|
||
- Backend: 11 telemetry unit tests, a Postgres-gated suite, and the new
|
||
routes in the gate tests.
|
||
- All 14 Postgres-gated tests (Phases 1 and 4) pass against a throwaway
|
||
`postgres:16-alpine`. This includes the parallel `go test ./...` run.
|
||
- Console: 64/64 suites, 1310 tests, lint clean, build green.
|
||
- **Live end-to-end run (2026-09-29, local, throwaway; all removed after).**
|
||
- Setup: throwaway Postgres and NATS; the real backend binary with no
|
||
`.env`; engine-shaped telemetry published over raw NATS.
|
||
- Six messages produced three rows. The redelivered task was ignored by the
|
||
unique index. The malformed event was dropped. The heartbeat was not
|
||
written to Postgres.
|
||
- `/admin/ai/insights` returned the right totals, failures and averages.
|
||
`/admin/ai/decisions` paged correctly and did not leak `context`.
|
||
- The Insights tab rendered the same figures with correct IST times.
|
||
- **Bugs the live run caught, fixed.**
|
||
1. The recorder stamped `utils.DBNow()` into a **timestamptz** column, which
|
||
AutoMigrate creates for new tables. A run received at 20:57 IST read back
|
||
as 02:27 the next day. It now uses `time.Now()`, and the window cutoffs
|
||
do too. `TestRecorderStampsARealInstant` fails with "5h30m off" if
|
||
DBNow comes back.
|
||
**Wider note:** `utils.DBNow` is only correct for the legacy
|
||
timestamp-WITHOUT-zone columns. Any table AutoMigrate creates fresh is
|
||
timestamptz, so audit other new tables before using DBNow in them.
|
||
2. Two Phase 1 Postgres fixtures still used `lookup_order`/`lookup_miler`,
|
||
which Phase 3 removed from the seed.
|
||
3. The Skills & Tools note still said the console skills were "not merged".
|
||
It now says they run on these settings, and that AI_engine does not
|
||
read them yet.
|
||
- **Deploy prerequisites.**
|
||
1. The backend's `NATS_URL` must point at the same NATS server AI_engine
|
||
publishes to (`NATS_HOST`/`NATS_PORT` there). Otherwise Insights shows
|
||
"Not receiving agent telemetry".
|
||
2. The Agents page still uses the 16–20 Sep snapshot. Moving it onto
|
||
`/admin/ai/insights` is a follow-up; it is dharaneesh's page.
|
||
3. To see AI_engine's LLM decisions in Insights, the engine would have to
|
||
POST them to `/internal/agent-decisions`. That is engine work (Phase 5).
|
||
|
||
### Phase 5 — Engine reads the registry
|
||
- The engine polls `GET /internal/ai/registry` (ETag, around 30 s) and applies
|
||
`enabled`, `autonomous`, `model` and thresholds without a restart. Today the
|
||
autonomy flags are read once at import.
|
||
- **Fix before any agent is shown as live:**
|
||
- `order_agent.py:182` (the enum does not exist)
|
||
- the JARVIS→ORDER payload key (`agent.py:279` vs `order_agent.py:144`)
|
||
- the DISPATCH→HUB id mismatch (`dispatch_agent.py:268` vs `hub_agent.py:151`)
|
||
- `release_vehicle_for_cancel` has no handler (`exception_agent.py:692`)
|
||
- messages without a `task_type` are silently dropped
|
||
- ORDER_AGENT still calls the renamed `crmbooking` route, with no auth header
|
||
- Mark HUB, FLEET and ROUTE_OPTIMIZER as `simulation` in the seed, or retire them.
|
||
- Add `pytest` and `pytest-asyncio` to `requirements.txt`.
|
||
|
||
#### Phase 5 — result (2026-09-29, uncommitted, not deployed)
|
||
- **Registry client** (`AI_engine/core/registry.py`).
|
||
- Polls `/internal/ai/registry` every `REGISTRY_POLL_SECONDS` (30 by default)
|
||
with `If-None-Match`. It is started from `main.py --production`.
|
||
- Precedence: once the registry has loaded, its value applies. Before that,
|
||
or if it never loads, the old env default applies.
|
||
- The last good copy survives 401/5xx/timeouts, so autonomy cannot flip
|
||
mid-shift.
|
||
- With no `INTERNAL_API_KEY`, it logs once and the engine runs on env
|
||
defaults.
|
||
- **What each agent now reads.**
|
||
|
||
| Agent | Skill gate | Autonomy | Other settings |
|
||
|---|---|---|---|
|
||
| Exception | `stall_response` | `EXCEPTION_AGENT` | `stallMinutes`, `reassignConfidence`, model |
|
||
| Dispatch | `assignment_failure_triage` | `DISPATCH_AGENT` | `realertEvery`, model |
|
||
| Express Dispatch | `express_batch_dispatch` | `EXPRESS_DISPATCH_AGENT` | `maxPerRider`, `maxRadiusKm`, `loadPenaltyKm` |
|
||
| Customer | `customer_notifications` | — | — |
|
||
|
||
A disabled skill means the agent logs the event and does nothing.
|
||
- **Model override.**
|
||
- `core/llm.request_params(model)` uses the agent's pinned model, or
|
||
`LLM_MODEL` when none is pinned.
|
||
- For a Haiku pin, thinking and effort are left out, because Haiku rejects
|
||
them (400).
|
||
- The console picker offers Opus 5.5, Sonnet 5.5 and Opus 4.8, plus "Engine
|
||
default". Haiku is left out because it is too weak for these decisions.
|
||
- **Decisions logged.**
|
||
- Stall and assignment-failure decisions are POSTed to
|
||
`/internal/agent-decisions` as `{decision_type, booking_id, context:{facts, model}, decision:{action, confidence}, reasoning}`.
|
||
- The post is fire-and-forget, so a slow backend never delays the reaction.
|
||
- These decisions now appear in Insights → Latest decisions.
|
||
- **Behaviour fix.** When the LLM is down and the agent is *not* autonomous,
|
||
it now escalates to a human. Before, it reassigned regardless of the
|
||
autonomy flag.
|
||
- **Message bugs fixed.**
|
||
1. `ORDER_STATUS_UPDATE` was added to the enum.
|
||
2. JARVIS→ORDER now sends `order_id`.
|
||
3. DISPATCH no longer forwards to HUB `prepare_receiving`. It sent a booking
|
||
id to a fictional-hub simulation.
|
||
4. FLEET handles `release_vehicle_for_cancel`, finding the vehicle by order
|
||
id.
|
||
5. CUSTOMER records `ORDER_CANCELLED` and `NOTIFICATION_SENT` instead of
|
||
dropping them. These come from simulated records, so no real customer
|
||
message is sent.
|
||
6. ORDER_AGENT refuses its backend calls and logs why. The seed now says it
|
||
is not connected. It stays `broken`.
|
||
- **Simulation agents.** HUB, FLEET and ROUTE_OPTIMIZER stay seeded as
|
||
`simulation`, and the Studio note says so.
|
||
- **Tests.**
|
||
- AI_engine: 95 unittest tests. 28 are new, in
|
||
`tests/test_registry_phase5.py`, all with no network. Three
|
||
`test_dispatch_agent` mocks were updated for the `model` argument.
|
||
- The only failures are the two modules that import pytest, and they failed
|
||
before this work. `pytest` and `pytest-asyncio` are now in
|
||
`requirements.txt` but are **not installed** in the venv.
|
||
- Console: Agent Studio suite green. Backend: `internal/ai/...` green.
|
||
- **Deploy prerequisites.**
|
||
1. The engine needs `GO_API_BASE_URL` and `INTERNAL_API_KEY`, the same key
|
||
the backend checks.
|
||
2. Before deploying, check the registry's current `enabled` and
|
||
`autonomous` values. On deploy they replace the env flags
|
||
(`AUTONOMOUS_REASSIGN` and the others).
|
||
|
||
### Phase 6 — Real Test playground
|
||
- Replace the `setTimeout` simulation with a backend endpoint that runs one prompt
|
||
through Claude tool-use, using the tools from the selected skill (their
|
||
`input_schema` from the registry).
|
||
- **Read tools execute. Write tools return a Proposal only** — the playground never
|
||
mutates production.
|
||
- Show the real trace: tool calls, arguments, results, latency and tokens. The
|
||
model comes from the registry.
|
||
#### Phase 6 — result (2026-09-29, uncommitted, not deployed)
|
||
- **Decisions:** use the official Go SDK, and redact personal data before
|
||
anything reaches Claude. **The SDK download was blocked by this machine's
|
||
permission check.** Everything else is built behind a `Model` interface.
|
||
The one missing piece is the ~80-line adapter from `playground.Request` to
|
||
`anthropic.MessageNewParams`. It needs
|
||
`go get github.com/anthropics/anthropic-sdk-go`, run or approved by a person.
|
||
Until then, `controllers.PlaygroundModel` is nil. The endpoint answers 503
|
||
`PLAYGROUND_NOT_CONFIGURED` and the Test tab says so; it never pretends to
|
||
run.
|
||
- **Backend** (`internal/ai/playground`, `controllers/aiPlaygroundController.go`).
|
||
- `POST /admin/ai/playground/run {agentid, skillid?, prompt}`, for Doormile
|
||
staff with roleid 1 only. The limit is 10 runs per user per 10 minutes,
|
||
prompts are at most 2000 characters, and a run times out after 120 s.
|
||
- The loop runs at most 6 turns with max_tokens 4096. The model comes from
|
||
the agent's registry pin, or `claude-opus-5-5` if none. The skill's tools
|
||
come from the registry, with `input_schema` taken from `inputschema`.
|
||
- Tool outcomes:
|
||
|
||
| Tool | Outcome |
|
||
|---|---|
|
||
| read, served by the backend (`get_booking_cache`, `scan_bookings`, `nearby_milers`) | **executed**, 5 s timeout |
|
||
| read, engine-only or external (`decide_*`, `sequence_stops`, `simulate_pricing_quote`, `list_express_*`) | **unavailable** |
|
||
| write / notify / event | **proposed**. Never executed; the model gets `{executed:false, proposal}` |
|
||
| not in the selected skill | **rejected** |
|
||
|
||
- **Redaction:**
|
||
- Executors select named non-personal columns only. There is no address,
|
||
name, phone or notes column, and coordinates are rounded to 2 dp.
|
||
- `Redact` then masks personal keys (name, phone, address, email, note,
|
||
reason, …) plus any email or Indian mobile number found in a string.
|
||
- Results are capped at 16 KB.
|
||
- The seed now gives `nearby_milers` (lat, lon, radius_km) and
|
||
`scan_bookings` (status, limit) real input schemas.
|
||
- **Console:**
|
||
- The Test tab calls the endpoint and shows the server's trace: each tool
|
||
call with its input, outcome, result and ms, plus the model, turns,
|
||
tokens and time.
|
||
- It has a skill picker, and "Test in Playground" preselects that skill.
|
||
Non-admins can't run it.
|
||
- 503, 429 and 403 errors each show a plain message.
|
||
- The setTimeout simulation is gone.
|
||
- **Tests:**
|
||
- Backend: 11 unit tests. They cover Prepare, every outcome, the turn limit,
|
||
model errors, truncation and redaction, including that dates are not
|
||
masked.
|
||
- A Postgres-gated test proves that personal columns in the table are never
|
||
selected. It passed on a throwaway `postgres:16-alpine`, which was removed
|
||
afterwards.
|
||
- Route tests: the gates, the 503 with no client, and bad input refused
|
||
before the model is called.
|
||
- Console: 5 new tests. Now 64/64 suites and 1315 tests; build green.
|
||
- **To switch it on:**
|
||
1. Add the SDK and the adapter.
|
||
2. Set `ANTHROPIC_API_KEY` on the backend.
|
||
3. Wire `controllers.PlaygroundModel` in `main.go`.
|
||
- **Earlier notes:**
|
||
1. **Official Go SDK or raw HTTP.** The Go SDK
|
||
(`github.com/anthropics/anthropic-sdk-go`) is not in the module cache, so
|
||
using it means a module download plus a new `go.mod` dependency.
|
||
2. **Data leaving for the Claude API.** The read tools (`scan_bookings`,
|
||
`get_booking_cache`, `nearby_milers`, `list_express_*`) return live
|
||
customer names, phones and addresses. Running them in the playground sends
|
||
that data to Anthropic. The options are: allow it; redact PII before it is
|
||
sent; or run against fixtures only.
|
||
- The backend also needs `ANTHROPIC_API_KEY`. Without it, the endpoint
|
||
returns 503 and the Test tab stays a labelled simulation.
|
||
|
||
---
|
||
|
||
## 6. Decisions
|
||
|
||
**Ratified 2026-09-29:** A = port, C = roleid 1 only. The Agents page is kept
|
||
and labelled. The secret check fails at boot in production only. B, D and E
|
||
follow the recommendations below unless changed.
|
||
|
||
| # | Decision | My recommendation |
|
||
|---|---|---|
|
||
| A | `feat/agentic-ops-layer`: merge or port? | **Port** the skills, tools and executors onto current main (Phase 3). Don't merge a 29-commit-stale branch with 10 conflicts. First confirm with dharaneesh that nothing newer exists elsewhere |
|
||
| B | Where does the registry live? | **`doormile_backend`/Postgres.** The engine has no API or persistence, and the console must not be the source of truth |
|
||
| C | Who may edit skills and autonomy? | **roleid 1 only.** Roles 3 and 4 read only. Today 1, 3 and 4 are identical everywhere, so this needs an explicit check |
|
||
| D | May the console toggle agent autonomy (auto-reassign riders, auto-notify customers)? | Yes, but only with a typed confirmation and an audit row. It stays **off** by default, matching compose |
|
||
| E | Delete or relabel the simulation agents (HUB, FLEET, ROUTE_OPTIMIZER)? | Seed them as `simulation`. Delete later if nobody objects |
|
||
|
||
---
|
||
|
||
## 7. Open questions (need checking, not guessing)
|
||
- **Is `booking.assignment_failed` published?** The engine's handoff doc says Go
|
||
doesn't publish it yet. Backend CLAUDE.md §4 names `publishAssignmentFailed`, and
|
||
the new retry window says `assignment_failed` fires after the first round. Check
|
||
which stream and subject it actually uses against what DISPATCH_AGENT binds to.
|
||
- **Which model is live?** The engine defaults to `LLM_MODEL=claude-opus-4-8`; the
|
||
console mock shows `claude-3-5-sonnet`. The registry's `model` field should hold
|
||
the id that is actually deployed. Set per agent, a cheaper model such as Haiku is
|
||
enough for the stall and assignment decisions.
|
||
- **Registry write safety:** admin handlers today are not tenant-guarded for
|
||
global data (pricing, hubs, app users). The registry endpoints must not copy that
|
||
pattern.
|