# Doormile Agent Platform — Plan & Agent Registry Status: **Phases 0–4 done (uncommitted, not deployed), Phase 5 next** · Written 2026-09-29 · Scope: `krow_talent_app` (the Doormile console), `doormile_backend` (Go API), `AI_engine` (Python agent swarm). This is the document `src/pages/doormile/settings/Settings.jsx` has been pointing at ("Phase 6 of docs/agent-platform-plan.md"). That file never existed until now, so the phase numbers below replace the ones the comment assumed. --- ## 1. Where things actually stand (verified 2026-09-29) ### Console — `krow_talent_app` - **Settings → Skills & Tools (Agent Studio) is a mock.** `agentRegistryData.js` holds 2 agents, 3 surfaces, 4 skills and 6 tools. Everything is saved to `localStorage` and nowhere else. The Test tab (`AgentPlayground.jsx:37`) is a `setTimeout` that always replies "succeeded with 100% confidence". Configure and Insights use hard-coded figures ("99.4%", "540 calls") and a model `claude-3-5-sonnet` that nothing reads. - **Three of the six mock tools are Krow leftovers:** `open_add_skill_training`, `open_add_training` and `query_learning_analytics` (workforce training, not logistics). Delete them; don't migrate them. - **The Agents page** (`/doormile/agents`) runs on a static snapshot in `src/lib/agentNetwork.js`, dated 16–20 Sep 2026. It is honest about being a snapshot, but its test (`tests/integration/agentsPage.test.jsx`) is stale: that is the 1 failing suite and 9 failing tests out of 45 suites. - **Unmerged branch `feat/agentic-ops-layer`** (dharaneesh, 2026-09-01) is the only real client-side agent work: - `SkillRegistry` with 8 rule-based skills: SlaGuardian, DoorstepStall, FleetBalancer, HighValueCod, RiderBatterySafety, HubCongestion, LateDispatch, CashExposure. - `tools.js` with 5 tools, following the rule "read-only tools execute, mutating tools return a Proposal". - AgentFactory, the ops briefing, proposal executors with a verify pass, and 20+ test files. - It is **29 commits behind main and conflicts in 10 files**. Its skill config is also stored in `localStorage`. - **The Home assistant** runs on the regex catalogue `src/lib/assistant/intents.js` (109 KB) and has no concept of skills. - **Other state:** lint shows 35 auto-fixable errors, and `Deliveries.jsx` has an uncommitted cosmetic change (a tidy-up of the batch filter). ### Backend — `doormile_backend` - Build, vet and test all pass (Go 1.26.4). There are 233 routes; CLAUDE.md is stale on counts, retry, routing, PIN auth and SMS (see its drift list). - **Nothing for an agent registry exists yet.** `internal/ai/` is an empty, untracked directory. - **What does exist:** - `models.AgentDecision` (table `agent_decisions`, pgvector 1536-dim — the 384-vs-1536 question is still open). - Internal routes, all behind `X-Internal-Key`: - `POST /internal/agent-decisions` - `GET /internal/agent-decisions/similar` - `PATCH /internal/agent-decisions/:id/outcome` - `/internal/express/{riders,bookings,assign}` - **Admin authorization is flat.** Roles 1, 3 and 4 get identical access, and no handler checks the role further. Many admin write handlers have no tenant guard. So any registry write endpoint must add its own role check (see §5). ### Engine — `AI_engine` - **It has no registry to read from.** Agents are hard-coded classes in `main.py:55-70`. `core/tool_registry.py` is an in-memory dict holding 3 tools, and none of them load in production. The two "skills" in `core/skills/` are stubs that nothing imports. - **The agent process has no HTTP API.** The Command Center (FastAPI, :8600) is a separate, unauthenticated NATS tap. - **Runs, escalations and decisions live in memory only.** They are capped at 500 and lost on restart. - **Tests:** 72 pass, but only with pytest installed by hand; `requirements.txt` lacks pytest. **Conclusion:** the registry must be **built**, not wired up. It belongs in `doormile_backend` (Postgres). The console reads it, and the engine reads it. Neither of them owns it. --- ## 2. The Agent Registry — what gets seeded This is the real inventory. Every row here becomes a seed row in Phase 1. "Status" is what the code does today, not what the docs claim. ### 2.1 Agents | id | Class / file | Purpose | Trigger | Writes to system | LLM | Status | |---|---|---|---|---|---|---| | `JARVIS` | `MasterAgent` · `core/agent.py:196` | Orchestrator, escalation inbox | `logistics.direct.JARVIS` | none | — | Partial. The inbox works; `orchestrate_order` sends the wrong payload key | | `DISPATCH_AGENT` | `agents/dispatch_agent.py:72` | Watches assignment outcomes, flags coverage gaps | JetStream `booking.assigned`, `booking.assignment_failed` | Alerts; customer notify only when autonomous | `decide_assignment_failure` | **Production-grade.** Idle until `assignment_failed` is confirmed as published (see §7) | | `EXCEPTION_AGENT` | `agents/exception_agent.py:106` | Stalled-rider detection and response | TRACKING `miler.location.updated`, `miler.stalled`, plus a 60 s DB sweep | `POST /internal/bookings/:id/reassign` (needs autonomy flag and confidence ≥ 0.7), `/internal/notify` | `decide_stall_response` | **Production-grade** (the stall path only) | | `EXPRESS_DISPATCH_AGENT` | `agents/express_dispatch_agent.py:80` | Tenant batch assign plus road sequencing | JetStream `express.dispatch_requested` | `POST /internal/express/assign` | — (greedy) | Implemented. Autonomy defaults to `true` in code and `false` in compose | | `CUSTOMER_AGENT` | `agents/customer_agent.py:50` | Customer notifications | Direct tasks, only from an autonomous Dispatch | `/internal/notify` | — | Implemented, rarely reached | | `ORDER_AGENT` | `agents/order_agent.py:15` | Order intake and validation | Direct tasks (no sender in production) | `/admin/crmbooking` — a **dead path, no auth header** | — | Broken. `update_status` crashes (`order_agent.py:182`) | | `HUB_AGENT` | `agents/hub_agent.py:35` | Hub capacity | Direct tasks | none | — | **Simulation** (8 fictional hubs) | | `FLEET_AGENT` | `agents/fleet_agent.py:34` | Vehicles | Direct tasks | none | — | **Simulation** (19 fake vehicles) | | `ROUTE_OPTIMIZER` | `agents/route_optimizer_agent.py:41` | Routing | Direct tasks | none | — | **Simulation** (haversine only) | | `CONSOLE_ASSISTANT` | `src/lib/assistant/*` (console) | Home chat: orders, bulk, assign, repeat | Operator prompt | Via proposals the operator confirms | RAG sidecar (off) | Live, regex-based | | `CONSOLE_OPS_AGENT` | `feat/agentic-ops-layer` | Ops briefing plus 8 monitoring skills | Page load / poll | Via proposals the operator confirms | — (rules) | **Unmerged** | The registry shows `status` as one of `live | partial | simulation | broken | unmerged | retired`. **The console must display this badge.** A simulation agent that looks live is the exact failure the Agents page comment warns about. ### 2.2 Tools (the real capabilities, not the stubs) `kind` is `read`, `write` or `notify`. Every `write` or `notify` tool has `requires_confirmation = true` unless its agent is explicitly autonomous. | name | Owner agent(s) | Kind | Target | Today lives at | |---|---|---|---|---| | `reassign_booking` | EXCEPTION | write | Go `/internal/bookings/:id/reassign` | `exception_agent.py:466` | | `notify_customer` | EXCEPTION, CUSTOMER | notify | Go `/internal/notify` | `exception_agent.py:476`, `customer_agent.py:349` | | `list_express_bookings` | EXPRESS | read | Go `/internal/express/bookings` | `express_dispatch_agent.py:351` | | `list_express_riders` | EXPRESS | read | Go `/internal/express/riders` | `express_dispatch_agent.py:342` | | `assign_express_batch` | EXPRESS | write | Go `/internal/express/assign` | `express_dispatch_agent.py:222` | | `sequence_stops` | EXPRESS | read (compute) | `routes.workolik.com /optimization/doormile/sequence` | `express_dispatch_agent.py:308` | | `get_booking_cache` | CUSTOMER | read | Go `/bookings/cache/:id` | `customer_agent.py:243` | | `nearby_milers` | DISPATCH, EXCEPTION | read | Redis GEO `milers:locations` | `dispatch_agent.py:170-240` | | `publish_miler_stalled` | EXCEPTION | write (event) | NATS `miler.stalled` | `exception_agent.py:377` | | `decide_stall_response` | EXCEPTION | read (LLM) | Claude | `core/llm.py:157` | | `decide_assignment_failure` | DISPATCH | read (LLM) | Claude | `core/llm.py:227` | | `record_agent_decision` | all | write | Go `/internal/agent-decisions` | Go route exists | | `diagnose_operations` | CONSOLE_OPS | read | Console queries | branch `tools.js` | | `lookup_order` | CONSOLE_OPS | read | `/admin/bookings` | branch `tools.js` | | `lookup_miler` | CONSOLE_OPS | read | `/admin/milers` | branch `tools.js` | | `propose_miler_reassignment` | CONSOLE_OPS | write → proposal | `/admin/bookings/:id/assign-miler` | branch `tools.js` | | `simulate_pricing_quote` | CONSOLE_OPS, CONSOLE_ASSISTANT | read | `/admin/pricing/simulate` | branch `tools.js` | | `create_single_order` | CONSOLE_ASSISTANT | write → proposal | `/admin/expressbooking` | current mock | | `rebalance_riders` | CONSOLE_OPS | write → proposal | `/hub/bookings/batch-assign` | current mock (no backing code) | **Not seeded:** - `ask_question`, `order_intake_skill`, `repeat_run_skill` (engine stubs that are never loaded). - The 3 Krow training tools. - ORDER_AGENT's `crmbooking` calls: that route was renamed to `expressbooking`, so those calls hit nothing. ### 2.3 Skills A skill is a named behaviour that belongs to one agent and uses one or more tools. It has a switch (`enabled`) and tunable `thresholds` (JSON, validated against a per-skill schema). | id | Agent | Tools | Source | |---|---|---|---| | `stall_response` | EXCEPTION | nearby_milers, decide_stall_response, reassign_booking, notify_customer | engine | | `assignment_failure_triage` | DISPATCH | nearby_milers, decide_assignment_failure | engine | | `express_batch_dispatch` | EXPRESS | list_express_*, assign_express_batch, sequence_stops | engine | | `sla_guardian` · `doorstep_stall` · `fleet_balancer` · `high_value_cod` · `rider_battery_safety` · `hub_congestion` · `late_dispatch` · `cash_exposure` | CONSOLE_OPS | branch `tools.js` set | branch (thresholds move from localStorage to the registry) | | `order_intake_auto_schedule` | CONSOLE_ASSISTANT | create_single_order, simulate_pricing_quote | current mock, real behaviour in `orderFlow.js` | | `dispatch_rebalance` | CONSOLE_OPS | rebalance_riders | current mock. **Keep disabled** until a backing endpoint exists | ### 2.4 Surfaces `/doormile/home` (assistant), `/doormile/control-x` (dispatch board), `/doormile/agents` (status board), and the ops banner (branch `AgentOperationsBanner`). `/doormile/dispatch` is only a redirect to Control X, so it is not a separate surface. --- ## 3. Schema (Phase 1) — ⚠ additive schema change, review before deploy New tables in `doormile_backend`, added through GORM AutoMigrate. They are additive only and touch no existing table. ``` ai_agents id text PK, name, class_ref, runtime ('engine'|'console'), purpose, trigger jsonb, status, autonomous bool, model text NULL, created_at, updated_at ai_tools name text PK, description, kind ('read'|'write'|'notify'), target text, input_schema jsonb, requires_confirmation bool, enabled bool, created_at, updated_at ai_skills id text PK, agent_id FK→ai_agents, title, category, description, sample_prompt, enabled bool, thresholds jsonb, thresholds_schema jsonb, version int, updated_by int NULL, updated_at ai_skill_tools skill_id FK, tool_name FK, PRIMARY KEY (skill_id, tool_name) ai_registry_audit id bigserial, entity, entity_id, field, old jsonb, new jsonb, changed_by int, changed_at ``` Rules: - **Tools and agents are code-defined.** The console can toggle `enabled` and `autonomous`, but it cannot invent a tool. A tool with no implementation is a lie on screen. "New skill" in the console therefore picks from existing tools only. - **Every write goes into `ai_registry_audit` in the same transaction.** - **Seeding is idempotent** (upsert by id) and lives in `migrations/`, not `scratch/`. - **Never store a secret.** Env var *names* only; the engine's `.env` stays out of the registry. --- ## 4. API contract (Phase 1) All handlers go in `controllers/aiRegistryController.go` and the logic in `internal/ai/registry`. Responses use `utils.OK` / `utils.List`, following the console conventions. | Method | Path | Auth | Notes | |---|---|---|---| | GET | `/admin/ai/agents` | admin (1,3,4), Doormile staff only | Includes a `status` badge and skill/tool counts | | GET | `/admin/ai/agents/:id` | same | Agent plus its skills and tools | | GET | `/admin/ai/skills` | same | `?agent=` filter | | GET | `/admin/ai/tools` | same | `?kind=` filter | | PATCH | `/admin/ai/skills/:id` | **roleid 1 only** | Allowed fields: `enabled`, `thresholds` (validated against schema). Bumps `version` and writes an audit row | | POST | `/admin/ai/skills` | roleid 1 only | New skill from existing tools only | | PATCH | `/admin/ai/agents/:id` | roleid 1 only | Allowed fields: `autonomous`, `model`. **Extra confirm in UI** — this changes what an agent does without a human | | GET | `/admin/ai/audit` | admin | Registry change history | | GET | `/admin/ai/decisions` | admin | Paged `agent_decisions` (Phase 4) | | GET | `/internal/ai/registry` | `X-Internal-Key` | The engine pulls its config from here (Phase 5). It sends an ETag so the engine can poll cheaply | "Doormile staff only" means `consoleTenantID == 0`. A tenant or client login gets 403. It does **not** get an empty list, because an empty list would hide a misconfiguration. --- ## 5. Phases Each phase ships on its own and leaves the system working. Nothing is committed or pushed without an explicit ask. ### Phase 0 — Prerequisites and hygiene (small, do first) 1. **Rotate secrets and untrack them.** - `doormile_backend`: `.env` and `doormile-abee7-*.json` are tracked; the `.gitignore` line `#.env` is commented out. - `AI_engine`: `.env` is tracked. - `config/config.go:60-71` has hard-coded fallback secrets. Fail at boot instead. - This is a **user action** (rotation needs the providers' consoles). I can do the untracking and ignore rules. 2. **Decide `feat/agentic-ops-layer`** (see §6, decision A). 3. **Console:** fix the stale `agentsPage.test.jsx`, run `lint:fix` for the 35 unused imports, and delete the 3 Krow training tools from `agentRegistryData.js`. 4. **Add a PREVIEW banner** on the Agent Studio tab itself. Today only a code comment says so; Agents shows a snapshot label, Agent Studio shows nothing. 5. **Backend:** remove the broken `.claude/skills/*` symlink stubs, `.agents/`, `skills.md` and `skills-lock.json`. The global plugin already provides these skills. **Done when:** tests are green, lint is clean, and no secret is in `git ls-files`. **Phase 0 result (2026-09-29, uncommitted):** - Console: lint clean; **46/46 suites, 1192 tests pass**. - Agent Studio: 4 Krow tools and 2 Krow skills removed (the plan said 3 tools; `open_miletruth_ai` was a fourth). `dispatch_rebalance` ships disabled. Storage keys moved to `_v3`. An on-screen Preview note was added. - Agents page: this was not just a stale test. The 24–25 Sep rebuild presented a simulation as live. Per decision, the design was kept and labelled: - a `Snapshot · 16–20 Sep 2026` stamp; - "Sample Activity — Simulation · not live data"; - no pulsing dot; - status counts taken from `networkStats()`; - "Autonomy gates on: 0 / 3", where the old "0 / 8" implied 8 gates. The tests were rewritten, keeping the honesty checks. - Backend: - **Reverted 2026-09-29 at Suriya's request:** these are back in git exactly as at HEAD: - the untracking of `.env` and the service-account key (the key is still committed, so rotation still stands); - the `.gitignore` edit; - the removal of `skills.md`, `skills-lock.json`, `.agents/` and `.claude/skills/`. AI_engine's `.env` is tracked again too. Do not redo any of this without asking. - `main.go` now refuses to boot when `ENV=production` and `JWT_SECRET_KEY`, `DB_PASSWORD` or `NATS_PASSWORD` is unset. - build, vet and test are green. - AI_engine: `.env` was untracked (it was already in `.gitignore`). - **Still yours:** - Rotate every secret that was committed. Git history still holds the values. - **Before the next backend deploy,** confirm that production sets all three secrets. If it has been running on a fallback, the new check stops it from starting. ### Phase 1 — Registry in the backend - Add the §3 tables, the idempotent seed from §2, and the §4 read endpoints plus PATCH/POST with the role check and audit trail. - **Tests:** a seed-idempotency test, a role test (roles 3 and 4 get 403 on PATCH; a tenant login gets 403 on GET), a threshold-schema validation test, and an audit test. - **Done when:** `go build/vet/test` is green and `curl` against a staging DB returns the §2 inventory. **Phase 1 result (2026-09-29, uncommitted, NOT deployed, no real DB touched):** - **Tables.** They follow the codebase's naming, not the names in §3: `aiagents`, `aitools`, `aiskills`, `aiskilltools`, `airegistryaudit`. The columns are as in §3, with a few changes: - The agent's trigger column is named `wakeon`. - `hasautonomygate` is new. Autonomy can only be set on Dispatch, Exception and Express. - `source` (engine/console/custom) is on skills. - There is no `enabled` flag on tools. - **Code.** - `internal/ai/registry`: seed, rules, store. - `controllers/aiRegistryController.go` - `middlewares/staff_only.go` (`DoormileStaffOnly`) - Routes are under `/admin/ai/*` and `/internal/ai/registry`, as in §4. - The seed runs in `migrations.Migrate`. It logs a failure and does not stop the boot. - **Seed.** 11 agents, 26 tools and 15 skills. - Four skills were added to §2.3 so that every tool belongs to a skill: `customer_notifications`, `ops_briefing`, and the 8 branch tools folded into their skills. - The console ops skills keep the branch ids and threshold keys, so Phase 3 is a 1:1 mapping. - `record_agent_decision` was dropped. It is a log the Go side writes, not a capability. - **Rules enforced server-side.** - Switching autonomy ON needs `confirm` set to the agent id. - Model ids must match `claude-*`, and only engine agents have one. - Thresholds are validated for range and step, and a patch is all-or-nothing. - Custom skills can be added to console agents only, and only from existing tools. - A patch that changes nothing does not bump the version or write an audit row. - **Tests.** - 26 unit tests and 6 HTTP gate tests always run. - 10 Postgres integration tests and 1 HTTP end-to-end test run only when `REGISTRY_TEST_DSN` is set. They need a throwaway database; each package uses its own schema. - All of them passed against a disposable `postgres:16-alpine` container. - **Bug the Postgres run caught.** With `enabled` tagged `default:true`, gorm dropped `false` from the INSERT, so `dispatch_rebalance` came up **enabled**. Fixed by removing the column default. - **Not yet proven.** The migration has not run against the real database; that happens on your next deploy. It is 5 new tables and touches nothing existing. ### Phase 2 — Console Agent Studio reads the registry - Replace `agentRegistryData.js` and its `localStorage` with React Query hooks (`useAiAgents`, `useAiSkills`, `useAiTools`) in `src/lib/doormileHooks.js`, and add the endpoints to `src/api/doormile/endpoints.js`. - Keep the existing components; only the data source changes. Also show the status badge, the tool `kind`, and a confirmation marker. - The skill toggle and "New skill" become real PATCH and POST calls. Add a `loading`/`error` state; remove the optimistic toast that claims success before the server answers. - Configure tab: model and autonomy from the registry. The temperature slider is **dropped**, because nothing reads it. - Insights and Test stay behind the PREVIEW banner until Phases 4 and 6. - Rewrite `tests/integration/agentStudio.test.jsx` against mocked hooks. - **Done when:** a toggle made in one browser shows in another, and survives a reload. **Phase 2 result (2026-09-29, uncommitted, NOT deployed):** - **Console.** - Endpoints were added to `api/doormile/endpoints.js`: `getAiAgents`, `getAiSkills`, `getAiTools`, `updateAiSkill`, `createAiSkill`, `updateAiAgent` and `getAiRegistryAudit`. - Hooks were added in `lib/doormileHooks.js`: `useAiAgents`, `useAiSkills`, `useAiTools` and three mutations. They share one `['doormile','ai']` key. - The adapter layer is `agentStudio/registryAdapters.js`, a set of pure functions. - `agentRegistryData.js` now keeps only the surfaces list and the selected-agent preference. - **Changes on screen.** - Skills are filtered to the selected agent; before, every agent's skills showed. - Agent status badges appear in the switcher. - The skill drawer has a real enable toggle, a threshold editor (range and step checked in the browser, then on the server), and shows source and version. - The tool table shows the kind, "Used by", the system each tool touches and where it is implemented. It no longer calls a read-only tool "Autonomous". - Configure shows the agent's record, a model picker (engine agents only) and an autonomy switch (gated agents only) with a typed confirmation. The fake temperature slider and GPT/DeepSeek list are gone. - Insights shows "No run data yet" instead of invented figures. - "New skill" is disabled on AI_engine agents and for anyone who is not an admin. - With no saved choice the page opens on the first agent that has skills, not on JARVIS, which has none. - **Backend addition.** `airegistryaudit.changedbyemail` records the token's email. The end-to-end run showed every audit row with `changedby = 0`: an admin login without an appusers row carries user id 0. - **Tests.** - Console: 18 Agent Studio tests (adapters plus the page with only HTTP mocked). The full suite is 46/46 suites and 1193 tests, and lint is clean. - Backend: everything is green with the database attached, including the parallel run that clashed before per-package schemas. - **End-to-end, in a real browser, fully local.** - Setup: a throwaway Postgres; the backend running with no `.env` and every host pinned to localhost; a second console on :5174; throwaway admin and manager logins. - Checked in the database: seed counts, the toggle, the threshold change, autonomy on (with confirmation) and off, and custom skill creation, each with its audit row. - Checked in the browser: the manager view is read-only. - Checked by direct API call: a manager's write gets 403. - Everything was removed afterwards: container, image, scripts, test credentials and temporary launch entries. ### Phase 3 — Land the ops-layer skills on main - Port the 8 rule-based skills, `tools.js`, the proposal executors and the banner from the branch onto current main. Resolve the 10 conflicts; the branch's `Deliveries.jsx` edits collide with today's uncommitted change. - `SkillRegistry` then reads `enabled`/`thresholds` from `/admin/ai/skills` and falls back to code defaults when offline. It stops reading `localStorage`. - **Keep the branch invariant:** write tools return Proposals, and a human confirms. - **Done when:** the branch's 20+ test files pass on main, and a threshold changed in Agent Studio changes the banner's output. **Phase 3 result (2026-09-29, uncommitted, NOT deployed):** - **Ported, as unstaged file copies (no merge).** - The 8 skill definitions. - `agent/{AgentFactory,signals,normalise,briefing,actions}.js`. - `SlaRemediationCard`. - `AgentOperationsBanner`, mounted on the Exceptions page. - The `opsBriefing` chat intent in `lib/assistant/intents.js`. - The "Needs attention" chip on the Exceptions context. - 11 branch test suites. - **Settings.** `SkillRegistry` now takes enabled and thresholds from `/admin/ai/skills` through `useSkillRegistrySync`, mounted once in `AdminLayout`. Nothing is kept in localStorage. It falls back to code defaults, and says so on the banner, when the registry cannot be read. - **Not ported.** - `tools.js`: nothing imported it. - `AgentStudioModal`: a second, localStorage-only settings UI. "Configure skills" goes to Settings → Skills & Tools instead. - `AgentDecisionDrawer` and the `/internal/agent-decisions` endpoints: they return 403 from the console. - The AI-panel "Autonomous Fleet Agent" card, and the branch's cosmetic edits. - **Defects found and fixed while porting.** 1. `OpenToast('success', msg)` has its arguments swapped; the signature is `(message, variant)`. Every successful action would have shown a red error toast reading "success". The branch's tests asserted the same wrong order. 2. The `assignMiler` executor posted to `/hub/bookings/batch-assign`, which is behind `HubStaffAuth` (role 6 only). Every console click would 403. It is now review-only, with the reason in `actions.js`. 3. **Three skills could never fire.** High-Value COD, Cash Exposure and Battery Safety read payment and battery fields that `/admin/bookings` rows do not carry. They would report a false all-clear. They ship **off** (`dataGap` in code, `Enabled: false` plus the reason in the seed). 4. The chat trigger was narrowed. It no longer claims "late/delayed orders" or "operations summary", which would have replaced existing answers. Routing tests pin both directions. 5. `SlaRemediationCard.test.jsx` could not have run: no `lucide-react` stub. - **Registry seed updated to match main.** - `CONSOLE_OPS_AGENT` is `live`. - Paths point at main. - The console tools are now the proposal verbs: `scan_bookings`, `notify_riders` (the only executor), `assign_riders` (review-only), and five review-only actions, each labelled REVIEW ONLY. - New backend tests guard the no-data skills and the review-only labels. - **Tests.** - Console: 64/64 suites, 1308 tests; lint is clean and the build is green. - Backend: build, vet and all tests are green. The Postgres-gated tests were not re-run; the seed change is data-only and unit-tested. - `Deliveries.jsx` was untouched; the uncommitted change there is still only Suriya's. - **Open items.** - Feeding the three off skills needs `/admin/bookings` to include payment amounts and mode, and the rider's battery. That is a backend response-shape change. - "Assign riders" needs an admin batch-assign route. `useBatchAssignBookings` has the same 403 problem, but nothing calls it. - `lib/assistant/CLAUDE.md` is now stale: it says proactive alerts were "not started" and that the `components/assistant` copies are live, but `AIPanel` imports `lib/assistant`. ### Phase 4 — Persist decisions and runs (makes Insights real) - A Go NATS consumer on `telemetry.task` plus the engine's LLM decisions, written to `agent_decisions`, plus a new `ai_agent_runs` table (⚠ schema change). - `/admin/ai/decisions` and run stats feed the Insights tab and the Agents page, replacing the `agentNetwork.js` snapshot. - **Resolve before relying on vector search:** the `context_embedding` 1536-vs-384 dimension question. **Phase 4 result (2026-09-29, uncommitted, NOT deployed):** - **Finding.** Nothing in doormile_backend or AI_engine writes `agent_decisions`. routemate (external) returns an `agent_decision_id` from `/decide-assignment`, so it presumably writes through `POST /internal/agent-decisions`. Whether production has rows is unverified. AI_engine's two LLM decisions (stall, assignment-failure) are **not** persisted anywhere; they appear only in logs. - **Backend.** - One new table, `aiagentruns`: append-only, unique on (agentid, taskid), pruned after 30 days. - `internal/ai/telemetry` queue-subscribes (`doormile-backend-telemetry`) to `telemetry.task`, and writes runs in batches (2 s / 200). The NATS callback never blocks: a full buffer drops the event, counts it and logs it. - `telemetry.agent` heartbeats go to Redis (`ai:agent:state:`, 5-minute TTL), not Postgres. - New endpoints, Doormile staff only: `GET /admin/ai/insights?days=1..30` (runs, failures and average time per agent; decisions by type and outcome; live state; a `receiving` flag) and `GET /admin/ai/decisions` (keyset-paged; the `context` column is excluded because it holds rider data). - Windows use the backend clock (`utils.DBNow`). The engine's naive timestamp is stored for display only. - **Console.** Insights shows those figures with a 24 h / 7 d / 30 d window. Silent agents appear as "silent" with zero runs rather than being left out. When `receiving` is false, the page says telemetry is not received rather than showing "0 runs" as if nothing happened. - **Tests.** - Backend: 11 telemetry unit tests, a Postgres-gated suite, and the new routes in the gate tests. - All 14 Postgres-gated tests (Phases 1 and 4) pass against a throwaway `postgres:16-alpine`. This includes the parallel `go test ./...` run. - Console: 64/64 suites, 1310 tests, lint clean, build green. - **Live end-to-end run (2026-09-29, local, throwaway; all removed after).** - Setup: throwaway Postgres and NATS; the real backend binary with no `.env`; engine-shaped telemetry published over raw NATS. - Six messages produced three rows. The redelivered task was ignored by the unique index. The malformed event was dropped. The heartbeat was not written to Postgres. - `/admin/ai/insights` returned the right totals, failures and averages. `/admin/ai/decisions` paged correctly and did not leak `context`. - The Insights tab rendered the same figures with correct IST times. - **Bugs the live run caught, fixed.** 1. The recorder stamped `utils.DBNow()` into a **timestamptz** column, which AutoMigrate creates for new tables. A run received at 20:57 IST read back as 02:27 the next day. It now uses `time.Now()`, and the window cutoffs do too. `TestRecorderStampsARealInstant` fails with "5h30m off" if DBNow comes back. **Wider note:** `utils.DBNow` is only correct for the legacy timestamp-WITHOUT-zone columns. Any table AutoMigrate creates fresh is timestamptz, so audit other new tables before using DBNow in them. 2. Two Phase 1 Postgres fixtures still used `lookup_order`/`lookup_miler`, which Phase 3 removed from the seed. 3. The Skills & Tools note still said the console skills were "not merged". It now says they run on these settings, and that AI_engine does not read them yet. - **Deploy prerequisites.** 1. The backend's `NATS_URL` must point at the same NATS server AI_engine publishes to (`NATS_HOST`/`NATS_PORT` there). Otherwise Insights shows "Not receiving agent telemetry". 2. The Agents page still uses the 16–20 Sep snapshot. Moving it onto `/admin/ai/insights` is a follow-up; it is dharaneesh's page. 3. To see AI_engine's LLM decisions in Insights, the engine would have to POST them to `/internal/agent-decisions`. That is engine work (Phase 5). ### Phase 5 — Engine reads the registry - The engine polls `GET /internal/ai/registry` (ETag, around 30 s) and applies `enabled`, `autonomous`, `model` and thresholds without a restart. Today the autonomy flags are read once at import. - **Fix before any agent is shown as live:** - `order_agent.py:182` (the enum does not exist) - the JARVIS→ORDER payload key (`agent.py:279` vs `order_agent.py:144`) - the DISPATCH→HUB id mismatch (`dispatch_agent.py:268` vs `hub_agent.py:151`) - `release_vehicle_for_cancel` has no handler (`exception_agent.py:692`) - messages without a `task_type` are silently dropped - ORDER_AGENT still calls the renamed `crmbooking` route, with no auth header - Mark HUB, FLEET and ROUTE_OPTIMIZER as `simulation` in the seed, or retire them. - Add `pytest` and `pytest-asyncio` to `requirements.txt`. #### Phase 5 — result (2026-09-29, uncommitted, not deployed) - **Registry client** (`AI_engine/core/registry.py`). - Polls `/internal/ai/registry` every `REGISTRY_POLL_SECONDS` (30 by default) with `If-None-Match`. It is started from `main.py --production`. - Precedence: once the registry has loaded, its value applies. Before that, or if it never loads, the old env default applies. - The last good copy survives 401/5xx/timeouts, so autonomy cannot flip mid-shift. - With no `INTERNAL_API_KEY`, it logs once and the engine runs on env defaults. - **What each agent now reads.** | Agent | Skill gate | Autonomy | Other settings | |---|---|---|---| | Exception | `stall_response` | `EXCEPTION_AGENT` | `stallMinutes`, `reassignConfidence`, model | | Dispatch | `assignment_failure_triage` | `DISPATCH_AGENT` | `realertEvery`, model | | Express Dispatch | `express_batch_dispatch` | `EXPRESS_DISPATCH_AGENT` | `maxPerRider`, `maxRadiusKm`, `loadPenaltyKm` | | Customer | `customer_notifications` | — | — | A disabled skill means the agent logs the event and does nothing. - **Model override.** - `core/llm.request_params(model)` uses the agent's pinned model, or `LLM_MODEL` when none is pinned. - For a Haiku pin, thinking and effort are left out, because Haiku rejects them (400). - The console picker offers Opus 5.5, Sonnet 5.5 and Opus 4.8, plus "Engine default". Haiku is left out because it is too weak for these decisions. - **Decisions logged.** - Stall and assignment-failure decisions are POSTed to `/internal/agent-decisions` as `{decision_type, booking_id, context:{facts, model}, decision:{action, confidence}, reasoning}`. - The post is fire-and-forget, so a slow backend never delays the reaction. - These decisions now appear in Insights → Latest decisions. - **Behaviour fix.** When the LLM is down and the agent is *not* autonomous, it now escalates to a human. Before, it reassigned regardless of the autonomy flag. - **Message bugs fixed.** 1. `ORDER_STATUS_UPDATE` was added to the enum. 2. JARVIS→ORDER now sends `order_id`. 3. DISPATCH no longer forwards to HUB `prepare_receiving`. It sent a booking id to a fictional-hub simulation. 4. FLEET handles `release_vehicle_for_cancel`, finding the vehicle by order id. 5. CUSTOMER records `ORDER_CANCELLED` and `NOTIFICATION_SENT` instead of dropping them. These come from simulated records, so no real customer message is sent. 6. ORDER_AGENT refuses its backend calls and logs why. The seed now says it is not connected. It stays `broken`. - **Simulation agents.** HUB, FLEET and ROUTE_OPTIMIZER stay seeded as `simulation`, and the Studio note says so. - **Tests.** - AI_engine: 95 unittest tests. 28 are new, in `tests/test_registry_phase5.py`, all with no network. Three `test_dispatch_agent` mocks were updated for the `model` argument. - The only failures are the two modules that import pytest, and they failed before this work. `pytest` and `pytest-asyncio` are now in `requirements.txt` but are **not installed** in the venv. - Console: Agent Studio suite green. Backend: `internal/ai/...` green. - **Deploy prerequisites.** 1. The engine needs `GO_API_BASE_URL` and `INTERNAL_API_KEY`, the same key the backend checks. 2. Before deploying, check the registry's current `enabled` and `autonomous` values. On deploy they replace the env flags (`AUTONOMOUS_REASSIGN` and the others). ### Phase 6 — Real Test playground - Replace the `setTimeout` simulation with a backend endpoint that runs one prompt through Claude tool-use, using the tools from the selected skill (their `input_schema` from the registry). - **Read tools execute. Write tools return a Proposal only** — the playground never mutates production. - Show the real trace: tool calls, arguments, results, latency and tokens. The model comes from the registry. #### Phase 6 — result (2026-09-29, uncommitted, not deployed) - **Decisions:** use the official Go SDK, and redact personal data before anything reaches Claude. **The SDK download was blocked by this machine's permission check.** Everything else is built behind a `Model` interface. The one missing piece is the ~80-line adapter from `playground.Request` to `anthropic.MessageNewParams`. It needs `go get github.com/anthropics/anthropic-sdk-go`, run or approved by a person. Until then, `controllers.PlaygroundModel` is nil. The endpoint answers 503 `PLAYGROUND_NOT_CONFIGURED` and the Test tab says so; it never pretends to run. - **Backend** (`internal/ai/playground`, `controllers/aiPlaygroundController.go`). - `POST /admin/ai/playground/run {agentid, skillid?, prompt}`, for Doormile staff with roleid 1 only. The limit is 10 runs per user per 10 minutes, prompts are at most 2000 characters, and a run times out after 120 s. - The loop runs at most 6 turns with max_tokens 4096. The model comes from the agent's registry pin, or `claude-opus-5-5` if none. The skill's tools come from the registry, with `input_schema` taken from `inputschema`. - Tool outcomes: | Tool | Outcome | |---|---| | read, served by the backend (`get_booking_cache`, `scan_bookings`, `nearby_milers`) | **executed**, 5 s timeout | | read, engine-only or external (`decide_*`, `sequence_stops`, `simulate_pricing_quote`, `list_express_*`) | **unavailable** | | write / notify / event | **proposed**. Never executed; the model gets `{executed:false, proposal}` | | not in the selected skill | **rejected** | - **Redaction:** - Executors select named non-personal columns only. There is no address, name, phone or notes column, and coordinates are rounded to 2 dp. - `Redact` then masks personal keys (name, phone, address, email, note, reason, …) plus any email or Indian mobile number found in a string. - Results are capped at 16 KB. - The seed now gives `nearby_milers` (lat, lon, radius_km) and `scan_bookings` (status, limit) real input schemas. - **Console:** - The Test tab calls the endpoint and shows the server's trace: each tool call with its input, outcome, result and ms, plus the model, turns, tokens and time. - It has a skill picker, and "Test in Playground" preselects that skill. Non-admins can't run it. - 503, 429 and 403 errors each show a plain message. - The setTimeout simulation is gone. - **Tests:** - Backend: 11 unit tests. They cover Prepare, every outcome, the turn limit, model errors, truncation and redaction, including that dates are not masked. - A Postgres-gated test proves that personal columns in the table are never selected. It passed on a throwaway `postgres:16-alpine`, which was removed afterwards. - Route tests: the gates, the 503 with no client, and bad input refused before the model is called. - Console: 5 new tests. Now 64/64 suites and 1315 tests; build green. - **To switch it on:** 1. Add the SDK and the adapter. 2. Set `ANTHROPIC_API_KEY` on the backend. 3. Wire `controllers.PlaygroundModel` in `main.go`. - **Earlier notes:** 1. **Official Go SDK or raw HTTP.** The Go SDK (`github.com/anthropics/anthropic-sdk-go`) is not in the module cache, so using it means a module download plus a new `go.mod` dependency. 2. **Data leaving for the Claude API.** The read tools (`scan_bookings`, `get_booking_cache`, `nearby_milers`, `list_express_*`) return live customer names, phones and addresses. Running them in the playground sends that data to Anthropic. The options are: allow it; redact PII before it is sent; or run against fixtures only. - The backend also needs `ANTHROPIC_API_KEY`. Without it, the endpoint returns 503 and the Test tab stays a labelled simulation. --- ## 6. Decisions **Ratified 2026-09-29:** A = port, C = roleid 1 only. The Agents page is kept and labelled. The secret check fails at boot in production only. B, D and E follow the recommendations below unless changed. | # | Decision | My recommendation | |---|---|---| | A | `feat/agentic-ops-layer`: merge or port? | **Port** the skills, tools and executors onto current main (Phase 3). Don't merge a 29-commit-stale branch with 10 conflicts. First confirm with dharaneesh that nothing newer exists elsewhere | | B | Where does the registry live? | **`doormile_backend`/Postgres.** The engine has no API or persistence, and the console must not be the source of truth | | C | Who may edit skills and autonomy? | **roleid 1 only.** Roles 3 and 4 read only. Today 1, 3 and 4 are identical everywhere, so this needs an explicit check | | D | May the console toggle agent autonomy (auto-reassign riders, auto-notify customers)? | Yes, but only with a typed confirmation and an audit row. It stays **off** by default, matching compose | | E | Delete or relabel the simulation agents (HUB, FLEET, ROUTE_OPTIMIZER)? | Seed them as `simulation`. Delete later if nobody objects | --- ## 7. Open questions (need checking, not guessing) - **Is `booking.assignment_failed` published?** The engine's handoff doc says Go doesn't publish it yet. Backend CLAUDE.md §4 names `publishAssignmentFailed`, and the new retry window says `assignment_failed` fires after the first round. Check which stream and subject it actually uses against what DISPATCH_AGENT binds to. - **Which model is live?** The engine defaults to `LLM_MODEL=claude-opus-4-8`; the console mock shows `claude-3-5-sonnet`. The registry's `model` field should hold the id that is actually deployed. Set per agent, a cheaper model such as Haiku is enough for the stall and assignment decisions. - **Registry write safety:** admin handlers today are not tenant-guarded for global data (pricing, hubs, app users). The registry endpoints must not copy that pattern.