42 KiB
Doormile Agent Platform — Plan & Agent Registry
Status: Phases 0–4 done (uncommitted, not deployed), Phase 5 next · Written 2026-09-29 · Scope: krow_talent_app (the
Doormile console), doormile_backend (Go API), AI_engine (Python agent swarm).
This is the document src/pages/doormile/settings/Settings.jsx has been pointing
at ("Phase 6 of docs/agent-platform-plan.md"). That file never existed until now,
so the phase numbers below replace the ones the comment assumed.
1. Where things actually stand (verified 2026-09-29)
Console — krow_talent_app
- Settings → Skills & Tools (Agent Studio) is a mock.
agentRegistryData.jsholds 2 agents, 3 surfaces, 4 skills and 6 tools. Everything is saved tolocalStorageand nowhere else. The Test tab (AgentPlayground.jsx:37) is asetTimeoutthat always replies "succeeded with 100% confidence". Configure and Insights use hard-coded figures ("99.4%", "540 calls") and a modelclaude-3-5-sonnetthat nothing reads. - Three of the six mock tools are Krow leftovers:
open_add_skill_training,open_add_trainingandquery_learning_analytics(workforce training, not logistics). Delete them; don't migrate them. - The Agents page (
/doormile/agents) runs on a static snapshot insrc/lib/agentNetwork.js, dated 16–20 Sep 2026. It is honest about being a snapshot, but its test (tests/integration/agentsPage.test.jsx) is stale: that is the 1 failing suite and 9 failing tests out of 45 suites. - Unmerged branch
feat/agentic-ops-layer(dharaneesh, 2026-09-01) is the only real client-side agent work:SkillRegistrywith 8 rule-based skills: SlaGuardian, DoorstepStall, FleetBalancer, HighValueCod, RiderBatterySafety, HubCongestion, LateDispatch, CashExposure.tools.jswith 5 tools, following the rule "read-only tools execute, mutating tools return a Proposal".- AgentFactory, the ops briefing, proposal executors with a verify pass, and 20+ test files.
- It is 29 commits behind main and conflicts in 10 files. Its skill config is
also stored in
localStorage.
- The Home assistant runs on the regex catalogue
src/lib/assistant/intents.js(109 KB) and has no concept of skills. - Other state: lint shows 35 auto-fixable errors, and
Deliveries.jsxhas an uncommitted cosmetic change (a tidy-up of the batch filter).
Backend — doormile_backend
- Build, vet and test all pass (Go 1.26.4). There are 233 routes; CLAUDE.md is stale on counts, retry, routing, PIN auth and SMS (see its drift list).
- Nothing for an agent registry exists yet.
internal/ai/is an empty, untracked directory. - What does exist:
models.AgentDecision(tableagent_decisions, pgvector 1536-dim — the 384-vs-1536 question is still open).- Internal routes, all behind
X-Internal-Key:POST /internal/agent-decisionsGET /internal/agent-decisions/similarPATCH /internal/agent-decisions/:id/outcome/internal/express/{riders,bookings,assign}
- Admin authorization is flat. Roles 1, 3 and 4 get identical access, and no handler checks the role further. Many admin write handlers have no tenant guard. So any registry write endpoint must add its own role check (see §5).
Engine — AI_engine
- It has no registry to read from. Agents are hard-coded classes in
main.py:55-70.core/tool_registry.pyis an in-memory dict holding 3 tools, and none of them load in production. The two "skills" incore/skills/are stubs that nothing imports. - The agent process has no HTTP API. The Command Center (FastAPI, :8600) is a separate, unauthenticated NATS tap.
- Runs, escalations and decisions live in memory only. They are capped at 500 and lost on restart.
- Tests: 72 pass, but only with pytest installed by hand;
requirements.txtlacks pytest.
Conclusion: the registry must be built, not wired up. It belongs in
doormile_backend (Postgres). The console reads it, and the engine reads it.
Neither of them owns it.
2. The Agent Registry — what gets seeded
This is the real inventory. Every row here becomes a seed row in Phase 1. "Status" is what the code does today, not what the docs claim.
2.1 Agents
| id | Class / file | Purpose | Trigger | Writes to system | LLM | Status |
|---|---|---|---|---|---|---|
JARVIS |
MasterAgent · core/agent.py:196 |
Orchestrator, escalation inbox | logistics.direct.JARVIS |
none | — | Partial. The inbox works; orchestrate_order sends the wrong payload key |
DISPATCH_AGENT |
agents/dispatch_agent.py:72 |
Watches assignment outcomes, flags coverage gaps | JetStream booking.assigned, booking.assignment_failed |
Alerts; customer notify only when autonomous | decide_assignment_failure |
Production-grade. Idle until assignment_failed is confirmed as published (see §7) |
EXCEPTION_AGENT |
agents/exception_agent.py:106 |
Stalled-rider detection and response | TRACKING miler.location.updated, miler.stalled, plus a 60 s DB sweep |
POST /internal/bookings/:id/reassign (needs autonomy flag and confidence ≥ 0.7), /internal/notify |
decide_stall_response |
Production-grade (the stall path only) |
EXPRESS_DISPATCH_AGENT |
agents/express_dispatch_agent.py:80 |
Tenant batch assign plus road sequencing | JetStream express.dispatch_requested |
POST /internal/express/assign |
— (greedy) | Implemented. Autonomy defaults to true in code and false in compose |
CUSTOMER_AGENT |
agents/customer_agent.py:50 |
Customer notifications | Direct tasks, only from an autonomous Dispatch | /internal/notify |
— | Implemented, rarely reached |
ORDER_AGENT |
agents/order_agent.py:15 |
Order intake and validation | Direct tasks (no sender in production) | /admin/crmbooking — a dead path, no auth header |
— | Broken. update_status crashes (order_agent.py:182) |
HUB_AGENT |
agents/hub_agent.py:35 |
Hub capacity | Direct tasks | none | — | Simulation (8 fictional hubs) |
FLEET_AGENT |
agents/fleet_agent.py:34 |
Vehicles | Direct tasks | none | — | Simulation (19 fake vehicles) |
ROUTE_OPTIMIZER |
agents/route_optimizer_agent.py:41 |
Routing | Direct tasks | none | — | Simulation (haversine only) |
CONSOLE_ASSISTANT |
src/lib/assistant/* (console) |
Home chat: orders, bulk, assign, repeat | Operator prompt | Via proposals the operator confirms | RAG sidecar (off) | Live, regex-based |
CONSOLE_OPS_AGENT |
feat/agentic-ops-layer |
Ops briefing plus 8 monitoring skills | Page load / poll | Via proposals the operator confirms | — (rules) | Unmerged |
The registry shows status as one of live | partial | simulation | broken | unmerged | retired. The console must display this badge. A simulation agent
that looks live is the exact failure the Agents page comment warns about.
2.2 Tools (the real capabilities, not the stubs)
kind is read, write or notify. Every write or notify tool has
requires_confirmation = true unless its agent is explicitly autonomous.
| name | Owner agent(s) | Kind | Target | Today lives at |
|---|---|---|---|---|
reassign_booking |
EXCEPTION | write | Go /internal/bookings/:id/reassign |
exception_agent.py:466 |
notify_customer |
EXCEPTION, CUSTOMER | notify | Go /internal/notify |
exception_agent.py:476, customer_agent.py:349 |
list_express_bookings |
EXPRESS | read | Go /internal/express/bookings |
express_dispatch_agent.py:351 |
list_express_riders |
EXPRESS | read | Go /internal/express/riders |
express_dispatch_agent.py:342 |
assign_express_batch |
EXPRESS | write | Go /internal/express/assign |
express_dispatch_agent.py:222 |
sequence_stops |
EXPRESS | read (compute) | routes.workolik.com /optimization/doormile/sequence |
express_dispatch_agent.py:308 |
get_booking_cache |
CUSTOMER | read | Go /bookings/cache/:id |
customer_agent.py:243 |
nearby_milers |
DISPATCH, EXCEPTION | read | Redis GEO milers:locations |
dispatch_agent.py:170-240 |
publish_miler_stalled |
EXCEPTION | write (event) | NATS miler.stalled |
exception_agent.py:377 |
decide_stall_response |
EXCEPTION | read (LLM) | Claude | core/llm.py:157 |
decide_assignment_failure |
DISPATCH | read (LLM) | Claude | core/llm.py:227 |
record_agent_decision |
all | write | Go /internal/agent-decisions |
Go route exists |
diagnose_operations |
CONSOLE_OPS | read | Console queries | branch tools.js |
lookup_order |
CONSOLE_OPS | read | /admin/bookings |
branch tools.js |
lookup_miler |
CONSOLE_OPS | read | /admin/milers |
branch tools.js |
propose_miler_reassignment |
CONSOLE_OPS | write → proposal | /admin/bookings/:id/assign-miler |
branch tools.js |
simulate_pricing_quote |
CONSOLE_OPS, CONSOLE_ASSISTANT | read | /admin/pricing/simulate |
branch tools.js |
create_single_order |
CONSOLE_ASSISTANT | write → proposal | /admin/expressbooking |
current mock |
rebalance_riders |
CONSOLE_OPS | write → proposal | /hub/bookings/batch-assign |
current mock (no backing code) |
Not seeded:
ask_question,order_intake_skill,repeat_run_skill(engine stubs that are never loaded).- The 3 Krow training tools.
- ORDER_AGENT's
crmbookingcalls: that route was renamed toexpressbooking, so those calls hit nothing.
2.3 Skills
A skill is a named behaviour that belongs to one agent and uses one or more tools.
It has a switch (enabled) and tunable thresholds (JSON, validated against a
per-skill schema).
| id | Agent | Tools | Source |
|---|---|---|---|
stall_response |
EXCEPTION | nearby_milers, decide_stall_response, reassign_booking, notify_customer | engine |
assignment_failure_triage |
DISPATCH | nearby_milers, decide_assignment_failure | engine |
express_batch_dispatch |
EXPRESS | list_express_*, assign_express_batch, sequence_stops | engine |
sla_guardian · doorstep_stall · fleet_balancer · high_value_cod · rider_battery_safety · hub_congestion · late_dispatch · cash_exposure |
CONSOLE_OPS | branch tools.js set |
branch (thresholds move from localStorage to the registry) |
order_intake_auto_schedule |
CONSOLE_ASSISTANT | create_single_order, simulate_pricing_quote | current mock, real behaviour in orderFlow.js |
dispatch_rebalance |
CONSOLE_OPS | rebalance_riders | current mock. Keep disabled until a backing endpoint exists |
2.4 Surfaces
/doormile/home (assistant), /doormile/control-x (dispatch board),
/doormile/agents (status board), and the ops banner (branch
AgentOperationsBanner). /doormile/dispatch is only a redirect to Control X, so
it is not a separate surface.
3. Schema (Phase 1) — ⚠ additive schema change, review before deploy
New tables in doormile_backend, added through GORM AutoMigrate. They are
additive only and touch no existing table.
ai_agents id text PK, name, class_ref, runtime ('engine'|'console'),
purpose, trigger jsonb, status, autonomous bool,
model text NULL, created_at, updated_at
ai_tools name text PK, description, kind ('read'|'write'|'notify'),
target text, input_schema jsonb, requires_confirmation bool,
enabled bool, created_at, updated_at
ai_skills id text PK, agent_id FK→ai_agents, title, category,
description, sample_prompt, enabled bool,
thresholds jsonb, thresholds_schema jsonb,
version int, updated_by int NULL, updated_at
ai_skill_tools skill_id FK, tool_name FK, PRIMARY KEY (skill_id, tool_name)
ai_registry_audit id bigserial, entity, entity_id, field, old jsonb, new jsonb,
changed_by int, changed_at
Rules:
- Tools and agents are code-defined. The console can toggle
enabledandautonomous, but it cannot invent a tool. A tool with no implementation is a lie on screen. "New skill" in the console therefore picks from existing tools only. - Every write goes into
ai_registry_auditin the same transaction. - Seeding is idempotent (upsert by id) and lives in
migrations/, notscratch/. - Never store a secret. Env var names only; the engine's
.envstays out of the registry.
4. API contract (Phase 1)
All handlers go in controllers/aiRegistryController.go and the logic in
internal/ai/registry. Responses use utils.OK / utils.List, following the
console conventions.
| Method | Path | Auth | Notes |
|---|---|---|---|
| GET | /admin/ai/agents |
admin (1,3,4), Doormile staff only | Includes a status badge and skill/tool counts |
| GET | /admin/ai/agents/:id |
same | Agent plus its skills and tools |
| GET | /admin/ai/skills |
same | ?agent= filter |
| GET | /admin/ai/tools |
same | ?kind= filter |
| PATCH | /admin/ai/skills/:id |
roleid 1 only | Allowed fields: enabled, thresholds (validated against schema). Bumps version and writes an audit row |
| POST | /admin/ai/skills |
roleid 1 only | New skill from existing tools only |
| PATCH | /admin/ai/agents/:id |
roleid 1 only | Allowed fields: autonomous, model. Extra confirm in UI — this changes what an agent does without a human |
| GET | /admin/ai/audit |
admin | Registry change history |
| GET | /admin/ai/decisions |
admin | Paged agent_decisions (Phase 4) |
| GET | /internal/ai/registry |
X-Internal-Key |
The engine pulls its config from here (Phase 5). It sends an ETag so the engine can poll cheaply |
"Doormile staff only" means consoleTenantID == 0. A tenant or client login gets
403. It does not get an empty list, because an empty list would hide a
misconfiguration.
5. Phases
Each phase ships on its own and leaves the system working. Nothing is committed or pushed without an explicit ask.
Phase 0 — Prerequisites and hygiene (small, do first)
- Rotate secrets and untrack them.
doormile_backend:.envanddoormile-abee7-*.jsonare tracked; the.gitignoreline#.envis commented out.AI_engine:.envis tracked.config/config.go:60-71has hard-coded fallback secrets. Fail at boot instead.- This is a user action (rotation needs the providers' consoles). I can do the untracking and ignore rules.
- Decide
feat/agentic-ops-layer(see §6, decision A). - Console: fix the stale
agentsPage.test.jsx, runlint:fixfor the 35 unused imports, and delete the 3 Krow training tools fromagentRegistryData.js. - Add a PREVIEW banner on the Agent Studio tab itself. Today only a code comment says so; Agents shows a snapshot label, Agent Studio shows nothing.
- Backend: remove the broken
.claude/skills/*symlink stubs,.agents/,skills.mdandskills-lock.json. The global plugin already provides these skills.
Done when: tests are green, lint is clean, and no secret is in git ls-files.
Phase 0 result (2026-09-29, uncommitted):
- Console: lint clean; 46/46 suites, 1192 tests pass.
- Agent Studio: 4 Krow tools and 2 Krow skills removed (the plan said 3 tools;
open_miletruth_aiwas a fourth).dispatch_rebalanceships disabled. Storage keys moved to_v3. An on-screen Preview note was added. - Agents page: this was not just a stale test. The 24–25 Sep rebuild presented
a simulation as live. Per decision, the design was kept and labelled:
- a
Snapshot · 16–20 Sep 2026stamp; - "Sample Activity — Simulation · not live data";
- no pulsing dot;
- status counts taken from
networkStats(); - "Autonomy gates on: 0 / 3", where the old "0 / 8" implied 8 gates. The tests were rewritten, keeping the honesty checks.
- a
- Backend:
- Reverted 2026-09-29 at Suriya's request: these are back in git
exactly as at HEAD:
- the untracking of
.envand the service-account key (the key is still committed, so rotation still stands); - the
.gitignoreedit; - the removal of
skills.md,skills-lock.json,.agents/and.claude/skills/. AI_engine's.envis tracked again too. Do not redo any of this without asking.
- the untracking of
main.gonow refuses to boot whenENV=productionandJWT_SECRET_KEY,DB_PASSWORDorNATS_PASSWORDis unset.- build, vet and test are green.
- Reverted 2026-09-29 at Suriya's request: these are back in git
exactly as at HEAD:
- AI_engine:
.envwas untracked (it was already in.gitignore). - Still yours:
- Rotate every secret that was committed. Git history still holds the values.
- Before the next backend deploy, confirm that production sets all three secrets. If it has been running on a fallback, the new check stops it from starting.
Phase 1 — Registry in the backend
- Add the §3 tables, the idempotent seed from §2, and the §4 read endpoints plus PATCH/POST with the role check and audit trail.
- Tests: a seed-idempotency test, a role test (roles 3 and 4 get 403 on PATCH; a tenant login gets 403 on GET), a threshold-schema validation test, and an audit test.
- Done when:
go build/vet/testis green andcurlagainst a staging DB returns the §2 inventory.
Phase 1 result (2026-09-29, uncommitted, NOT deployed, no real DB touched):
- Tables. They follow the codebase's naming, not the names in §3:
aiagents,aitools,aiskills,aiskilltools,airegistryaudit. The columns are as in §3, with a few changes:- The agent's trigger column is named
wakeon. hasautonomygateis new. Autonomy can only be set on Dispatch, Exception and Express.source(engine/console/custom) is on skills.- There is no
enabledflag on tools.
- The agent's trigger column is named
- Code.
internal/ai/registry: seed, rules, store.controllers/aiRegistryController.gomiddlewares/staff_only.go(DoormileStaffOnly)- Routes are under
/admin/ai/*and/internal/ai/registry, as in §4. - The seed runs in
migrations.Migrate. It logs a failure and does not stop the boot.
- Seed. 11 agents, 26 tools and 15 skills.
- Four skills were added to §2.3 so that every tool belongs to a skill:
customer_notifications,ops_briefing, and the 8 branch tools folded into their skills. - The console ops skills keep the branch ids and threshold keys, so Phase 3 is a 1:1 mapping.
record_agent_decisionwas dropped. It is a log the Go side writes, not a capability.
- Four skills were added to §2.3 so that every tool belongs to a skill:
- Rules enforced server-side.
- Switching autonomy ON needs
confirmset to the agent id. - Model ids must match
claude-*, and only engine agents have one. - Thresholds are validated for range and step, and a patch is all-or-nothing.
- Custom skills can be added to console agents only, and only from existing tools.
- A patch that changes nothing does not bump the version or write an audit row.
- Switching autonomy ON needs
- Tests.
- 26 unit tests and 6 HTTP gate tests always run.
- 10 Postgres integration tests and 1 HTTP end-to-end test run only when
REGISTRY_TEST_DSNis set. They need a throwaway database; each package uses its own schema. - All of them passed against a disposable
postgres:16-alpinecontainer.
- Bug the Postgres run caught. With
enabledtaggeddefault:true, gorm droppedfalsefrom the INSERT, sodispatch_rebalancecame up enabled. Fixed by removing the column default. - Not yet proven. The migration has not run against the real database; that happens on your next deploy. It is 5 new tables and touches nothing existing.
Phase 2 — Console Agent Studio reads the registry
- Replace
agentRegistryData.jsand itslocalStoragewith React Query hooks (useAiAgents,useAiSkills,useAiTools) insrc/lib/doormileHooks.js, and add the endpoints tosrc/api/doormile/endpoints.js. - Keep the existing components; only the data source changes. Also show the status
badge, the tool
kind, and a confirmation marker. - The skill toggle and "New skill" become real PATCH and POST calls. Add a
loading/errorstate; remove the optimistic toast that claims success before the server answers. - Configure tab: model and autonomy from the registry. The temperature slider is dropped, because nothing reads it.
- Insights and Test stay behind the PREVIEW banner until Phases 4 and 6.
- Rewrite
tests/integration/agentStudio.test.jsxagainst mocked hooks. - Done when: a toggle made in one browser shows in another, and survives a reload.
Phase 2 result (2026-09-29, uncommitted, NOT deployed):
- Console.
- Endpoints were added to
api/doormile/endpoints.js:getAiAgents,getAiSkills,getAiTools,updateAiSkill,createAiSkill,updateAiAgentandgetAiRegistryAudit. - Hooks were added in
lib/doormileHooks.js:useAiAgents,useAiSkills,useAiToolsand three mutations. They share one['doormile','ai']key. - The adapter layer is
agentStudio/registryAdapters.js, a set of pure functions. agentRegistryData.jsnow keeps only the surfaces list and the selected-agent preference.
- Endpoints were added to
- Changes on screen.
- Skills are filtered to the selected agent; before, every agent's skills showed.
- Agent status badges appear in the switcher.
- The skill drawer has a real enable toggle, a threshold editor (range and step checked in the browser, then on the server), and shows source and version.
- The tool table shows the kind, "Used by", the system each tool touches and where it is implemented. It no longer calls a read-only tool "Autonomous".
- Configure shows the agent's record, a model picker (engine agents only) and an autonomy switch (gated agents only) with a typed confirmation. The fake temperature slider and GPT/DeepSeek list are gone.
- Insights shows "No run data yet" instead of invented figures.
- "New skill" is disabled on AI_engine agents and for anyone who is not an admin.
- With no saved choice the page opens on the first agent that has skills, not on JARVIS, which has none.
- Backend addition.
airegistryaudit.changedbyemailrecords the token's email. The end-to-end run showed every audit row withchangedby = 0: an admin login without an appusers row carries user id 0. - Tests.
- Console: 18 Agent Studio tests (adapters plus the page with only HTTP mocked). The full suite is 46/46 suites and 1193 tests, and lint is clean.
- Backend: everything is green with the database attached, including the parallel run that clashed before per-package schemas.
- End-to-end, in a real browser, fully local.
- Setup: a throwaway Postgres; the backend running with no
.envand every host pinned to localhost; a second console on :5174; throwaway admin and manager logins. - Checked in the database: seed counts, the toggle, the threshold change, autonomy on (with confirmation) and off, and custom skill creation, each with its audit row.
- Checked in the browser: the manager view is read-only.
- Checked by direct API call: a manager's write gets 403.
- Everything was removed afterwards: container, image, scripts, test credentials and temporary launch entries.
- Setup: a throwaway Postgres; the backend running with no
Phase 3 — Land the ops-layer skills on main
- Port the 8 rule-based skills,
tools.js, the proposal executors and the banner from the branch onto current main. Resolve the 10 conflicts; the branch'sDeliveries.jsxedits collide with today's uncommitted change. SkillRegistrythen readsenabled/thresholdsfrom/admin/ai/skillsand falls back to code defaults when offline. It stops readinglocalStorage.- Keep the branch invariant: write tools return Proposals, and a human confirms.
- Done when: the branch's 20+ test files pass on main, and a threshold changed in Agent Studio changes the banner's output.
Phase 3 result (2026-09-29, uncommitted, NOT deployed):
- Ported, as unstaged file copies (no merge).
- The 8 skill definitions.
agent/{AgentFactory,signals,normalise,briefing,actions}.js.SlaRemediationCard.AgentOperationsBanner, mounted on the Exceptions page.- The
opsBriefingchat intent inlib/assistant/intents.js. - The "Needs attention" chip on the Exceptions context.
- 11 branch test suites.
- Settings.
SkillRegistrynow takes enabled and thresholds from/admin/ai/skillsthroughuseSkillRegistrySync, mounted once inAdminLayout. Nothing is kept in localStorage. It falls back to code defaults, and says so on the banner, when the registry cannot be read. - Not ported.
tools.js: nothing imported it.AgentStudioModal: a second, localStorage-only settings UI. "Configure skills" goes to Settings → Skills & Tools instead.AgentDecisionDrawerand the/internal/agent-decisionsendpoints: they return 403 from the console.- The AI-panel "Autonomous Fleet Agent" card, and the branch's cosmetic edits.
- Defects found and fixed while porting.
OpenToast('success', msg)has its arguments swapped; the signature is(message, variant). Every successful action would have shown a red error toast reading "success". The branch's tests asserted the same wrong order.- The
assignMilerexecutor posted to/hub/bookings/batch-assign, which is behindHubStaffAuth(role 6 only). Every console click would 403. It is now review-only, with the reason inactions.js. - Three skills could never fire. High-Value COD, Cash Exposure and
Battery Safety read payment and battery fields that
/admin/bookingsrows do not carry. They would report a false all-clear. They ship off (dataGapin code,Enabled: falseplus the reason in the seed). - The chat trigger was narrowed. It no longer claims "late/delayed orders" or "operations summary", which would have replaced existing answers. Routing tests pin both directions.
SlaRemediationCard.test.jsxcould not have run: nolucide-reactstub.
- Registry seed updated to match main.
CONSOLE_OPS_AGENTislive.- Paths point at main.
- The console tools are now the proposal verbs:
scan_bookings,notify_riders(the only executor),assign_riders(review-only), and five review-only actions, each labelled REVIEW ONLY. - New backend tests guard the no-data skills and the review-only labels.
- Tests.
- Console: 64/64 suites, 1308 tests; lint is clean and the build is green.
- Backend: build, vet and all tests are green. The Postgres-gated tests were not re-run; the seed change is data-only and unit-tested.
Deliveries.jsxwas untouched; the uncommitted change there is still only Suriya's.
- Open items.
- Feeding the three off skills needs
/admin/bookingsto include payment amounts and mode, and the rider's battery. That is a backend response-shape change. - "Assign riders" needs an admin batch-assign route.
useBatchAssignBookingshas the same 403 problem, but nothing calls it. lib/assistant/CLAUDE.mdis now stale: it says proactive alerts were "not started" and that thecomponents/assistantcopies are live, butAIPanelimportslib/assistant.
- Feeding the three off skills needs
Phase 4 — Persist decisions and runs (makes Insights real)
- A Go NATS consumer on
telemetry.taskplus the engine's LLM decisions, written toagent_decisions, plus a newai_agent_runstable (⚠ schema change). /admin/ai/decisionsand run stats feed the Insights tab and the Agents page, replacing theagentNetwork.jssnapshot.- Resolve before relying on vector search: the
context_embedding1536-vs-384 dimension question.
Phase 4 result (2026-09-29, uncommitted, NOT deployed):
- Finding. Nothing in doormile_backend or AI_engine writes
agent_decisions. routemate (external) returns anagent_decision_idfrom/decide-assignment, so it presumably writes throughPOST /internal/agent-decisions. Whether production has rows is unverified. AI_engine's two LLM decisions (stall, assignment-failure) are not persisted anywhere; they appear only in logs. - Backend.
- One new table,
aiagentruns: append-only, unique on (agentid, taskid), pruned after 30 days. internal/ai/telemetryqueue-subscribes (doormile-backend-telemetry) totelemetry.task, and writes runs in batches (2 s / 200). The NATS callback never blocks: a full buffer drops the event, counts it and logs it.telemetry.agentheartbeats go to Redis (ai:agent:state:<id>, 5-minute TTL), not Postgres.- New endpoints, Doormile staff only:
GET /admin/ai/insights?days=1..30(runs, failures and average time per agent; decisions by type and outcome; live state; areceivingflag) andGET /admin/ai/decisions(keyset-paged; thecontextcolumn is excluded because it holds rider data). - Windows use the backend clock (
utils.DBNow). The engine's naive timestamp is stored for display only.
- One new table,
- Console. Insights shows those figures with a 24 h / 7 d / 30 d window.
Silent agents appear as "silent" with zero runs rather than being left out.
When
receivingis false, the page says telemetry is not received rather than showing "0 runs" as if nothing happened. - Tests.
- Backend: 11 telemetry unit tests, a Postgres-gated suite, and the new routes in the gate tests.
- All 14 Postgres-gated tests (Phases 1 and 4) pass against a throwaway
postgres:16-alpine. This includes the parallelgo test ./...run. - Console: 64/64 suites, 1310 tests, lint clean, build green.
- Live end-to-end run (2026-09-29, local, throwaway; all removed after).
- Setup: throwaway Postgres and NATS; the real backend binary with no
.env; engine-shaped telemetry published over raw NATS. - Six messages produced three rows. The redelivered task was ignored by the unique index. The malformed event was dropped. The heartbeat was not written to Postgres.
/admin/ai/insightsreturned the right totals, failures and averages./admin/ai/decisionspaged correctly and did not leakcontext.- The Insights tab rendered the same figures with correct IST times.
- Setup: throwaway Postgres and NATS; the real backend binary with no
- Bugs the live run caught, fixed.
- The recorder stamped
utils.DBNow()into a timestamptz column, which AutoMigrate creates for new tables. A run received at 20:57 IST read back as 02:27 the next day. It now usestime.Now(), and the window cutoffs do too.TestRecorderStampsARealInstantfails with "5h30m off" if DBNow comes back. Wider note:utils.DBNowis only correct for the legacy timestamp-WITHOUT-zone columns. Any table AutoMigrate creates fresh is timestamptz, so audit other new tables before using DBNow in them. - Two Phase 1 Postgres fixtures still used
lookup_order/lookup_miler, which Phase 3 removed from the seed. - The Skills & Tools note still said the console skills were "not merged". It now says they run on these settings, and that AI_engine does not read them yet.
- The recorder stamped
- Deploy prerequisites.
- The backend's
NATS_URLmust point at the same NATS server AI_engine publishes to (NATS_HOST/NATS_PORTthere). Otherwise Insights shows "Not receiving agent telemetry". - The Agents page still uses the 16–20 Sep snapshot. Moving it onto
/admin/ai/insightsis a follow-up; it is dharaneesh's page. - To see AI_engine's LLM decisions in Insights, the engine would have to
POST them to
/internal/agent-decisions. That is engine work (Phase 5).
- The backend's
Phase 5 — Engine reads the registry
- The engine polls
GET /internal/ai/registry(ETag, around 30 s) and appliesenabled,autonomous,modeland thresholds without a restart. Today the autonomy flags are read once at import. - Fix before any agent is shown as live:
order_agent.py:182(the enum does not exist)- the JARVIS→ORDER payload key (
agent.py:279vsorder_agent.py:144) - the DISPATCH→HUB id mismatch (
dispatch_agent.py:268vshub_agent.py:151) release_vehicle_for_cancelhas no handler (exception_agent.py:692)- messages without a
task_typeare silently dropped - ORDER_AGENT still calls the renamed
crmbookingroute, with no auth header
- Mark HUB, FLEET and ROUTE_OPTIMIZER as
simulationin the seed, or retire them. - Add
pytestandpytest-asynciotorequirements.txt.
Phase 5 — result (2026-09-29, uncommitted, not deployed)
-
Registry client (
AI_engine/core/registry.py).- Polls
/internal/ai/registryeveryREGISTRY_POLL_SECONDS(30 by default) withIf-None-Match. It is started frommain.py --production. - Precedence: once the registry has loaded, its value applies. Before that, or if it never loads, the old env default applies.
- The last good copy survives 401/5xx/timeouts, so autonomy cannot flip mid-shift.
- With no
INTERNAL_API_KEY, it logs once and the engine runs on env defaults.
- Polls
-
What each agent now reads.
Agent Skill gate Autonomy Other settings Exception stall_responseEXCEPTION_AGENTstallMinutes,reassignConfidence, modelDispatch assignment_failure_triageDISPATCH_AGENTrealertEvery, modelExpress Dispatch express_batch_dispatchEXPRESS_DISPATCH_AGENTmaxPerRider,maxRadiusKm,loadPenaltyKmCustomer customer_notifications— — A disabled skill means the agent logs the event and does nothing.
-
Model override.
core/llm.request_params(model)uses the agent's pinned model, orLLM_MODELwhen none is pinned.- For a Haiku pin, thinking and effort are left out, because Haiku rejects them (400).
- The console picker offers Opus 5.5, Sonnet 5.5 and Opus 4.8, plus "Engine default". Haiku is left out because it is too weak for these decisions.
-
Decisions logged.
- Stall and assignment-failure decisions are POSTed to
/internal/agent-decisionsas{decision_type, booking_id, context:{facts, model}, decision:{action, confidence}, reasoning}. - The post is fire-and-forget, so a slow backend never delays the reaction.
- These decisions now appear in Insights → Latest decisions.
- Stall and assignment-failure decisions are POSTed to
-
Behaviour fix. When the LLM is down and the agent is not autonomous, it now escalates to a human. Before, it reassigned regardless of the autonomy flag.
-
Message bugs fixed.
ORDER_STATUS_UPDATEwas added to the enum.- JARVIS→ORDER now sends
order_id. - DISPATCH no longer forwards to HUB
prepare_receiving. It sent a booking id to a fictional-hub simulation. - FLEET handles
release_vehicle_for_cancel, finding the vehicle by order id. - CUSTOMER records
ORDER_CANCELLEDandNOTIFICATION_SENTinstead of dropping them. These come from simulated records, so no real customer message is sent. - ORDER_AGENT refuses its backend calls and logs why. The seed now says it
is not connected. It stays
broken.
-
Simulation agents. HUB, FLEET and ROUTE_OPTIMIZER stay seeded as
simulation, and the Studio note says so. -
Tests.
- AI_engine: 95 unittest tests. 28 are new, in
tests/test_registry_phase5.py, all with no network. Threetest_dispatch_agentmocks were updated for themodelargument. - The only failures are the two modules that import pytest, and they failed
before this work.
pytestandpytest-asyncioare now inrequirements.txtbut are not installed in the venv. - Console: Agent Studio suite green. Backend:
internal/ai/...green.
- AI_engine: 95 unittest tests. 28 are new, in
-
Deploy prerequisites.
- The engine needs
GO_API_BASE_URLandINTERNAL_API_KEY, the same key the backend checks. - Before deploying, check the registry's current
enabledandautonomousvalues. On deploy they replace the env flags (AUTONOMOUS_REASSIGNand the others).
- The engine needs
Phase 6 — Real Test playground
- Replace the
setTimeoutsimulation with a backend endpoint that runs one prompt through Claude tool-use, using the tools from the selected skill (theirinput_schemafrom the registry). - Read tools execute. Write tools return a Proposal only — the playground never mutates production.
- Show the real trace: tool calls, arguments, results, latency and tokens. The model comes from the registry.
Phase 6 — result (2026-09-29, uncommitted, not deployed)
- Decisions: use the official Go SDK, and redact personal data before
anything reaches Claude. The SDK download was blocked by this machine's
permission check. Everything else is built behind a
Modelinterface. The one missing piece is the ~80-line adapter fromplayground.Requesttoanthropic.MessageNewParams. It needsgo get github.com/anthropics/anthropic-sdk-go, run or approved by a person. Until then,controllers.PlaygroundModelis nil. The endpoint answers 503PLAYGROUND_NOT_CONFIGUREDand the Test tab says so; it never pretends to run. - Backend (
internal/ai/playground,controllers/aiPlaygroundController.go).-
POST /admin/ai/playground/run {agentid, skillid?, prompt}, for Doormile staff with roleid 1 only. The limit is 10 runs per user per 10 minutes, prompts are at most 2000 characters, and a run times out after 120 s. -
The loop runs at most 6 turns with max_tokens 4096. The model comes from the agent's registry pin, or
claude-opus-5-5if none. The skill's tools come from the registry, withinput_schemataken frominputschema. -
Tool outcomes:
Tool Outcome read, served by the backend ( get_booking_cache,scan_bookings,nearby_milers)executed, 5 s timeout read, engine-only or external ( decide_*,sequence_stops,simulate_pricing_quote,list_express_*)unavailable write / notify / event proposed. Never executed; the model gets {executed:false, proposal}not in the selected skill rejected -
Redaction:
- Executors select named non-personal columns only. There is no address, name, phone or notes column, and coordinates are rounded to 2 dp.
Redactthen masks personal keys (name, phone, address, email, note, reason, …) plus any email or Indian mobile number found in a string.- Results are capped at 16 KB.
-
The seed now gives
nearby_milers(lat, lon, radius_km) andscan_bookings(status, limit) real input schemas.
-
- Console:
- The Test tab calls the endpoint and shows the server's trace: each tool call with its input, outcome, result and ms, plus the model, turns, tokens and time.
- It has a skill picker, and "Test in Playground" preselects that skill. Non-admins can't run it.
- 503, 429 and 403 errors each show a plain message.
- The setTimeout simulation is gone.
- Tests:
- Backend: 11 unit tests. They cover Prepare, every outcome, the turn limit, model errors, truncation and redaction, including that dates are not masked.
- A Postgres-gated test proves that personal columns in the table are never
selected. It passed on a throwaway
postgres:16-alpine, which was removed afterwards. - Route tests: the gates, the 503 with no client, and bad input refused before the model is called.
- Console: 5 new tests. Now 64/64 suites and 1315 tests; build green.
- To switch it on:
- Add the SDK and the adapter.
- Set
ANTHROPIC_API_KEYon the backend. - Wire
controllers.PlaygroundModelinmain.go.
- Earlier notes:
- Official Go SDK or raw HTTP. The Go SDK
(
github.com/anthropics/anthropic-sdk-go) is not in the module cache, so using it means a module download plus a newgo.moddependency. - Data leaving for the Claude API. The read tools (
scan_bookings,get_booking_cache,nearby_milers,list_express_*) return live customer names, phones and addresses. Running them in the playground sends that data to Anthropic. The options are: allow it; redact PII before it is sent; or run against fixtures only.
- The backend also needs
ANTHROPIC_API_KEY. Without it, the endpoint returns 503 and the Test tab stays a labelled simulation.
- Official Go SDK or raw HTTP. The Go SDK
(
6. Decisions
Ratified 2026-09-29: A = port, C = roleid 1 only. The Agents page is kept and labelled. The secret check fails at boot in production only. B, D and E follow the recommendations below unless changed.
| # | Decision | My recommendation |
|---|---|---|
| A | feat/agentic-ops-layer: merge or port? |
Port the skills, tools and executors onto current main (Phase 3). Don't merge a 29-commit-stale branch with 10 conflicts. First confirm with dharaneesh that nothing newer exists elsewhere |
| B | Where does the registry live? | doormile_backend/Postgres. The engine has no API or persistence, and the console must not be the source of truth |
| C | Who may edit skills and autonomy? | roleid 1 only. Roles 3 and 4 read only. Today 1, 3 and 4 are identical everywhere, so this needs an explicit check |
| D | May the console toggle agent autonomy (auto-reassign riders, auto-notify customers)? | Yes, but only with a typed confirmation and an audit row. It stays off by default, matching compose |
| E | Delete or relabel the simulation agents (HUB, FLEET, ROUTE_OPTIMIZER)? | Seed them as simulation. Delete later if nobody objects |
7. Open questions (need checking, not guessing)
- Is
booking.assignment_failedpublished? The engine's handoff doc says Go doesn't publish it yet. Backend CLAUDE.md §4 namespublishAssignmentFailed, and the new retry window saysassignment_failedfires after the first round. Check which stream and subject it actually uses against what DISPATCH_AGENT binds to. - Which model is live? The engine defaults to
LLM_MODEL=claude-opus-4-8; the console mock showsclaude-3-5-sonnet. The registry'smodelfield should hold the id that is actually deployed. Set per agent, a cheaper model such as Haiku is enough for the stall and assignment decisions. - Registry write safety: admin handlers today are not tenant-guarded for global data (pricing, hubs, app users). The registry endpoints must not copy that pattern.