Files
doormilxpress_astryx/docs/agent-platform-plan.md

42 KiB
Raw Permalink Blame History

Doormile Agent Platform — Plan & Agent Registry

Status: Phases 0–4 done (uncommitted, not deployed), Phase 5 next · Written 2026-09-29 · Scope: krow_talent_app (the Doormile console), doormile_backend (Go API), AI_engine (Python agent swarm).

This is the document src/pages/doormile/settings/Settings.jsx has been pointing at ("Phase 6 of docs/agent-platform-plan.md"). That file never existed until now, so the phase numbers below replace the ones the comment assumed.


1. Where things actually stand (verified 2026-09-29)

Console — krow_talent_app

  • Settings → Skills & Tools (Agent Studio) is a mock. agentRegistryData.js holds 2 agents, 3 surfaces, 4 skills and 6 tools. Everything is saved to localStorage and nowhere else. The Test tab (AgentPlayground.jsx:37) is a setTimeout that always replies "succeeded with 100% confidence". Configure and Insights use hard-coded figures ("99.4%", "540 calls") and a model claude-3-5-sonnet that nothing reads.
  • Three of the six mock tools are Krow leftovers: open_add_skill_training, open_add_training and query_learning_analytics (workforce training, not logistics). Delete them; don't migrate them.
  • The Agents page (/doormile/agents) runs on a static snapshot in src/lib/agentNetwork.js, dated 16–20 Sep 2026. It is honest about being a snapshot, but its test (tests/integration/agentsPage.test.jsx) is stale: that is the 1 failing suite and 9 failing tests out of 45 suites.
  • Unmerged branch feat/agentic-ops-layer (dharaneesh, 2026-09-01) is the only real client-side agent work:
    • SkillRegistry with 8 rule-based skills: SlaGuardian, DoorstepStall, FleetBalancer, HighValueCod, RiderBatterySafety, HubCongestion, LateDispatch, CashExposure.
    • tools.js with 5 tools, following the rule "read-only tools execute, mutating tools return a Proposal".
    • AgentFactory, the ops briefing, proposal executors with a verify pass, and 20+ test files.
    • It is 29 commits behind main and conflicts in 10 files. Its skill config is also stored in localStorage.
  • The Home assistant runs on the regex catalogue src/lib/assistant/intents.js (109 KB) and has no concept of skills.
  • Other state: lint shows 35 auto-fixable errors, and Deliveries.jsx has an uncommitted cosmetic change (a tidy-up of the batch filter).

Backend — doormile_backend

  • Build, vet and test all pass (Go 1.26.4). There are 233 routes; CLAUDE.md is stale on counts, retry, routing, PIN auth and SMS (see its drift list).
  • Nothing for an agent registry exists yet. internal/ai/ is an empty, untracked directory.
  • What does exist:
    • models.AgentDecision (table agent_decisions, pgvector 1536-dim — the 384-vs-1536 question is still open).
    • Internal routes, all behind X-Internal-Key:
      • POST /internal/agent-decisions
      • GET /internal/agent-decisions/similar
      • PATCH /internal/agent-decisions/:id/outcome
      • /internal/express/{riders,bookings,assign}
  • Admin authorization is flat. Roles 1, 3 and 4 get identical access, and no handler checks the role further. Many admin write handlers have no tenant guard. So any registry write endpoint must add its own role check (see §5).

Engine — AI_engine

  • It has no registry to read from. Agents are hard-coded classes in main.py:55-70. core/tool_registry.py is an in-memory dict holding 3 tools, and none of them load in production. The two "skills" in core/skills/ are stubs that nothing imports.
  • The agent process has no HTTP API. The Command Center (FastAPI, :8600) is a separate, unauthenticated NATS tap.
  • Runs, escalations and decisions live in memory only. They are capped at 500 and lost on restart.
  • Tests: 72 pass, but only with pytest installed by hand; requirements.txt lacks pytest.

Conclusion: the registry must be built, not wired up. It belongs in doormile_backend (Postgres). The console reads it, and the engine reads it. Neither of them owns it.


2. The Agent Registry — what gets seeded

This is the real inventory. Every row here becomes a seed row in Phase 1. "Status" is what the code does today, not what the docs claim.

2.1 Agents

id Class / file Purpose Trigger Writes to system LLM Status
JARVIS MasterAgent · core/agent.py:196 Orchestrator, escalation inbox logistics.direct.JARVIS none — Partial. The inbox works; orchestrate_order sends the wrong payload key
DISPATCH_AGENT agents/dispatch_agent.py:72 Watches assignment outcomes, flags coverage gaps JetStream booking.assigned, booking.assignment_failed Alerts; customer notify only when autonomous decide_assignment_failure Production-grade. Idle until assignment_failed is confirmed as published (see §7)
EXCEPTION_AGENT agents/exception_agent.py:106 Stalled-rider detection and response TRACKING miler.location.updated, miler.stalled, plus a 60 s DB sweep POST /internal/bookings/:id/reassign (needs autonomy flag and confidence ≥ 0.7), /internal/notify decide_stall_response Production-grade (the stall path only)
EXPRESS_DISPATCH_AGENT agents/express_dispatch_agent.py:80 Tenant batch assign plus road sequencing JetStream express.dispatch_requested POST /internal/express/assign — (greedy) Implemented. Autonomy defaults to true in code and false in compose
CUSTOMER_AGENT agents/customer_agent.py:50 Customer notifications Direct tasks, only from an autonomous Dispatch /internal/notify — Implemented, rarely reached
ORDER_AGENT agents/order_agent.py:15 Order intake and validation Direct tasks (no sender in production) /admin/crmbooking — a dead path, no auth header — Broken. update_status crashes (order_agent.py:182)
HUB_AGENT agents/hub_agent.py:35 Hub capacity Direct tasks none — Simulation (8 fictional hubs)
FLEET_AGENT agents/fleet_agent.py:34 Vehicles Direct tasks none — Simulation (19 fake vehicles)
ROUTE_OPTIMIZER agents/route_optimizer_agent.py:41 Routing Direct tasks none — Simulation (haversine only)
CONSOLE_ASSISTANT src/lib/assistant/* (console) Home chat: orders, bulk, assign, repeat Operator prompt Via proposals the operator confirms RAG sidecar (off) Live, regex-based
CONSOLE_OPS_AGENT feat/agentic-ops-layer Ops briefing plus 8 monitoring skills Page load / poll Via proposals the operator confirms — (rules) Unmerged

The registry shows status as one of live | partial | simulation | broken | unmerged | retired. The console must display this badge. A simulation agent that looks live is the exact failure the Agents page comment warns about.

2.2 Tools (the real capabilities, not the stubs)

kind is read, write or notify. Every write or notify tool has requires_confirmation = true unless its agent is explicitly autonomous.

name Owner agent(s) Kind Target Today lives at
reassign_booking EXCEPTION write Go /internal/bookings/:id/reassign exception_agent.py:466
notify_customer EXCEPTION, CUSTOMER notify Go /internal/notify exception_agent.py:476, customer_agent.py:349
list_express_bookings EXPRESS read Go /internal/express/bookings express_dispatch_agent.py:351
list_express_riders EXPRESS read Go /internal/express/riders express_dispatch_agent.py:342
assign_express_batch EXPRESS write Go /internal/express/assign express_dispatch_agent.py:222
sequence_stops EXPRESS read (compute) routes.workolik.com /optimization/doormile/sequence express_dispatch_agent.py:308
get_booking_cache CUSTOMER read Go /bookings/cache/:id customer_agent.py:243
nearby_milers DISPATCH, EXCEPTION read Redis GEO milers:locations dispatch_agent.py:170-240
publish_miler_stalled EXCEPTION write (event) NATS miler.stalled exception_agent.py:377
decide_stall_response EXCEPTION read (LLM) Claude core/llm.py:157
decide_assignment_failure DISPATCH read (LLM) Claude core/llm.py:227
record_agent_decision all write Go /internal/agent-decisions Go route exists
diagnose_operations CONSOLE_OPS read Console queries branch tools.js
lookup_order CONSOLE_OPS read /admin/bookings branch tools.js
lookup_miler CONSOLE_OPS read /admin/milers branch tools.js
propose_miler_reassignment CONSOLE_OPS write → proposal /admin/bookings/:id/assign-miler branch tools.js
simulate_pricing_quote CONSOLE_OPS, CONSOLE_ASSISTANT read /admin/pricing/simulate branch tools.js
create_single_order CONSOLE_ASSISTANT write → proposal /admin/expressbooking current mock
rebalance_riders CONSOLE_OPS write → proposal /hub/bookings/batch-assign current mock (no backing code)

Not seeded:

  • ask_question, order_intake_skill, repeat_run_skill (engine stubs that are never loaded).
  • The 3 Krow training tools.
  • ORDER_AGENT's crmbooking calls: that route was renamed to expressbooking, so those calls hit nothing.

2.3 Skills

A skill is a named behaviour that belongs to one agent and uses one or more tools. It has a switch (enabled) and tunable thresholds (JSON, validated against a per-skill schema).

id Agent Tools Source
stall_response EXCEPTION nearby_milers, decide_stall_response, reassign_booking, notify_customer engine
assignment_failure_triage DISPATCH nearby_milers, decide_assignment_failure engine
express_batch_dispatch EXPRESS list_express_*, assign_express_batch, sequence_stops engine
sla_guardian · doorstep_stall · fleet_balancer · high_value_cod · rider_battery_safety · hub_congestion · late_dispatch · cash_exposure CONSOLE_OPS branch tools.js set branch (thresholds move from localStorage to the registry)
order_intake_auto_schedule CONSOLE_ASSISTANT create_single_order, simulate_pricing_quote current mock, real behaviour in orderFlow.js
dispatch_rebalance CONSOLE_OPS rebalance_riders current mock. Keep disabled until a backing endpoint exists

2.4 Surfaces

/doormile/home (assistant), /doormile/control-x (dispatch board), /doormile/agents (status board), and the ops banner (branch AgentOperationsBanner). /doormile/dispatch is only a redirect to Control X, so it is not a separate surface.


3. Schema (Phase 1) — ⚠ additive schema change, review before deploy

New tables in doormile_backend, added through GORM AutoMigrate. They are additive only and touch no existing table.

ai_agents        id text PK, name, class_ref, runtime ('engine'|'console'),
                 purpose, trigger jsonb, status, autonomous bool,
                 model text NULL, created_at, updated_at
ai_tools         name text PK, description, kind ('read'|'write'|'notify'),
                 target text, input_schema jsonb, requires_confirmation bool,
                 enabled bool, created_at, updated_at
ai_skills        id text PK, agent_id FK→ai_agents, title, category,
                 description, sample_prompt, enabled bool,
                 thresholds jsonb, thresholds_schema jsonb,
                 version int, updated_by int NULL, updated_at
ai_skill_tools   skill_id FK, tool_name FK, PRIMARY KEY (skill_id, tool_name)
ai_registry_audit id bigserial, entity, entity_id, field, old jsonb, new jsonb,
                 changed_by int, changed_at

Rules:

  • Tools and agents are code-defined. The console can toggle enabled and autonomous, but it cannot invent a tool. A tool with no implementation is a lie on screen. "New skill" in the console therefore picks from existing tools only.
  • Every write goes into ai_registry_audit in the same transaction.
  • Seeding is idempotent (upsert by id) and lives in migrations/, not scratch/.
  • Never store a secret. Env var names only; the engine's .env stays out of the registry.

4. API contract (Phase 1)

All handlers go in controllers/aiRegistryController.go and the logic in internal/ai/registry. Responses use utils.OK / utils.List, following the console conventions.

Method Path Auth Notes
GET /admin/ai/agents admin (1,3,4), Doormile staff only Includes a status badge and skill/tool counts
GET /admin/ai/agents/:id same Agent plus its skills and tools
GET /admin/ai/skills same ?agent= filter
GET /admin/ai/tools same ?kind= filter
PATCH /admin/ai/skills/:id roleid 1 only Allowed fields: enabled, thresholds (validated against schema). Bumps version and writes an audit row
POST /admin/ai/skills roleid 1 only New skill from existing tools only
PATCH /admin/ai/agents/:id roleid 1 only Allowed fields: autonomous, model. Extra confirm in UI — this changes what an agent does without a human
GET /admin/ai/audit admin Registry change history
GET /admin/ai/decisions admin Paged agent_decisions (Phase 4)
GET /internal/ai/registry X-Internal-Key The engine pulls its config from here (Phase 5). It sends an ETag so the engine can poll cheaply

"Doormile staff only" means consoleTenantID == 0. A tenant or client login gets 403. It does not get an empty list, because an empty list would hide a misconfiguration.


5. Phases

Each phase ships on its own and leaves the system working. Nothing is committed or pushed without an explicit ask.

Phase 0 — Prerequisites and hygiene (small, do first)

  1. Rotate secrets and untrack them.
    • doormile_backend: .env and doormile-abee7-*.json are tracked; the .gitignore line #.env is commented out.
    • AI_engine: .env is tracked.
    • config/config.go:60-71 has hard-coded fallback secrets. Fail at boot instead.
    • This is a user action (rotation needs the providers' consoles). I can do the untracking and ignore rules.
  2. Decide feat/agentic-ops-layer (see §6, decision A).
  3. Console: fix the stale agentsPage.test.jsx, run lint:fix for the 35 unused imports, and delete the 3 Krow training tools from agentRegistryData.js.
  4. Add a PREVIEW banner on the Agent Studio tab itself. Today only a code comment says so; Agents shows a snapshot label, Agent Studio shows nothing.
  5. Backend: remove the broken .claude/skills/* symlink stubs, .agents/, skills.md and skills-lock.json. The global plugin already provides these skills.

Done when: tests are green, lint is clean, and no secret is in git ls-files.

Phase 0 result (2026-09-29, uncommitted):

  • Console: lint clean; 46/46 suites, 1192 tests pass.
  • Agent Studio: 4 Krow tools and 2 Krow skills removed (the plan said 3 tools; open_miletruth_ai was a fourth). dispatch_rebalance ships disabled. Storage keys moved to _v3. An on-screen Preview note was added.
  • Agents page: this was not just a stale test. The 24–25 Sep rebuild presented a simulation as live. Per decision, the design was kept and labelled:
    • a Snapshot · 16–20 Sep 2026 stamp;
    • "Sample Activity — Simulation · not live data";
    • no pulsing dot;
    • status counts taken from networkStats();
    • "Autonomy gates on: 0 / 3", where the old "0 / 8" implied 8 gates. The tests were rewritten, keeping the honesty checks.
  • Backend:
    • Reverted 2026-09-29 at Suriya's request: these are back in git exactly as at HEAD:
      • the untracking of .env and the service-account key (the key is still committed, so rotation still stands);
      • the .gitignore edit;
      • the removal of skills.md, skills-lock.json, .agents/ and .claude/skills/. AI_engine's .env is tracked again too. Do not redo any of this without asking.
    • main.go now refuses to boot when ENV=production and JWT_SECRET_KEY, DB_PASSWORD or NATS_PASSWORD is unset.
    • build, vet and test are green.
  • AI_engine: .env was untracked (it was already in .gitignore).
  • Still yours:
    • Rotate every secret that was committed. Git history still holds the values.
    • Before the next backend deploy, confirm that production sets all three secrets. If it has been running on a fallback, the new check stops it from starting.

Phase 1 — Registry in the backend

  • Add the §3 tables, the idempotent seed from §2, and the §4 read endpoints plus PATCH/POST with the role check and audit trail.
  • Tests: a seed-idempotency test, a role test (roles 3 and 4 get 403 on PATCH; a tenant login gets 403 on GET), a threshold-schema validation test, and an audit test.
  • Done when: go build/vet/test is green and curl against a staging DB returns the §2 inventory.

Phase 1 result (2026-09-29, uncommitted, NOT deployed, no real DB touched):

  • Tables. They follow the codebase's naming, not the names in §3: aiagents, aitools, aiskills, aiskilltools, airegistryaudit. The columns are as in §3, with a few changes:
    • The agent's trigger column is named wakeon.
    • hasautonomygate is new. Autonomy can only be set on Dispatch, Exception and Express.
    • source (engine/console/custom) is on skills.
    • There is no enabled flag on tools.
  • Code.
    • internal/ai/registry: seed, rules, store.
    • controllers/aiRegistryController.go
    • middlewares/staff_only.go (DoormileStaffOnly)
    • Routes are under /admin/ai/* and /internal/ai/registry, as in §4.
    • The seed runs in migrations.Migrate. It logs a failure and does not stop the boot.
  • Seed. 11 agents, 26 tools and 15 skills.
    • Four skills were added to §2.3 so that every tool belongs to a skill: customer_notifications, ops_briefing, and the 8 branch tools folded into their skills.
    • The console ops skills keep the branch ids and threshold keys, so Phase 3 is a 1:1 mapping.
    • record_agent_decision was dropped. It is a log the Go side writes, not a capability.
  • Rules enforced server-side.
    • Switching autonomy ON needs confirm set to the agent id.
    • Model ids must match claude-*, and only engine agents have one.
    • Thresholds are validated for range and step, and a patch is all-or-nothing.
    • Custom skills can be added to console agents only, and only from existing tools.
    • A patch that changes nothing does not bump the version or write an audit row.
  • Tests.
    • 26 unit tests and 6 HTTP gate tests always run.
    • 10 Postgres integration tests and 1 HTTP end-to-end test run only when REGISTRY_TEST_DSN is set. They need a throwaway database; each package uses its own schema.
    • All of them passed against a disposable postgres:16-alpine container.
  • Bug the Postgres run caught. With enabled tagged default:true, gorm dropped false from the INSERT, so dispatch_rebalance came up enabled. Fixed by removing the column default.
  • Not yet proven. The migration has not run against the real database; that happens on your next deploy. It is 5 new tables and touches nothing existing.

Phase 2 — Console Agent Studio reads the registry

  • Replace agentRegistryData.js and its localStorage with React Query hooks (useAiAgents, useAiSkills, useAiTools) in src/lib/doormileHooks.js, and add the endpoints to src/api/doormile/endpoints.js.
  • Keep the existing components; only the data source changes. Also show the status badge, the tool kind, and a confirmation marker.
  • The skill toggle and "New skill" become real PATCH and POST calls. Add a loading/error state; remove the optimistic toast that claims success before the server answers.
  • Configure tab: model and autonomy from the registry. The temperature slider is dropped, because nothing reads it.
  • Insights and Test stay behind the PREVIEW banner until Phases 4 and 6.
  • Rewrite tests/integration/agentStudio.test.jsx against mocked hooks.
  • Done when: a toggle made in one browser shows in another, and survives a reload.

Phase 2 result (2026-09-29, uncommitted, NOT deployed):

  • Console.
    • Endpoints were added to api/doormile/endpoints.js: getAiAgents, getAiSkills, getAiTools, updateAiSkill, createAiSkill, updateAiAgent and getAiRegistryAudit.
    • Hooks were added in lib/doormileHooks.js: useAiAgents, useAiSkills, useAiTools and three mutations. They share one ['doormile','ai'] key.
    • The adapter layer is agentStudio/registryAdapters.js, a set of pure functions.
    • agentRegistryData.js now keeps only the surfaces list and the selected-agent preference.
  • Changes on screen.
    • Skills are filtered to the selected agent; before, every agent's skills showed.
    • Agent status badges appear in the switcher.
    • The skill drawer has a real enable toggle, a threshold editor (range and step checked in the browser, then on the server), and shows source and version.
    • The tool table shows the kind, "Used by", the system each tool touches and where it is implemented. It no longer calls a read-only tool "Autonomous".
    • Configure shows the agent's record, a model picker (engine agents only) and an autonomy switch (gated agents only) with a typed confirmation. The fake temperature slider and GPT/DeepSeek list are gone.
    • Insights shows "No run data yet" instead of invented figures.
    • "New skill" is disabled on AI_engine agents and for anyone who is not an admin.
    • With no saved choice the page opens on the first agent that has skills, not on JARVIS, which has none.
  • Backend addition. airegistryaudit.changedbyemail records the token's email. The end-to-end run showed every audit row with changedby = 0: an admin login without an appusers row carries user id 0.
  • Tests.
    • Console: 18 Agent Studio tests (adapters plus the page with only HTTP mocked). The full suite is 46/46 suites and 1193 tests, and lint is clean.
    • Backend: everything is green with the database attached, including the parallel run that clashed before per-package schemas.
  • End-to-end, in a real browser, fully local.
    • Setup: a throwaway Postgres; the backend running with no .env and every host pinned to localhost; a second console on :5174; throwaway admin and manager logins.
    • Checked in the database: seed counts, the toggle, the threshold change, autonomy on (with confirmation) and off, and custom skill creation, each with its audit row.
    • Checked in the browser: the manager view is read-only.
    • Checked by direct API call: a manager's write gets 403.
    • Everything was removed afterwards: container, image, scripts, test credentials and temporary launch entries.

Phase 3 — Land the ops-layer skills on main

  • Port the 8 rule-based skills, tools.js, the proposal executors and the banner from the branch onto current main. Resolve the 10 conflicts; the branch's Deliveries.jsx edits collide with today's uncommitted change.
  • SkillRegistry then reads enabled/thresholds from /admin/ai/skills and falls back to code defaults when offline. It stops reading localStorage.
  • Keep the branch invariant: write tools return Proposals, and a human confirms.
  • Done when: the branch's 20+ test files pass on main, and a threshold changed in Agent Studio changes the banner's output.

Phase 3 result (2026-09-29, uncommitted, NOT deployed):

  • Ported, as unstaged file copies (no merge).
    • The 8 skill definitions.
    • agent/{AgentFactory,signals,normalise,briefing,actions}.js.
    • SlaRemediationCard.
    • AgentOperationsBanner, mounted on the Exceptions page.
    • The opsBriefing chat intent in lib/assistant/intents.js.
    • The "Needs attention" chip on the Exceptions context.
    • 11 branch test suites.
  • Settings. SkillRegistry now takes enabled and thresholds from /admin/ai/skills through useSkillRegistrySync, mounted once in AdminLayout. Nothing is kept in localStorage. It falls back to code defaults, and says so on the banner, when the registry cannot be read.
  • Not ported.
    • tools.js: nothing imported it.
    • AgentStudioModal: a second, localStorage-only settings UI. "Configure skills" goes to Settings → Skills & Tools instead.
    • AgentDecisionDrawer and the /internal/agent-decisions endpoints: they return 403 from the console.
    • The AI-panel "Autonomous Fleet Agent" card, and the branch's cosmetic edits.
  • Defects found and fixed while porting.
    1. OpenToast('success', msg) has its arguments swapped; the signature is (message, variant). Every successful action would have shown a red error toast reading "success". The branch's tests asserted the same wrong order.
    2. The assignMiler executor posted to /hub/bookings/batch-assign, which is behind HubStaffAuth (role 6 only). Every console click would 403. It is now review-only, with the reason in actions.js.
    3. Three skills could never fire. High-Value COD, Cash Exposure and Battery Safety read payment and battery fields that /admin/bookings rows do not carry. They would report a false all-clear. They ship off (dataGap in code, Enabled: false plus the reason in the seed).
    4. The chat trigger was narrowed. It no longer claims "late/delayed orders" or "operations summary", which would have replaced existing answers. Routing tests pin both directions.
    5. SlaRemediationCard.test.jsx could not have run: no lucide-react stub.
  • Registry seed updated to match main.
    • CONSOLE_OPS_AGENT is live.
    • Paths point at main.
    • The console tools are now the proposal verbs: scan_bookings, notify_riders (the only executor), assign_riders (review-only), and five review-only actions, each labelled REVIEW ONLY.
    • New backend tests guard the no-data skills and the review-only labels.
  • Tests.
    • Console: 64/64 suites, 1308 tests; lint is clean and the build is green.
    • Backend: build, vet and all tests are green. The Postgres-gated tests were not re-run; the seed change is data-only and unit-tested.
    • Deliveries.jsx was untouched; the uncommitted change there is still only Suriya's.
  • Open items.
    • Feeding the three off skills needs /admin/bookings to include payment amounts and mode, and the rider's battery. That is a backend response-shape change.
    • "Assign riders" needs an admin batch-assign route. useBatchAssignBookings has the same 403 problem, but nothing calls it.
    • lib/assistant/CLAUDE.md is now stale: it says proactive alerts were "not started" and that the components/assistant copies are live, but AIPanel imports lib/assistant.

Phase 4 — Persist decisions and runs (makes Insights real)

  • A Go NATS consumer on telemetry.task plus the engine's LLM decisions, written to agent_decisions, plus a new ai_agent_runs table (⚠ schema change).
  • /admin/ai/decisions and run stats feed the Insights tab and the Agents page, replacing the agentNetwork.js snapshot.
  • Resolve before relying on vector search: the context_embedding 1536-vs-384 dimension question.

Phase 4 result (2026-09-29, uncommitted, NOT deployed):

  • Finding. Nothing in doormile_backend or AI_engine writes agent_decisions. routemate (external) returns an agent_decision_id from /decide-assignment, so it presumably writes through POST /internal/agent-decisions. Whether production has rows is unverified. AI_engine's two LLM decisions (stall, assignment-failure) are not persisted anywhere; they appear only in logs.
  • Backend.
    • One new table, aiagentruns: append-only, unique on (agentid, taskid), pruned after 30 days.
    • internal/ai/telemetry queue-subscribes (doormile-backend-telemetry) to telemetry.task, and writes runs in batches (2 s / 200). The NATS callback never blocks: a full buffer drops the event, counts it and logs it.
    • telemetry.agent heartbeats go to Redis (ai:agent:state:<id>, 5-minute TTL), not Postgres.
    • New endpoints, Doormile staff only: GET /admin/ai/insights?days=1..30 (runs, failures and average time per agent; decisions by type and outcome; live state; a receiving flag) and GET /admin/ai/decisions (keyset-paged; the context column is excluded because it holds rider data).
    • Windows use the backend clock (utils.DBNow). The engine's naive timestamp is stored for display only.
  • Console. Insights shows those figures with a 24 h / 7 d / 30 d window. Silent agents appear as "silent" with zero runs rather than being left out. When receiving is false, the page says telemetry is not received rather than showing "0 runs" as if nothing happened.
  • Tests.
    • Backend: 11 telemetry unit tests, a Postgres-gated suite, and the new routes in the gate tests.
    • All 14 Postgres-gated tests (Phases 1 and 4) pass against a throwaway postgres:16-alpine. This includes the parallel go test ./... run.
    • Console: 64/64 suites, 1310 tests, lint clean, build green.
  • Live end-to-end run (2026-09-29, local, throwaway; all removed after).
    • Setup: throwaway Postgres and NATS; the real backend binary with no .env; engine-shaped telemetry published over raw NATS.
    • Six messages produced three rows. The redelivered task was ignored by the unique index. The malformed event was dropped. The heartbeat was not written to Postgres.
    • /admin/ai/insights returned the right totals, failures and averages. /admin/ai/decisions paged correctly and did not leak context.
    • The Insights tab rendered the same figures with correct IST times.
  • Bugs the live run caught, fixed.
    1. The recorder stamped utils.DBNow() into a timestamptz column, which AutoMigrate creates for new tables. A run received at 20:57 IST read back as 02:27 the next day. It now uses time.Now(), and the window cutoffs do too. TestRecorderStampsARealInstant fails with "5h30m off" if DBNow comes back. Wider note: utils.DBNow is only correct for the legacy timestamp-WITHOUT-zone columns. Any table AutoMigrate creates fresh is timestamptz, so audit other new tables before using DBNow in them.
    2. Two Phase 1 Postgres fixtures still used lookup_order/lookup_miler, which Phase 3 removed from the seed.
    3. The Skills & Tools note still said the console skills were "not merged". It now says they run on these settings, and that AI_engine does not read them yet.
  • Deploy prerequisites.
    1. The backend's NATS_URL must point at the same NATS server AI_engine publishes to (NATS_HOST/NATS_PORT there). Otherwise Insights shows "Not receiving agent telemetry".
    2. The Agents page still uses the 16–20 Sep snapshot. Moving it onto /admin/ai/insights is a follow-up; it is dharaneesh's page.
    3. To see AI_engine's LLM decisions in Insights, the engine would have to POST them to /internal/agent-decisions. That is engine work (Phase 5).

Phase 5 — Engine reads the registry

  • The engine polls GET /internal/ai/registry (ETag, around 30 s) and applies enabled, autonomous, model and thresholds without a restart. Today the autonomy flags are read once at import.
  • Fix before any agent is shown as live:
    • order_agent.py:182 (the enum does not exist)
    • the JARVIS→ORDER payload key (agent.py:279 vs order_agent.py:144)
    • the DISPATCH→HUB id mismatch (dispatch_agent.py:268 vs hub_agent.py:151)
    • release_vehicle_for_cancel has no handler (exception_agent.py:692)
    • messages without a task_type are silently dropped
    • ORDER_AGENT still calls the renamed crmbooking route, with no auth header
  • Mark HUB, FLEET and ROUTE_OPTIMIZER as simulation in the seed, or retire them.
  • Add pytest and pytest-asyncio to requirements.txt.

Phase 5 — result (2026-09-29, uncommitted, not deployed)

  • Registry client (AI_engine/core/registry.py).

    • Polls /internal/ai/registry every REGISTRY_POLL_SECONDS (30 by default) with If-None-Match. It is started from main.py --production.
    • Precedence: once the registry has loaded, its value applies. Before that, or if it never loads, the old env default applies.
    • The last good copy survives 401/5xx/timeouts, so autonomy cannot flip mid-shift.
    • With no INTERNAL_API_KEY, it logs once and the engine runs on env defaults.
  • What each agent now reads.

    Agent Skill gate Autonomy Other settings
    Exception stall_response EXCEPTION_AGENT stallMinutes, reassignConfidence, model
    Dispatch assignment_failure_triage DISPATCH_AGENT realertEvery, model
    Express Dispatch express_batch_dispatch EXPRESS_DISPATCH_AGENT maxPerRider, maxRadiusKm, loadPenaltyKm
    Customer customer_notifications — —

    A disabled skill means the agent logs the event and does nothing.

  • Model override.

    • core/llm.request_params(model) uses the agent's pinned model, or LLM_MODEL when none is pinned.
    • For a Haiku pin, thinking and effort are left out, because Haiku rejects them (400).
    • The console picker offers Opus 5.5, Sonnet 5.5 and Opus 4.8, plus "Engine default". Haiku is left out because it is too weak for these decisions.
  • Decisions logged.

    • Stall and assignment-failure decisions are POSTed to /internal/agent-decisions as {decision_type, booking_id, context:{facts, model}, decision:{action, confidence}, reasoning}.
    • The post is fire-and-forget, so a slow backend never delays the reaction.
    • These decisions now appear in Insights → Latest decisions.
  • Behaviour fix. When the LLM is down and the agent is not autonomous, it now escalates to a human. Before, it reassigned regardless of the autonomy flag.

  • Message bugs fixed.

    1. ORDER_STATUS_UPDATE was added to the enum.
    2. JARVIS→ORDER now sends order_id.
    3. DISPATCH no longer forwards to HUB prepare_receiving. It sent a booking id to a fictional-hub simulation.
    4. FLEET handles release_vehicle_for_cancel, finding the vehicle by order id.
    5. CUSTOMER records ORDER_CANCELLED and NOTIFICATION_SENT instead of dropping them. These come from simulated records, so no real customer message is sent.
    6. ORDER_AGENT refuses its backend calls and logs why. The seed now says it is not connected. It stays broken.
  • Simulation agents. HUB, FLEET and ROUTE_OPTIMIZER stay seeded as simulation, and the Studio note says so.

  • Tests.

    • AI_engine: 95 unittest tests. 28 are new, in tests/test_registry_phase5.py, all with no network. Three test_dispatch_agent mocks were updated for the model argument.
    • The only failures are the two modules that import pytest, and they failed before this work. pytest and pytest-asyncio are now in requirements.txt but are not installed in the venv.
    • Console: Agent Studio suite green. Backend: internal/ai/... green.
  • Deploy prerequisites.

    1. The engine needs GO_API_BASE_URL and INTERNAL_API_KEY, the same key the backend checks.
    2. Before deploying, check the registry's current enabled and autonomous values. On deploy they replace the env flags (AUTONOMOUS_REASSIGN and the others).

Phase 6 — Real Test playground

  • Replace the setTimeout simulation with a backend endpoint that runs one prompt through Claude tool-use, using the tools from the selected skill (their input_schema from the registry).
  • Read tools execute. Write tools return a Proposal only — the playground never mutates production.
  • Show the real trace: tool calls, arguments, results, latency and tokens. The model comes from the registry.

Phase 6 — result (2026-09-29, uncommitted, not deployed)

  • Decisions: use the official Go SDK, and redact personal data before anything reaches Claude. The SDK download was blocked by this machine's permission check. Everything else is built behind a Model interface. The one missing piece is the ~80-line adapter from playground.Request to anthropic.MessageNewParams. It needs go get github.com/anthropics/anthropic-sdk-go, run or approved by a person. Until then, controllers.PlaygroundModel is nil. The endpoint answers 503 PLAYGROUND_NOT_CONFIGURED and the Test tab says so; it never pretends to run.
  • Backend (internal/ai/playground, controllers/aiPlaygroundController.go).
    • POST /admin/ai/playground/run {agentid, skillid?, prompt}, for Doormile staff with roleid 1 only. The limit is 10 runs per user per 10 minutes, prompts are at most 2000 characters, and a run times out after 120 s.

    • The loop runs at most 6 turns with max_tokens 4096. The model comes from the agent's registry pin, or claude-opus-5-5 if none. The skill's tools come from the registry, with input_schema taken from inputschema.

    • Tool outcomes:

      Tool Outcome
      read, served by the backend (get_booking_cache, scan_bookings, nearby_milers) executed, 5 s timeout
      read, engine-only or external (decide_*, sequence_stops, simulate_pricing_quote, list_express_*) unavailable
      write / notify / event proposed. Never executed; the model gets {executed:false, proposal}
      not in the selected skill rejected
    • Redaction:

      • Executors select named non-personal columns only. There is no address, name, phone or notes column, and coordinates are rounded to 2 dp.
      • Redact then masks personal keys (name, phone, address, email, note, reason, …) plus any email or Indian mobile number found in a string.
      • Results are capped at 16 KB.
    • The seed now gives nearby_milers (lat, lon, radius_km) and scan_bookings (status, limit) real input schemas.

  • Console:
    • The Test tab calls the endpoint and shows the server's trace: each tool call with its input, outcome, result and ms, plus the model, turns, tokens and time.
    • It has a skill picker, and "Test in Playground" preselects that skill. Non-admins can't run it.
    • 503, 429 and 403 errors each show a plain message.
    • The setTimeout simulation is gone.
  • Tests:
    • Backend: 11 unit tests. They cover Prepare, every outcome, the turn limit, model errors, truncation and redaction, including that dates are not masked.
    • A Postgres-gated test proves that personal columns in the table are never selected. It passed on a throwaway postgres:16-alpine, which was removed afterwards.
    • Route tests: the gates, the 503 with no client, and bad input refused before the model is called.
    • Console: 5 new tests. Now 64/64 suites and 1315 tests; build green.
  • To switch it on:
    1. Add the SDK and the adapter.
    2. Set ANTHROPIC_API_KEY on the backend.
    3. Wire controllers.PlaygroundModel in main.go.
  • Earlier notes:
    1. Official Go SDK or raw HTTP. The Go SDK (github.com/anthropics/anthropic-sdk-go) is not in the module cache, so using it means a module download plus a new go.mod dependency.
    2. Data leaving for the Claude API. The read tools (scan_bookings, get_booking_cache, nearby_milers, list_express_*) return live customer names, phones and addresses. Running them in the playground sends that data to Anthropic. The options are: allow it; redact PII before it is sent; or run against fixtures only.
    • The backend also needs ANTHROPIC_API_KEY. Without it, the endpoint returns 503 and the Test tab stays a labelled simulation.

6. Decisions

Ratified 2026-09-29: A = port, C = roleid 1 only. The Agents page is kept and labelled. The secret check fails at boot in production only. B, D and E follow the recommendations below unless changed.

# Decision My recommendation
A feat/agentic-ops-layer: merge or port? Port the skills, tools and executors onto current main (Phase 3). Don't merge a 29-commit-stale branch with 10 conflicts. First confirm with dharaneesh that nothing newer exists elsewhere
B Where does the registry live? doormile_backend/Postgres. The engine has no API or persistence, and the console must not be the source of truth
C Who may edit skills and autonomy? roleid 1 only. Roles 3 and 4 read only. Today 1, 3 and 4 are identical everywhere, so this needs an explicit check
D May the console toggle agent autonomy (auto-reassign riders, auto-notify customers)? Yes, but only with a typed confirmation and an audit row. It stays off by default, matching compose
E Delete or relabel the simulation agents (HUB, FLEET, ROUTE_OPTIMIZER)? Seed them as simulation. Delete later if nobody objects

7. Open questions (need checking, not guessing)

  • Is booking.assignment_failed published? The engine's handoff doc says Go doesn't publish it yet. Backend CLAUDE.md §4 names publishAssignmentFailed, and the new retry window says assignment_failed fires after the first round. Check which stream and subject it actually uses against what DISPATCH_AGENT binds to.
  • Which model is live? The engine defaults to LLM_MODEL=claude-opus-4-8; the console mock shows claude-3-5-sonnet. The registry's model field should hold the id that is actually deployed. Set per agent, a cheaper model such as Haiku is enough for the stall and assignment decisions.
  • Registry write safety: admin handlers today are not tenant-guarded for global data (pricing, hubs, app users). The registry endpoints must not copy that pattern.