Files
doormilxpress_astryx/docs/agent-platform-phase7-plan.md

38 KiB
Raw Blame History

Doormile Agent Platform — Phase 7: Decision Memory, Deployment, Executors

Status: mostly implemented 2026-10-08 (uncommitted, not deployed, no migration run) Written 2026-10-08 · Scope: krow_talent_app, doormile_backend, AI_engine, kubernetes


Status (2026-10-08)

Verified: backend go build/go vet clean, 20 packages pass · engine 148 tests (3 pre-existing import errors) · console 1397 pass (13 pre-existing agentStudio.test.jsx failures) · kubectl kustomize manifests/doormile renders clean. Nothing committed. No migration has run. No SQL in this work has ever executed against a database.

Item State
A1 config/secrets manifest, envFrom, probes, resources built — placeholders unreplaced
A2 secrets out of git not done — documented in manifests/doormile/SECRETS.md, including the kustomize-clobbers-hand-created-secrets trap. Five more placeholder keys were added.
A3 anthropic>=0.49.0 · A4 claude-opus-5-5 done
B0.1 CREATE EXTENSION vector · B0.2 /similar → POST done
B1 provider (OpenAI text-embedding-3-small, 1536) done — my call, flagged
B2 core/embeddings.py + decision wiring done
B3 outcome sweeper + rules + 14 tests done
B4 core/memory.py + both agents done
B5 retention/prune done — ivfflat lists re-tune still needs a row count
B6 tenantid scoping done
B2.5 historical backfill done — internal/ai/outcomes/backfill.go
B7 persist console findings done — aiskillfindings, 4 routes, findingReport.js, 15 tests
C1 health surface + ai-engine.yaml + Dockerfile/compose port done — image not built or pushed
C2 kustomization + deploy-doormile.sh done
C3 probes/resources/PDB done — StatefulSet→Deployment not done
C4 ingress built — apply order matters, see the file header
C5 NATS port note done — docs/ARCHITECTURE.md
C6 engine replica safety see correction 1 — was already safe for a different reason
D1 admin batch-assign + console executor done
D2 alert_low_battery_rider done
D3 five review-only tools 1 of 5 done — trigger_auto_dispatch turned out to BE batch-assign under another name. The other four need endpoints that do not exist.
Agents-page snapshot → live endpoint backend done — engine serves GET /agents/status; the console still reads the snapshot, so the swap is one fetch away

Two things this plan got wrong

Correction 1 — C6. The engine's stall dedup was never per-pod. This plan claimed ExceptionAgent kept its dedup in an in-process TTL set and that moving it to Redis was the prerequisite for scaling out. Wrong: _claim_stall / _claim_stall_handled (agents/exception_agent.py:336) have been Redis SET NX with a TTL all along — cross-pod, self-expiring, durable across restarts. The comment there saying "replaces the old unbounded in-memory set" describes what it REPLACED; I read it as current state.

The real blocker is different and still real: core/message_bus.py:284 and :312 bind durable push consumers with fixed names, and a durable push consumer admits one active subscriber unless created with a deliver group. A second replica does not duplicate work — it fails to bind. Scaling out means giving those subscriptions a deliver group (or moving to pull consumers). replicas: 1 stands, for the corrected reason.

Correction 2 — the registry was already honest about the simulated agents. This plan said the three fictional-data agents were "presented as live" and should be marked. They already were: seed.go has StatusSimulation for HUB_AGENT, FLEET_AGENT and ROUTE_OPTIMIZER, with purposes reading "an in-memory fleet of 19 fake vehicles" and "8 hard-coded fictional hubs". The console snapshot (src/lib/agentNetwork.js) was also honest about activity ("Zero tasks received", publishes: 'Nothing on the bus.', api: 'No.') but silent on the data being invented — so the three entries now say so, and the two surfaces agree.

I overstated that surface as "telling anyone something untrue". It was incomplete, not false.


Track A — Unblock (do this first; nothing else matters until it lands)

A1. INTERNAL_API_KEY is missing from the cluster — everything engine↔backend is 401

middlewares/internal_auth.go:13 reads INTERNAL_API_KEY and fails closed when unset:

expected := os.Getenv("INTERNAL_API_KEY")
if expected == "" || c.Get("X-Internal-Key") != expected {

kubernetes/manifests/doormile/miletruth.yaml never sets it. The engine's docker-compose.yml does pass it. So the engine sends a key the backend rejects, and in-cluster every /api/v1/internal/* call 401s:

  • GET /internal/ai/registry — the registry poll (this is why "engine reads registry in prod" was never confirmable)
  • POST /internal/agent-decisions — the decision log, so Insights sees nothing
  • GET /internal/express/riders, /express/bookings, POST /express/assign
  • POST /internal/bookings/:id/reassign, POST /internal/notify

Step 1 — find out whether git matches reality. The kubernetes repo's log is full of Restore… / Recreate… commits, so the manifest may already be fiction:

kubectl -n doormile get statefulset doormile -o jsonpath='{range .spec.template.spec.containers[0].env[*]}{.name}{"\n"}{end}' | sort

config/config.go reads 29 env vars. If that list is much shorter, the live cluster has hand-applied drift and the next kubectl apply -f wipes it.

Step 2 — close the gap in git, not by hand. Add to doormile-secrets (stringData) and reference from the StatefulSet. Absent from the manifest today and read by config.go:

INTERNAL_API_KEY, JWT_SECRET_KEY, APP_PORT, ENV, TRUSTED_PROXIES, AI_LAYER_BASE_URL, ROUTE_OPTIMIZER_URL, GEOCODER_URL, GEOCODER_EMAIL, SMTP_HOST, SMTP_PORT, SMTP_USER, SMTP_PASSWORD, SMTP_FROM, CLIENT_ONBOARDING_OWNERS, PLAYGROUND_LLM_API_KEY, PLAYGROUND_LLM_BASE_URL, PLAYGROUND_LLM_MODEL, REDIS_USER.

Secrets vs ConfigMap: INTERNAL_API_KEY, JWT_SECRET_KEY, SMTP_PASSWORD, PLAYGROUND_LLM_API_KEY are secrets. The rest belong in a doormile-config ConfigMap so a URL change is not a secret edit.

Files

File Change
kubernetes/manifests/doormile/miletruth.yaml modify — add the 19 env vars to the StatefulSet + doormile-secrets
kubernetes/manifests/doormile/doormile-config.yaml new — ConfigMap for the non-secret vars
AI_engine/.env (server-local, untracked content) modify — same INTERNAL_API_KEY value

Read-only, no change: doormile_backend/middlewares/internal_auth.go:13, doormile_backend/config/config.go:60-102 (the contract being satisfied).

INTERNAL_API_KEY must be byte-identical in the backend Secret and the engine's env. Generate once (openssl rand -hex 32), set both.

Acceptance: from inside the cluster, curl -H "X-Internal-Key: $KEY" http://doormile-service.doormile:8081/api/v1/internal/ai/registry returns 200 with a body and an ETag; a second call with If-None-Match returns 304. Backend logs show ai playground: enabled. GET /admin/ai/status reports the engine as following the registry (aiRegistryController.go:208 counts the poll).

A2. Plaintext secrets are committed

miletruth.yaml:19-21 has DB_PASSWORD, REDIS_PASSWORD, NATS_PASSWORD as literal in git (the same password reused for all three). Same class as the committed Firebase keys from the CX audit. Pick one and apply it before A1 adds more secrets to that file: Sealed Secrets, SOPS, or an out-of-git kubectl create secret with the manifest holding only a reference.

Do not delete or untrack anything here without per-file approval — see the standing rule. Rotating that shared password is a separate decision; this item is only about stopping new secrets entering git.

Files

File Change
kubernetes/manifests/doormile/miletruth.yaml modify — stringData block becomes a reference
kubernetes/manifests/doormile/sealed-secrets.yaml new — if Sealed Secrets is the chosen route
kubernetes/docs/DEPLOY.md modify — document how secrets are supplied now

Same pattern already in git at kubernetes/manifests/core/core-secrets.yaml and manifests/nearle/nearle-secrets.yaml — whatever is chosen should cover those too, but that is outside this phase.

A3. anthropic>=0.21.0 floor is wrong

AI_engine/requirements.txt pins anthropic>=0.21.0, but core/llm.py:117 sends thinking: {"type": "adaptive"} and output_config.effort — parameters that old SDK does not know. There is no lockfile, so a fresh pip install in a rebuilt image can resolve to something that 400s on every LLM call.

Bump to a current floor and pin the image build. openai>=1.0.0 is already present (unused today — Track B uses it).

Files

File Change
AI_engine/requirements.txt modify — raise the anthropic floor
AI_engine/Dockerfile modify — optional: pip install from a lockfile instead
AI_engine/requirements-dev.txt check — may pin the same packages

A4. LLM_MODEL default is a generation behind

core/llm.py:22 defaults to claude-opus-4-8 ($5/$25 per MTok). claude-opus-5-5 is both cheaper ($4/$20) and more capable. One-line change; the registry's per-agent model pin already overrides it (registry.model(agent_id)), so this only moves the floor.

Files

File Change
AI_engine/core/llm.py:22 modify — default LLM_MODEL
AI_engine/core/llm.py:11 modify — the docstring naming the default
AI_engine/docker-compose.yml:45 modify — LLM_MODEL fallback
AI_engine/tests/test_llm.py check — may assert the old default

Note core/llm.py:109 special-cases Haiku 4.5 (rejects adaptive thinking and effort). Leave that branch alone; it is correct.


Track B — RAG decision memory

The skeleton exists and is wired to nothing. Four pieces are already built:

Piece Where
context_embedding vector(1536) on agent_decisions migrations/migrate.go:92
ivfflat cosine index, lists = 100 migrations/migrate.go:98
POST /internal/agent-decisions accepts context_embedding controllers/agentDecisionController.go:23
cosine kNN query + PATCH /:id/outcome agentDecisionController.go:74,120

Nothing produces an embedding, nothing calls /similar, nothing records an outcome. There are exactly two decision types to cover — assignment_failure (dispatch_agent.py:335) and stall_response (exception_agent.py:456).

B0. Three defects to fix before writing any new code

B0.1 — CREATE EXTENSION vector appears nowhere in the repo. migrate.go:92 runs ALTER TABLE agent_decisions ADD COLUMN … vector(1536) and logs failure non-fatally. If the extension is not installed on logistics, both the column and the index silently fail and every retrieval 500s.

SELECT extname, extversion FROM pg_extension WHERE extname = 'vector';

If absent, add before line 92 in Migrate():

if res := db.Exec(`CREATE EXTENSION IF NOT EXISTS vector`); res.Error != nil {
    utils.Error("❌ pgvector extension unavailable — decision memory disabled", "error", res.Error)
}

Needs the pgvector extension available on the server and a role with rights to create it. On a managed Postgres this may be an admin action, not a migration — check before assuming.

B0.2 — /similar is a GET that requires a JSON body. routes.go:594 registers it as GET, and agentDecisionController.go:84 calls BodyParser demanding an embedding array. nginx and most HTTP clients drop GET bodies — and doormile is served through host nginx (see C4). Change to:

internal.Post("/agent-decisions/similar", controllers.FindSimilarDecisions)

No caller exists yet, so this breaks nothing.

B0.3 — /similar filters WHERE outcome IS NOT NULL, and nothing writes outcomes. Even with embeddings flowing it returns zero rows forever. B2 is not optional — it is the half that makes retrieval worth anything.

Files

File Change
doormile_backend/migrations/migrate.go:92 modify — add CREATE EXTENSION above the ALTER TABLE
doormile_backend/routes/routes.go:594 modify — internal.Get → internal.Post for /agent-decisions/similar
doormile_backend/controllers/agentDecisionController.go:74 modify — comment the method change; body parsing already correct
doormile_backend/routes/routes_ai_registry_pg_test.go modify — add coverage for the POST shape

B1. Embedding provider — decision required

Anthropic has no embeddings endpoint, so this needs a second provider.

Option Dims Column change Notes
OpenAI text-embedding-3-small 1536 none ~$0.02/MTok. openai>=1.0.0 already a dependency. Column was sized for it.
Voyage voyage-3 1024 yes Anthropic-recommended; new key, new vendor.
Local bge-small / MiniLM 384 yes Free, no egress, no key. +~400MB RAM, model in image, slower cold start.

Recommendation: OpenAI text-embedding-3-small. The schema already matches, the dependency is already there, and the second-provider line is already crossed (the playground runs on Groq). Revisit if data egress is a constraint — then take the local model and migrate the column to vector(384).

B2. Write path — AI_engine/core/embeddings.py

One function, modelled on core/decisions.py's fire-and-forget discipline:

async def embed(text: str) -> Optional[List[float]]:
    """None on any failure. An embedding must never delay an agent's reaction."""

Requirements:

  • Embed the same dict build_payload already stores. Serialise the facts dict deterministically (sorted keys) so the embedded text and the stored context cannot drift apart. Add a _context_text(facts) helper and test it pure, the way build_payload is tested.
  • Fail open — return None, log once, never raise into the agent path.
  • Small LRU cache: repeated stalls on one booking produce near-identical facts.
  • Gate on EMBEDDINGS_ENABLED and a registry skill flag, so it can be switched off from Agent Studio without a redeploy (registry.skill_enabled(...) already exists).

Then extend build_payload/record_decision (core/decisions.py:33,53) to carry context_embedding. Both call sites (dispatch_agent.py:335, exception_agent.py:456) keep their signatures.

Acceptance: after one stall, SELECT count(*) FROM agent_decisions WHERE context_embedding IS NOT NULL > 0.

Files

File Change
AI_engine/core/embeddings.py new — embed(), _context_text(), LRU cache, fail-open
AI_engine/core/decisions.py:33 modify — build_payload carries context_embedding
AI_engine/core/decisions.py:53 modify — record_decision awaits the embed before posting
AI_engine/config/system_config.py modify — EMBEDDINGS_ENABLED, provider key, model name
AI_engine/docker-compose.yml modify — pass the embedding env through
AI_engine/tests/test_embeddings.py new — _context_text determinism, fail-open returns None
AI_engine/tests/test_registry_phase5.py:194 modify — asserts the record_decision payload shape

Not touched: agents/dispatch_agent.py:335 and agents/exception_agent.py:456 keep their call signatures — the embedding is added inside decisions.py, so neither agent changes.

B3. Outcome loop — the part that makes retrieval useful

PATCH /internal/agent-decisions/:id/outcome exists with no caller. Define, per decision type, what "it worked" means. Starting proposal:

Type Outcome = success when failure when
stall_response booking reaches Delivered within its SLA window after the decision SLA breached, or cancelled
assignment_failure booking gets an assignment within N minutes of the decision still unassigned after N, or cancelled

Implement as a backend sweeper following the established pattern in internal/assignment/sweeper.go:88 — ticker, recover() per tick, and the Redis lock that keeps one replica sweeping (sweeper.go:104). Wire it in main.go beside go assignment.StartPendingSweeper() (main.go:244).

Two gotchas from the existing code:

  • Use Receivedat-style true instants, not utils.DBNow — models/ai_runs.go:26 documents that DBNow returns IST digits labelled UTC and is 5h30m off for timestamptz. The same trap applies to outcome_recorded_at.
  • Leave outcome NULL while undecided. /similar already treats NULL as "no evidence yet", which is correct.

Acceptance: rows acquire non-NULL outcome within one sweep interval of their window closing, and /insights decision-outcome counts stop being all pending.

Files

File Change
doormile_backend/internal/ai/outcomes/sweeper.go new — ticker + Redis lock, modelled on internal/assignment/sweeper.go:88
doormile_backend/internal/ai/outcomes/rules.go new — the per-decision-type success/failure predicates
doormile_backend/internal/ai/outcomes/sweeper_test.go new — interval, window, and both predicates
doormile_backend/main.go:244 modify — go outcomes.StartOutcomeSweeper() beside the pending sweeper
doormile_backend/models/agentdecision.go modify — only if B6 adds Tenantid

Reference, not modified: internal/assignment/sweeper.go:88-110 (the ticker + recover() + Redis single-replica lock pattern to copy), models/ai_runs.go:26 (the utils.DBNow timezone trap to avoid).

B4. Read path — precedent in the prompt

Before the LLM call in exception_agent / dispatch_agent, fetch the top-5 similar resolved decisions and include them as precedent — "the last 5 comparable situations and whether the action worked."

  • Feature-flag it on a registry skill so it is switchable from Agent Studio.
  • Hard timeout (~300ms) with fail-open to today's prompt. The current behaviour is the floor; this can only raise it. (core/registry.py already follows this rule; ragRouter.js:9-14 documents the same discipline on the console side.)
  • Keep precedent out of the structured-output schema. It informs the prompt; it must not become a field the model can invent.

Acceptance: an eval run shows the decision quality moving. AI_engine/evals/ already has the harness and cases (stall_cases.jsonl, assignment_cases.jsonl) — extend those rather than judging by eye.

Files

File Change
AI_engine/core/memory.py new — recall(decision_type, embedding, k) → POST /internal/agent-decisions/similar, timeout + fail-open
AI_engine/core/llm.py:157 modify — decide_stall_response accepts optional precedent
AI_engine/core/llm.py:227 modify — decide_assignment_failure accepts optional precedent
AI_engine/agents/exception_agent.py:456 modify — recall before the decide call
AI_engine/agents/dispatch_agent.py:335 modify — recall before the decide call
AI_engine/tests/test_memory.py new — timeout fails open, empty recall changes nothing
AI_engine/evals/stall_eval.py modify — run with and without precedent
AI_engine/evals/assignment_eval.py modify — same
AI_engine/evals/stall_cases.jsonl modify — cases where precedent should change the answer
AI_engine/evals/assignment_cases.jsonl modify — same
doormile_backend/internal/ai/registry/seed.go:114 modify — seed a recall_similar_decisions read tool + the skill flag that gates B4

The registry seed row matters: without it the feature cannot be switched off from Agent Studio, which is the whole point of gating it on registry.skill_enabled.

B5. Hygiene

  • agent_decisions has no retention. aiagentruns purges at 30 days (telemetry/recorder.go:28,128); decisions grow forever, and this is the table retrieval scans. Decide a window — longer than 30 days, since old precedent is the point. Suggest 180 days, or keep resolved rows and purge unresolved ones.
  • Re-tune lists = 100. That is right for roughly 100k–1M rows. Below ~10k it over-partitions and recall drops. Check count(*) once embeddings flow; consider HNSW instead if the pgvector version supports it.

Files

File Change
doormile_backend/internal/ai/outcomes/sweeper.go modify — fold the decision purge into the same tick
doormile_backend/migrations/migrate.go:98 modify — index tuning, once row count is known

B6. Tenant isolation — decide before B4 ships, not after

agent_decisions has no tenant column. If retrieved precedent crosses tenants, one client's operational history shapes decisions made for another. Given that console logins are already unscoped on tenantid NULL (doormile-console-logins-unscoped), this needs deciding up front.

Recommendation: add tenantid to agent_decisions, have the engine populate it, and filter in the /similar query. Cheap now, expensive after the table fills.

Files

File Change
doormile_backend/models/agentdecision.go modify — add Tenantid *uint64 with an index
doormile_backend/controllers/agentDecisionController.go:18 modify — accept tenant_id on create
doormile_backend/controllers/agentDecisionController.go:104 modify — add AND tenantid = ? to the kNN query
AI_engine/core/decisions.py:33 modify — build_payload carries the tenant
AI_engine/agents/exception_agent.py · dispatch_agent.py modify — source the tenant from the booking facts
doormile_backend/routes/routes_ai_registry_pg_test.go modify — a cross-tenant recall must return nothing

Nullable, because the engine will not always know the tenant. Decide whether a NULL tenant row is recallable by everyone or by no one — given doormile-console-logins-unscoped, by no one is the safer default.


Track C — Kubernetes

C1. AI_engine is not in Kubernetes at all

No manifest, no kustomization, no deploy script, and docker-compose.yml uses build: . with no registry push. It runs on a VM by compose.

Prerequisite: the engine has no HTTP server in production mode. main.py --production starts no listener, so there is no liveness/readiness target and no /metrics. Without it, Kubernetes can only restart on process exit — a NATS-disconnected engine looks healthy forever.

fastapi and uvicorn are already in requirements.txt. Add a small surface in production_mode():

  • GET /healthz — process alive (event loop responsive)
  • GET /readyz — NATS connected and registry.loaded is true
  • GET /metrics — optional; decisions recorded, tool calls, LLM failures

Readiness must include registry.loaded, otherwise a pod that cannot reach the backend serves traffic on env defaults while reporting healthy — exactly the A1 failure mode, invisible again.

Then: push the image to a registry (compose builds locally), and write manifests/doormile/ai-engine.yaml as a Deployment (it is stateless; replicas: 1 to start — the agents are not yet idempotent across replicas, see C6).

Files

File Change
AI_engine/core/health.py new — the aiohttp/FastAPI surface (/healthz, /readyz, /metrics)
AI_engine/main.py:130 modify — start the health server inside production_mode() beside registry.run(...)
AI_engine/main.py (print_help) modify — document the health port
AI_engine/Dockerfile modify — EXPOSE the health port
AI_engine/docker-compose.yml modify — publish the port so compose and k8s behave alike
AI_engine/tests/test_health.py new — /readyz is red while registry.loaded is false
kubernetes/manifests/doormile/ai-engine.yaml new — Deployment + Service + probes + resources

fastapi and uvicorn are already in requirements.txt — no new dependency. Readiness must check registry.loaded (core/registry.py:45), not just the process, or A1's failure mode becomes invisible again.

C2. doormile/ has no kustomization.yaml

alaska/, core/ and nearle/ all have one. doormile/ does not, and there is no deploy-doormile.sh alongside deploy-core-stack.sh / deploy-nearle-stack.sh. That is why it drifts. Add both.

Files

File Change
kubernetes/manifests/doormile/kustomization.yaml new — list miletruth, config, ai-engine, pdb
kubernetes/deploy-doormile.sh new — copy the shape of deploy-nearle-stack.sh
kubernetes/scripts/sync_manifests.py check — may need the new namespace registering
kubernetes/docs/DEPLOY_CHECKLIST.md modify — add the doormile stack

Pattern to copy: manifests/nearle/kustomization.yaml + deploy-nearle-stack.sh.

C3. The doormile StatefulSet has no probes, resources or PDB

replicas: 3 with no resource requests means the scheduler can stack all three on one node, and no readinessProbe means a pod receives traffic before Postgres/Redis/NATS are connected. core/ has worker-pdb.yaml; doormile has nothing.

Add resources.requests/limits, a readinessProbe and livenessProbe against the backend's health route, and a PodDisruptionBudget (minAvailable: 2).

Also: a StatefulSet for a stateless Go API is the wrong kind — it gives serial rollouts and no benefit. Switching to Deployment is low-risk and makes deploys faster. Not urgent; flagging because it is why rollouts feel slow.

Files

File Change
kubernetes/manifests/doormile/miletruth.yaml:23 modify — resources, readinessProbe, livenessProbe
kubernetes/manifests/doormile/doormile-pdb.yaml new — minAvailable: 2, copy manifests/core/worker-pdb.yaml

No backend change needed — the probe targets already exist and are correct: GET /api/v1/health (routes/routes.go:46, unauthenticated, always 200) for liveness, and GET /api/v1/ready (routes/routes.go:50) for readiness, which already returns 503 when Postgres or Redis is unreachable (routes.go:68-70). Point the probes at those; do not write new ones.

Note /ready reports Redis GEO status without gating on it — deliberate, per the comment at routes.go:73. A readinessProbe on /ready therefore will not pull a pod out of service for a broken rider search, which is the intended behaviour.

C4. No ingress for doormile

nearle and alaska are on manifests/core/ingress-unified.yaml. doormile is NodePort 30830 plus host nginx (conf/nginx-doormile.conf) — half-migrated. This also makes B0.2 (GET-with-body) a certainty rather than a risk.

Files

File Change
kubernetes/manifests/core/ingress-unified.yaml modify — add a doormile rule (needs a ReferenceGrant if it stays cross-namespace, cf. manifests/nearle/nearle-reference-grant.yaml)
kubernetes/manifests/doormile/miletruth.yaml:91 modify — NodePort → ClusterIP once the ingress serves it
kubernetes/conf/nginx-doormile.conf modify — retire or repoint, only after the ingress is verified

Do these in that order. Flipping the Service type before the ingress works takes the API offline.

C5. NATS is outside the cluster on two ports

Backend uses nats://66.116.226.161:4223; core-config.yaml:10 uses :4222. Worth a line in docs/ARCHITECTURE.md on which port is which and why, before the engine joins and needs to pick one.

C6. Decide replica safety before scaling the engine

The telemetry recorder is already replica-safe (NATS queue group + uq_aiagentruns_agent_task). The agents are not obviously so: exception_agent has an in-process TTL dedup set (exception_agent.py:339), which is per-pod. Two engine replicas would each decide on the same stall. Keep replicas: 1 until dedup moves to Redis.

Files (only if scaling past 1 replica)

File Change
AI_engine/agents/exception_agent.py:339 modify — TTL set → Redis SET NX EX
AI_engine/tests/test_stall_dedup.py modify — covers the current in-process behaviour
kubernetes/manifests/doormile/ai-engine.yaml modify — raise replicas

C5 (the NATS port note) is documentation only: kubernetes/docs/ARCHITECTURE.md.


Track D — Executor backlog

9 of 22 seeded tools are marked REVIEW ONLY … No executor in internal/ai/registry/seed.go. The registry is honest about it and the console renders them disabled. This is the feature list, in value order.

D1. assign_riders — needs an admin auto-assign route (not just auth)

Correcting an earlier assumption: this is not a free auth fix. actions.js:96-107 already explains why — POST /admin/bookings/:id/assign-miler (routes.go:413 → adminController.go:2965) requires a chosen rider per booking ({mileruserid}), and a finding does not pick one. The hub route that does pick (POST /hub/bookings/:id/auto-assign, routes.go:520) is behind HubStaffAuth and 403s for every console login.

The cleanest route is an admin batch-assign, better than the per-booking auto-assign first considered. HubBatchAssign (controllers/hubController.go:1963) already does exactly what a finding needs: it takes bookingids[], picks riders via Redis GEO + scoring, and commits server-side in one call. Its only hub-specific parts — c.Locals("hubid") and hubPincodePrefix(hubID) (hubController.go:1964-1966) — are used solely as a fallback when bookingids is empty (hubController.go:1981-1985). The console always passes explicit ids, so that branch never runs.

So: extract the body into a shared helper taking (bookingIDs, capPerRider, actorID, scopeFn) and have both the hub route and a new admin route call it. scopeBookingsToOwnTenant (hubController.go:1986) already works for admin logins.

The console side is then nearly free — batchAssignBookings already exists at src/api/doormile/endpoints.js:554 and already sends {bookingids, max_per_rider}. It just points at the hub URL that 403s. One URL change.

Option (b), having the skill pick a rider via nearby_milers and call the existing assign-miler, is worse: more console work and it puts solver logic in the browser.

Files

File Change
doormile_backend/controllers/hubController.go:1963 modify — extract the shared assign helper out of HubBatchAssign
doormile_backend/controllers/adminController.go add AdminBatchAssign calling that helper
doormile_backend/routes/routes.go:413 modify — register adminAuth.Post("/bookings/batch-assign", …)
krow_talent_app/src/api/doormile/endpoints.js:554 modify — /hub/bookings/batch-assign → /admin/bookings/batch-assign
krow_talent_app/src/lib/assistant/agent/actions.js:122 modify — add the assignMiler executor
krow_talent_app/src/lib/assistant/agent/actions.js:96-107 modify — delete the "deliberately NOT an executor" note
krow_talent_app/tests/lib/agentActions.test.js modify — asserts the current executor set
doormile_backend/internal/ai/registry/seed.go modify — assign_riders description stops saying REVIEW ONLY

Check before starting: endpoints.js:547-553 warns that Doormile-native batch assign commits with no preview/reconcile step and leaves multi-stop riders unsequenced. An agent-proposed assignment firing straight to commit is a behaviour decision, not just a wiring one — confirm that is wanted.

This takes the console from 1 working verb to 2 and makes the highest-severity SLA finding actionable.

D2. alert_low_battery_rider — nearly free

seed.go already targets POST /admin/milers/:id/notify, which is the endpoint notify_riders already uses successfully. This is a message-text change and an EXECUTORS entry, not a new capability.

Files

File Change
krow_talent_app/src/lib/assistant/agent/actions.js:122 modify — add the alertLowBatteryRider executor
krow_talent_app/src/lib/assistant/skills/definitions/RiderBatterySafetySkill.js check — confirm the proposal carries milerId
krow_talent_app/tests/lib/agentActions.test.js modify — same assertion as D1
doormile_backend/internal/ai/registry/seed.go modify — drop REVIEW ONLY from the description

No backend change. notifyMiler already exists at src/api/doormile/endpoints.js:389 → POST /admin/milers/:id/notify, which is the endpoint seed.go already names as the target.

D3. The rest need endpoints that do not exist

enforce_otp_verification, dispatch_hub_idle_parcels, trigger_auto_dispatch, enforce_cash_handoff, rebalance_riders — all marked Target: "none yet". Each is a product decision first. Not in this phase.


Two more loose ends

  • src/lib/agentNetwork.js is a hand-maintained snapshot dated 16–20 Sep, and the file says so honestly. Once C1 gives the engine an HTTP surface, add GET /agents/status and swap the source. The file is deliberately shaped like that response, so it is a change of source, not a rewrite.
  • src/lib/assistant/ragRouter.js points at a services/ai sidecar that does not exist in any repo. VITE_AI_URL appears nowhere, so isRagEnabled() is permanently false and the module is dead code. Do not conflate this with Track B — ragRouter is semantic intent routing for the console assistant, not decision memory. Lower value. Leave it dormant (it is correctly fail-open) or decide to build the sidecar as its own piece of work.

Consolidated file manifest

14 new files, 46 modified, across 4 repos. Per-item detail is in the tracks above.

doormile_backend — 4 new, 17 modified

File New? Items
internal/ai/outcomes/sweeper.go new B3, B5
internal/ai/outcomes/rules.go new B3
internal/ai/outcomes/sweeper_test.go new B3
migrations/migrate.go B0.1 (:92), B5 (:98)
routes/routes.go B0.2 (:594), D1 (:413)
controllers/agentDecisionController.go B0.2 (:74), B6 (:18, :104)
controllers/hubController.go D1 (:1963 — extract helper)
controllers/adminController.go D1 (add AdminBatchAssign)
models/agentdecision.go B6
internal/ai/registry/seed.go B4 (:114), D1, D2
main.go B3 (:244)
routes/routes_ai_registry_pg_test.go B0.2, B6

AI_engine — 5 new, 20 modified

File New? Items
core/embeddings.py new B2
core/memory.py new B4
core/health.py new C1
tests/test_embeddings.py new B2
tests/test_memory.py · tests/test_health.py new B2, C1
core/decisions.py B2 (:33, :53), B6
core/llm.py A4 (:11, :22), B4 (:157, :227)
core/registry.py — read-only (loaded consumed by C1)
agents/exception_agent.py B4 (:456), B6, C6 (:339)
agents/dispatch_agent.py B4 (:335), B6
main.py C1 (:130, print_help)
config/system_config.py B2
requirements.txt · Dockerfile · docker-compose.yml A3, A4, C1
evals/stall_eval.py · assignment_eval.py · both .jsonl B4
tests/test_registry_phase5.py · test_llm.py · test_stall_dedup.py B2, A4, C6

kubernetes — 5 new, 7 modified

File New? Items
manifests/doormile/doormile-config.yaml new A1
manifests/doormile/ai-engine.yaml new C1, C6
manifests/doormile/kustomization.yaml new C2
manifests/doormile/doormile-pdb.yaml new C3
deploy-doormile.sh new C2
manifests/doormile/miletruth.yaml A1, A2, C3 (:23), C4 (:91)
manifests/core/ingress-unified.yaml C4
conf/nginx-doormile.conf C4 (last)
scripts/sync_manifests.py C2
docs/DEPLOY.md · DEPLOY_CHECKLIST.md · ARCHITECTURE.md A2, C2, C5

krow_talent_app — 0 new, 4 modified

The console barely changes. Everything it needs already exists.

File Items
src/api/doormile/endpoints.js D1 (:554 — one URL)
src/lib/assistant/agent/actions.js D1 (:96-107, :122), D2 (:122)
tests/lib/agentActions.test.js D1, D2
src/lib/assistant/skills/definitions/RiderBatterySafetySkill.js D2 — check only

Untouched on purpose: src/lib/agentNetwork.js (until C1 ships an endpoint) and src/lib/assistant/ragRouter.js (dormant, out of scope — see loose ends).

Files deliberately not touched

  • src/lib / src/utils duplicate trees — both have live importers, do not consolidate.
  • AI_engine/core/tool_registry.py — the 2-vs-22 tool gap is a separate decision, not this phase.
  • AI_engine/customer_portal/, dashboard/ — not on any path this phase touches.

Decisions needed before coding

  1. Embedding provider — OpenAI 1536 (no column change), Voyage 1024, or local 384? (B1)
  2. Is pgvector installed on logistics? If creating extensions needs an admin, that is a prerequisite, not a migration. (B0.1)
  3. Outcome definitions — are the two in B3 right, and what is N for assignment_failure?
  4. tenantid on agent_decisions — add it now, or accept cross-tenant precedent? (B6)
  5. assign_riders — route (a) admin auto-assign, or (b) console picks the rider? (D1)
  6. Decision retention window — 180 days, or keep-resolved-purge-unresolved? (B5)

Sequencing

A1 ──> A3, A4 ──┬──> B0 ──> B2 ──> B3 ──> B4 ──> B5, B6
                │
                └──> C1(health) ──> C1(manifest) ──> C2 ──> C3 ──> C4
A2 (independent, before A1 adds more secrets to git)
D1, D2 (independent of everything above)

A1 first and alone. Until the internal key is set, the engine is not talking to the backend, so every Track B acceptance check would fail for the wrong reason. Run the kubectl check in A1 before writing any code — if git does not match the cluster, that changes the shape of Track C.

Cheapest real step up a level, once A1 is in: D1 + D2. Two executors, no new LLM work, and it moves autonomy off the floor.


Standing constraints

  • Nothing in this plan is committed, pushed, or deployed without being asked.
  • No migrations run against a real database without being asked.
  • No files deleted or untracked without per-file approval.
  • src/lib and src/utils in the console both have live importers — neither tree is dead, do not consolidate them as part of this work.