38 KiB
Doormile Agent Platform — Phase 7: Decision Memory, Deployment, Executors
Status: mostly implemented 2026-10-08 (uncommitted, not deployed, no migration run)
Written 2026-10-08 · Scope: krow_talent_app, doormile_backend, AI_engine, kubernetes
Status (2026-10-08)
Verified: backend go build/go vet clean, 20 packages pass · engine 148 tests
(3 pre-existing import errors) · console 1397 pass (13 pre-existing
agentStudio.test.jsx failures) · kubectl kustomize manifests/doormile renders
clean. Nothing committed. No migration has run. No SQL in this work has ever
executed against a database.
| Item | State |
|---|---|
A1 config/secrets manifest, envFrom, probes, resources |
built — placeholders unreplaced |
| A2 secrets out of git | not done — documented in manifests/doormile/SECRETS.md, including the kustomize-clobbers-hand-created-secrets trap. Five more placeholder keys were added. |
A3 anthropic>=0.49.0 · A4 claude-opus-5-5 |
done |
B0.1 CREATE EXTENSION vector · B0.2 /similar → POST |
done |
B1 provider (OpenAI text-embedding-3-small, 1536) |
done — my call, flagged |
B2 core/embeddings.py + decision wiring |
done |
| B3 outcome sweeper + rules + 14 tests | done |
B4 core/memory.py + both agents |
done |
| B5 retention/prune | done — ivfflat lists re-tune still needs a row count |
B6 tenantid scoping |
done |
| B2.5 historical backfill | done — internal/ai/outcomes/backfill.go |
| B7 persist console findings | done — aiskillfindings, 4 routes, findingReport.js, 15 tests |
C1 health surface + ai-engine.yaml + Dockerfile/compose port |
done — image not built or pushed |
C2 kustomization + deploy-doormile.sh |
done |
| C3 probes/resources/PDB | done — StatefulSet→Deployment not done |
| C4 ingress | built — apply order matters, see the file header |
| C5 NATS port note | done — docs/ARCHITECTURE.md |
| C6 engine replica safety | see correction 1 — was already safe for a different reason |
| D1 admin batch-assign + console executor | done |
D2 alert_low_battery_rider |
done |
| D3 five review-only tools | 1 of 5 done — trigger_auto_dispatch turned out to BE batch-assign under another name. The other four need endpoints that do not exist. |
| Agents-page snapshot → live endpoint | backend done — engine serves GET /agents/status; the console still reads the snapshot, so the swap is one fetch away |
Two things this plan got wrong
Correction 1 — C6. The engine's stall dedup was never per-pod.
This plan claimed ExceptionAgent kept its dedup in an in-process TTL set and
that moving it to Redis was the prerequisite for scaling out. Wrong:
_claim_stall / _claim_stall_handled (agents/exception_agent.py:336) have
been Redis SET NX with a TTL all along — cross-pod, self-expiring, durable
across restarts. The comment there saying "replaces the old unbounded in-memory
set" describes what it REPLACED; I read it as current state.
The real blocker is different and still real: core/message_bus.py:284 and
:312 bind durable push consumers with fixed names, and a durable push
consumer admits one active subscriber unless created with a deliver group. A
second replica does not duplicate work — it fails to bind. Scaling out means
giving those subscriptions a deliver group (or moving to pull consumers).
replicas: 1 stands, for the corrected reason.
Correction 2 — the registry was already honest about the simulated agents.
This plan said the three fictional-data agents were "presented as live" and
should be marked. They already were: seed.go has StatusSimulation for
HUB_AGENT, FLEET_AGENT and ROUTE_OPTIMIZER, with purposes reading "an
in-memory fleet of 19 fake vehicles" and "8 hard-coded fictional hubs". The
console snapshot (src/lib/agentNetwork.js) was also honest about activity
("Zero tasks received", publishes: 'Nothing on the bus.', api: 'No.') but
silent on the data being invented — so the three entries now say so, and the two
surfaces agree.
I overstated that surface as "telling anyone something untrue". It was incomplete, not false.
Track A — Unblock (do this first; nothing else matters until it lands)
A1. INTERNAL_API_KEY is missing from the cluster — everything engine↔backend is 401
middlewares/internal_auth.go:13 reads INTERNAL_API_KEY and fails closed when
unset:
expected := os.Getenv("INTERNAL_API_KEY")
if expected == "" || c.Get("X-Internal-Key") != expected {
kubernetes/manifests/doormile/miletruth.yaml never sets it. The engine's
docker-compose.yml does pass it. So the engine sends a key the backend
rejects, and in-cluster every /api/v1/internal/* call 401s:
GET /internal/ai/registry— the registry poll (this is why "engine reads registry in prod" was never confirmable)POST /internal/agent-decisions— the decision log, so Insights sees nothingGET /internal/express/riders,/express/bookings,POST /express/assignPOST /internal/bookings/:id/reassign,POST /internal/notify
Step 1 — find out whether git matches reality. The kubernetes repo's log is
full of Restore… / Recreate… commits, so the manifest may already be fiction:
kubectl -n doormile get statefulset doormile -o jsonpath='{range .spec.template.spec.containers[0].env[*]}{.name}{"\n"}{end}' | sort
config/config.go reads 29 env vars. If that list is much shorter, the live
cluster has hand-applied drift and the next kubectl apply -f wipes it.
Step 2 — close the gap in git, not by hand. Add to doormile-secrets
(stringData) and reference from the StatefulSet. Absent from the manifest today
and read by config.go:
INTERNAL_API_KEY, JWT_SECRET_KEY, APP_PORT, ENV, TRUSTED_PROXIES,
AI_LAYER_BASE_URL, ROUTE_OPTIMIZER_URL, GEOCODER_URL, GEOCODER_EMAIL,
SMTP_HOST, SMTP_PORT, SMTP_USER, SMTP_PASSWORD, SMTP_FROM,
CLIENT_ONBOARDING_OWNERS, PLAYGROUND_LLM_API_KEY, PLAYGROUND_LLM_BASE_URL,
PLAYGROUND_LLM_MODEL, REDIS_USER.
Secrets vs ConfigMap: INTERNAL_API_KEY, JWT_SECRET_KEY, SMTP_PASSWORD,
PLAYGROUND_LLM_API_KEY are secrets. The rest belong in a doormile-config
ConfigMap so a URL change is not a secret edit.
Files
| File | Change |
|---|---|
kubernetes/manifests/doormile/miletruth.yaml |
modify — add the 19 env vars to the StatefulSet + doormile-secrets |
kubernetes/manifests/doormile/doormile-config.yaml |
new — ConfigMap for the non-secret vars |
AI_engine/.env (server-local, untracked content) |
modify — same INTERNAL_API_KEY value |
Read-only, no change: doormile_backend/middlewares/internal_auth.go:13,
doormile_backend/config/config.go:60-102 (the contract being satisfied).
INTERNAL_API_KEYmust be byte-identical in the backend Secret and the engine's env. Generate once (openssl rand -hex 32), set both.
Acceptance: from inside the cluster,
curl -H "X-Internal-Key: $KEY" http://doormile-service.doormile:8081/api/v1/internal/ai/registry
returns 200 with a body and an ETag; a second call with If-None-Match returns
304. Backend logs show ai playground: enabled. GET /admin/ai/status reports
the engine as following the registry (aiRegistryController.go:208 counts the
poll).
A2. Plaintext secrets are committed
miletruth.yaml:19-21 has DB_PASSWORD, REDIS_PASSWORD, NATS_PASSWORD as
literal in git (the same password reused for all three). Same class as the committed Firebase keys from the
CX audit. Pick one and apply it before A1 adds more secrets to that file:
Sealed Secrets, SOPS, or an out-of-git kubectl create secret with the manifest
holding only a reference.
Do not delete or untrack anything here without per-file approval — see the standing rule. Rotating that shared password is a separate decision; this item is only about stopping new secrets entering git.
Files
| File | Change |
|---|---|
kubernetes/manifests/doormile/miletruth.yaml |
modify — stringData block becomes a reference |
kubernetes/manifests/doormile/sealed-secrets.yaml |
new — if Sealed Secrets is the chosen route |
kubernetes/docs/DEPLOY.md |
modify — document how secrets are supplied now |
Same pattern already in git at kubernetes/manifests/core/core-secrets.yaml and
manifests/nearle/nearle-secrets.yaml — whatever is chosen should cover those too,
but that is outside this phase.
A3. anthropic>=0.21.0 floor is wrong
AI_engine/requirements.txt pins anthropic>=0.21.0, but core/llm.py:117
sends thinking: {"type": "adaptive"} and output_config.effort — parameters
that old SDK does not know. There is no lockfile, so a fresh pip install in a
rebuilt image can resolve to something that 400s on every LLM call.
Bump to a current floor and pin the image build. openai>=1.0.0 is already
present (unused today — Track B uses it).
Files
| File | Change |
|---|---|
AI_engine/requirements.txt |
modify — raise the anthropic floor |
AI_engine/Dockerfile |
modify — optional: pip install from a lockfile instead |
AI_engine/requirements-dev.txt |
check — may pin the same packages |
A4. LLM_MODEL default is a generation behind
core/llm.py:22 defaults to claude-opus-4-8 ($5/$25 per MTok). claude-opus-5-5
is both cheaper ($4/$20) and more capable. One-line change; the registry's
per-agent model pin already overrides it (registry.model(agent_id)), so this
only moves the floor.
Files
| File | Change |
|---|---|
AI_engine/core/llm.py:22 |
modify — default LLM_MODEL |
AI_engine/core/llm.py:11 |
modify — the docstring naming the default |
AI_engine/docker-compose.yml:45 |
modify — LLM_MODEL fallback |
AI_engine/tests/test_llm.py |
check — may assert the old default |
Note core/llm.py:109 special-cases Haiku 4.5 (rejects adaptive thinking and
effort). Leave that branch alone; it is correct.
Track B — RAG decision memory
The skeleton exists and is wired to nothing. Four pieces are already built:
| Piece | Where |
|---|---|
context_embedding vector(1536) on agent_decisions |
migrations/migrate.go:92 |
ivfflat cosine index, lists = 100 |
migrations/migrate.go:98 |
POST /internal/agent-decisions accepts context_embedding |
controllers/agentDecisionController.go:23 |
cosine kNN query + PATCH /:id/outcome |
agentDecisionController.go:74,120 |
Nothing produces an embedding, nothing calls /similar, nothing records an
outcome. There are exactly two decision types to cover —
assignment_failure (dispatch_agent.py:335) and stall_response
(exception_agent.py:456).
B0. Three defects to fix before writing any new code
B0.1 — CREATE EXTENSION vector appears nowhere in the repo.
migrate.go:92 runs ALTER TABLE agent_decisions ADD COLUMN … vector(1536) and
logs failure non-fatally. If the extension is not installed on logistics,
both the column and the index silently fail and every retrieval 500s.
SELECT extname, extversion FROM pg_extension WHERE extname = 'vector';
If absent, add before line 92 in Migrate():
if res := db.Exec(`CREATE EXTENSION IF NOT EXISTS vector`); res.Error != nil {
utils.Error("❌ pgvector extension unavailable — decision memory disabled", "error", res.Error)
}
Needs the pgvector extension available on the server and a role with rights to
create it. On a managed Postgres this may be an admin action, not a migration —
check before assuming.
B0.2 — /similar is a GET that requires a JSON body.
routes.go:594 registers it as GET, and agentDecisionController.go:84 calls
BodyParser demanding an embedding array. nginx and most HTTP clients drop GET
bodies — and doormile is served through host nginx (see C4). Change to:
internal.Post("/agent-decisions/similar", controllers.FindSimilarDecisions)
No caller exists yet, so this breaks nothing.
B0.3 — /similar filters WHERE outcome IS NOT NULL, and nothing writes
outcomes. Even with embeddings flowing it returns zero rows forever. B2 is
not optional — it is the half that makes retrieval worth anything.
Files
| File | Change |
|---|---|
doormile_backend/migrations/migrate.go:92 |
modify — add CREATE EXTENSION above the ALTER TABLE |
doormile_backend/routes/routes.go:594 |
modify — internal.Get → internal.Post for /agent-decisions/similar |
doormile_backend/controllers/agentDecisionController.go:74 |
modify — comment the method change; body parsing already correct |
doormile_backend/routes/routes_ai_registry_pg_test.go |
modify — add coverage for the POST shape |
B1. Embedding provider — decision required
Anthropic has no embeddings endpoint, so this needs a second provider.
| Option | Dims | Column change | Notes |
|---|---|---|---|
OpenAI text-embedding-3-small |
1536 | none | ~$0.02/MTok. openai>=1.0.0 already a dependency. Column was sized for it. |
Voyage voyage-3 |
1024 | yes | Anthropic-recommended; new key, new vendor. |
Local bge-small / MiniLM |
384 | yes | Free, no egress, no key. +~400MB RAM, model in image, slower cold start. |
Recommendation: OpenAI text-embedding-3-small. The schema already matches,
the dependency is already there, and the second-provider line is already crossed
(the playground runs on Groq). Revisit if data egress is a constraint — then take
the local model and migrate the column to vector(384).
B2. Write path — AI_engine/core/embeddings.py
One function, modelled on core/decisions.py's fire-and-forget discipline:
async def embed(text: str) -> Optional[List[float]]:
"""None on any failure. An embedding must never delay an agent's reaction."""
Requirements:
- Embed the same dict
build_payloadalready stores. Serialise thefactsdict deterministically (sorted keys) so the embedded text and the storedcontextcannot drift apart. Add a_context_text(facts)helper and test it pure, the waybuild_payloadis tested. - Fail open — return
None, log once, never raise into the agent path. - Small LRU cache: repeated stalls on one booking produce near-identical facts.
- Gate on
EMBEDDINGS_ENABLEDand a registry skill flag, so it can be switched off from Agent Studio without a redeploy (registry.skill_enabled(...)already exists).
Then extend build_payload/record_decision (core/decisions.py:33,53) to carry
context_embedding. Both call sites (dispatch_agent.py:335,
exception_agent.py:456) keep their signatures.
Acceptance: after one stall,
SELECT count(*) FROM agent_decisions WHERE context_embedding IS NOT NULL > 0.
Files
| File | Change |
|---|---|
AI_engine/core/embeddings.py |
new — embed(), _context_text(), LRU cache, fail-open |
AI_engine/core/decisions.py:33 |
modify — build_payload carries context_embedding |
AI_engine/core/decisions.py:53 |
modify — record_decision awaits the embed before posting |
AI_engine/config/system_config.py |
modify — EMBEDDINGS_ENABLED, provider key, model name |
AI_engine/docker-compose.yml |
modify — pass the embedding env through |
AI_engine/tests/test_embeddings.py |
new — _context_text determinism, fail-open returns None |
AI_engine/tests/test_registry_phase5.py:194 |
modify — asserts the record_decision payload shape |
Not touched: agents/dispatch_agent.py:335 and agents/exception_agent.py:456
keep their call signatures — the embedding is added inside decisions.py, so
neither agent changes.
B3. Outcome loop — the part that makes retrieval useful
PATCH /internal/agent-decisions/:id/outcome exists with no caller. Define, per
decision type, what "it worked" means. Starting proposal:
| Type | Outcome = success when |
failure when |
|---|---|---|
stall_response |
booking reaches Delivered within its SLA window after the decision |
SLA breached, or cancelled |
assignment_failure |
booking gets an assignment within N minutes of the decision | still unassigned after N, or cancelled |
Implement as a backend sweeper following the established pattern in
internal/assignment/sweeper.go:88 — ticker, recover() per tick, and the Redis
lock that keeps one replica sweeping (sweeper.go:104). Wire it in main.go
beside go assignment.StartPendingSweeper() (main.go:244).
Two gotchas from the existing code:
- Use
Receivedat-style true instants, notutils.DBNow—models/ai_runs.go:26documents thatDBNowreturns IST digits labelled UTC and is 5h30m off fortimestamptz. The same trap applies tooutcome_recorded_at. - Leave
outcomeNULL while undecided./similaralready treats NULL as "no evidence yet", which is correct.
Acceptance: rows acquire non-NULL outcome within one sweep interval of
their window closing, and /insights decision-outcome counts stop being all
pending.
Files
| File | Change |
|---|---|
doormile_backend/internal/ai/outcomes/sweeper.go |
new — ticker + Redis lock, modelled on internal/assignment/sweeper.go:88 |
doormile_backend/internal/ai/outcomes/rules.go |
new — the per-decision-type success/failure predicates |
doormile_backend/internal/ai/outcomes/sweeper_test.go |
new — interval, window, and both predicates |
doormile_backend/main.go:244 |
modify — go outcomes.StartOutcomeSweeper() beside the pending sweeper |
doormile_backend/models/agentdecision.go |
modify — only if B6 adds Tenantid |
Reference, not modified: internal/assignment/sweeper.go:88-110 (the ticker +
recover() + Redis single-replica lock pattern to copy),
models/ai_runs.go:26 (the utils.DBNow timezone trap to avoid).
B4. Read path — precedent in the prompt
Before the LLM call in exception_agent / dispatch_agent, fetch the top-5
similar resolved decisions and include them as precedent — "the last 5
comparable situations and whether the action worked."
- Feature-flag it on a registry skill so it is switchable from Agent Studio.
- Hard timeout (~300ms) with fail-open to today's prompt. The current behaviour
is the floor; this can only raise it. (
core/registry.pyalready follows this rule;ragRouter.js:9-14documents the same discipline on the console side.) - Keep precedent out of the structured-output schema. It informs the prompt; it must not become a field the model can invent.
Acceptance: an eval run shows the decision quality moving. AI_engine/evals/
already has the harness and cases (stall_cases.jsonl,
assignment_cases.jsonl) — extend those rather than judging by eye.
Files
| File | Change |
|---|---|
AI_engine/core/memory.py |
new — recall(decision_type, embedding, k) → POST /internal/agent-decisions/similar, timeout + fail-open |
AI_engine/core/llm.py:157 |
modify — decide_stall_response accepts optional precedent |
AI_engine/core/llm.py:227 |
modify — decide_assignment_failure accepts optional precedent |
AI_engine/agents/exception_agent.py:456 |
modify — recall before the decide call |
AI_engine/agents/dispatch_agent.py:335 |
modify — recall before the decide call |
AI_engine/tests/test_memory.py |
new — timeout fails open, empty recall changes nothing |
AI_engine/evals/stall_eval.py |
modify — run with and without precedent |
AI_engine/evals/assignment_eval.py |
modify — same |
AI_engine/evals/stall_cases.jsonl |
modify — cases where precedent should change the answer |
AI_engine/evals/assignment_cases.jsonl |
modify — same |
doormile_backend/internal/ai/registry/seed.go:114 |
modify — seed a recall_similar_decisions read tool + the skill flag that gates B4 |
The registry seed row matters: without it the feature cannot be switched off from
Agent Studio, which is the whole point of gating it on registry.skill_enabled.
B5. Hygiene
agent_decisionshas no retention.aiagentrunspurges at 30 days (telemetry/recorder.go:28,128); decisions grow forever, and this is the table retrieval scans. Decide a window — longer than 30 days, since old precedent is the point. Suggest 180 days, or keep resolved rows and purge unresolved ones.- Re-tune
lists = 100. That is right for roughly 100k–1M rows. Below ~10k it over-partitions and recall drops. Checkcount(*)once embeddings flow; consider HNSW instead if the pgvector version supports it.
Files
| File | Change |
|---|---|
doormile_backend/internal/ai/outcomes/sweeper.go |
modify — fold the decision purge into the same tick |
doormile_backend/migrations/migrate.go:98 |
modify — index tuning, once row count is known |
B6. Tenant isolation — decide before B4 ships, not after
agent_decisions has no tenant column. If retrieved precedent crosses
tenants, one client's operational history shapes decisions made for another.
Given that console logins are already unscoped on tenantid NULL
(doormile-console-logins-unscoped), this needs deciding up front.
Recommendation: add tenantid to agent_decisions, have the engine populate it,
and filter in the /similar query. Cheap now, expensive after the table fills.
Files
| File | Change |
|---|---|
doormile_backend/models/agentdecision.go |
modify — add Tenantid *uint64 with an index |
doormile_backend/controllers/agentDecisionController.go:18 |
modify — accept tenant_id on create |
doormile_backend/controllers/agentDecisionController.go:104 |
modify — add AND tenantid = ? to the kNN query |
AI_engine/core/decisions.py:33 |
modify — build_payload carries the tenant |
AI_engine/agents/exception_agent.py · dispatch_agent.py |
modify — source the tenant from the booking facts |
doormile_backend/routes/routes_ai_registry_pg_test.go |
modify — a cross-tenant recall must return nothing |
Nullable, because the engine will not always know the tenant. Decide whether a
NULL tenant row is recallable by everyone or by no one — given
doormile-console-logins-unscoped, by no one is the safer default.
Track C — Kubernetes
C1. AI_engine is not in Kubernetes at all
No manifest, no kustomization, no deploy script, and docker-compose.yml uses
build: . with no registry push. It runs on a VM by compose.
Prerequisite: the engine has no HTTP server in production mode.
main.py --production starts no listener, so there is no liveness/readiness
target and no /metrics. Without it, Kubernetes can only restart on process
exit — a NATS-disconnected engine looks healthy forever.
fastapi and uvicorn are already in requirements.txt. Add a small surface in
production_mode():
GET /healthz— process alive (event loop responsive)GET /readyz— NATS connected andregistry.loadedis trueGET /metrics— optional; decisions recorded, tool calls, LLM failures
Readiness must include registry.loaded, otherwise a pod that cannot reach the
backend serves traffic on env defaults while reporting healthy — exactly the A1
failure mode, invisible again.
Then: push the image to a registry (compose builds locally), and write
manifests/doormile/ai-engine.yaml as a Deployment (it is stateless;
replicas: 1 to start — the agents are not yet idempotent across replicas, see
C6).
Files
| File | Change |
|---|---|
AI_engine/core/health.py |
new — the aiohttp/FastAPI surface (/healthz, /readyz, /metrics) |
AI_engine/main.py:130 |
modify — start the health server inside production_mode() beside registry.run(...) |
AI_engine/main.py (print_help) |
modify — document the health port |
AI_engine/Dockerfile |
modify — EXPOSE the health port |
AI_engine/docker-compose.yml |
modify — publish the port so compose and k8s behave alike |
AI_engine/tests/test_health.py |
new — /readyz is red while registry.loaded is false |
kubernetes/manifests/doormile/ai-engine.yaml |
new — Deployment + Service + probes + resources |
fastapi and uvicorn are already in requirements.txt — no new dependency.
Readiness must check registry.loaded (core/registry.py:45), not just the
process, or A1's failure mode becomes invisible again.
C2. doormile/ has no kustomization.yaml
alaska/, core/ and nearle/ all have one. doormile/ does not, and there is
no deploy-doormile.sh alongside deploy-core-stack.sh / deploy-nearle-stack.sh.
That is why it drifts. Add both.
Files
| File | Change |
|---|---|
kubernetes/manifests/doormile/kustomization.yaml |
new — list miletruth, config, ai-engine, pdb |
kubernetes/deploy-doormile.sh |
new — copy the shape of deploy-nearle-stack.sh |
kubernetes/scripts/sync_manifests.py |
check — may need the new namespace registering |
kubernetes/docs/DEPLOY_CHECKLIST.md |
modify — add the doormile stack |
Pattern to copy: manifests/nearle/kustomization.yaml + deploy-nearle-stack.sh.
C3. The doormile StatefulSet has no probes, resources or PDB
replicas: 3 with no resource requests means the scheduler can stack all
three on one node, and no readinessProbe means a pod receives traffic before
Postgres/Redis/NATS are connected. core/ has worker-pdb.yaml; doormile has
nothing.
Add resources.requests/limits, a readinessProbe and livenessProbe against the
backend's health route, and a PodDisruptionBudget (minAvailable: 2).
Also: a StatefulSet for a stateless Go API is the wrong kind — it gives serial
rollouts and no benefit. Switching to Deployment is low-risk and makes deploys
faster. Not urgent; flagging because it is why rollouts feel slow.
Files
| File | Change |
|---|---|
kubernetes/manifests/doormile/miletruth.yaml:23 |
modify — resources, readinessProbe, livenessProbe |
kubernetes/manifests/doormile/doormile-pdb.yaml |
new — minAvailable: 2, copy manifests/core/worker-pdb.yaml |
No backend change needed — the probe targets already exist and are correct:
GET /api/v1/health (routes/routes.go:46, unauthenticated, always 200) for
liveness, and GET /api/v1/ready (routes/routes.go:50) for readiness, which
already returns 503 when Postgres or Redis is unreachable
(routes.go:68-70). Point the probes at those; do not write new ones.
Note /ready reports Redis GEO status without gating on it — deliberate, per the
comment at routes.go:73. A readinessProbe on /ready therefore will not pull a
pod out of service for a broken rider search, which is the intended behaviour.
C4. No ingress for doormile
nearle and alaska are on manifests/core/ingress-unified.yaml. doormile is
NodePort 30830 plus host nginx (conf/nginx-doormile.conf) — half-migrated. This
also makes B0.2 (GET-with-body) a certainty rather than a risk.
Files
| File | Change |
|---|---|
kubernetes/manifests/core/ingress-unified.yaml |
modify — add a doormile rule (needs a ReferenceGrant if it stays cross-namespace, cf. manifests/nearle/nearle-reference-grant.yaml) |
kubernetes/manifests/doormile/miletruth.yaml:91 |
modify — NodePort → ClusterIP once the ingress serves it |
kubernetes/conf/nginx-doormile.conf |
modify — retire or repoint, only after the ingress is verified |
Do these in that order. Flipping the Service type before the ingress works takes the API offline.
C5. NATS is outside the cluster on two ports
Backend uses nats://66.116.226.161:4223; core-config.yaml:10 uses :4222.
Worth a line in docs/ARCHITECTURE.md on which port is which and why, before the
engine joins and needs to pick one.
C6. Decide replica safety before scaling the engine
The telemetry recorder is already replica-safe (NATS queue group +
uq_aiagentruns_agent_task). The agents are not obviously so: exception_agent
has an in-process TTL dedup set (exception_agent.py:339), which is per-pod. Two
engine replicas would each decide on the same stall. Keep replicas: 1 until
dedup moves to Redis.
Files (only if scaling past 1 replica)
| File | Change |
|---|---|
AI_engine/agents/exception_agent.py:339 |
modify — TTL set → Redis SET NX EX |
AI_engine/tests/test_stall_dedup.py |
modify — covers the current in-process behaviour |
kubernetes/manifests/doormile/ai-engine.yaml |
modify — raise replicas |
C5 (the NATS port note) is documentation only: kubernetes/docs/ARCHITECTURE.md.
Track D — Executor backlog
9 of 22 seeded tools are marked REVIEW ONLY … No executor in
internal/ai/registry/seed.go. The registry is honest about it and the console
renders them disabled. This is the feature list, in value order.
D1. assign_riders — needs an admin auto-assign route (not just auth)
Correcting an earlier assumption: this is not a free auth fix.
actions.js:96-107 already explains why — POST /admin/bookings/:id/assign-miler
(routes.go:413 → adminController.go:2965) requires a chosen rider per
booking ({mileruserid}), and a finding does not pick one. The hub route that
does pick (POST /hub/bookings/:id/auto-assign, routes.go:520) is behind
HubStaffAuth and 403s for every console login.
The cleanest route is an admin batch-assign, better than the per-booking
auto-assign first considered. HubBatchAssign (controllers/hubController.go:1963)
already does exactly what a finding needs: it takes bookingids[], picks riders
via Redis GEO + scoring, and commits server-side in one call. Its only hub-specific
parts — c.Locals("hubid") and hubPincodePrefix(hubID) (hubController.go:1964-1966)
— are used solely as a fallback when bookingids is empty (hubController.go:1981-1985).
The console always passes explicit ids, so that branch never runs.
So: extract the body into a shared helper taking (bookingIDs, capPerRider, actorID, scopeFn)
and have both the hub route and a new admin route call it. scopeBookingsToOwnTenant
(hubController.go:1986) already works for admin logins.
The console side is then nearly free — batchAssignBookings already exists at
src/api/doormile/endpoints.js:554 and already sends {bookingids, max_per_rider}.
It just points at the hub URL that 403s. One URL change.
Option (b), having the skill pick a rider via nearby_milers and call the
existing assign-miler, is worse: more console work and it puts solver logic in
the browser.
Files
| File | Change |
|---|---|
doormile_backend/controllers/hubController.go:1963 |
modify — extract the shared assign helper out of HubBatchAssign |
doormile_backend/controllers/adminController.go |
add AdminBatchAssign calling that helper |
doormile_backend/routes/routes.go:413 |
modify — register adminAuth.Post("/bookings/batch-assign", …) |
krow_talent_app/src/api/doormile/endpoints.js:554 |
modify — /hub/bookings/batch-assign → /admin/bookings/batch-assign |
krow_talent_app/src/lib/assistant/agent/actions.js:122 |
modify — add the assignMiler executor |
krow_talent_app/src/lib/assistant/agent/actions.js:96-107 |
modify — delete the "deliberately NOT an executor" note |
krow_talent_app/tests/lib/agentActions.test.js |
modify — asserts the current executor set |
doormile_backend/internal/ai/registry/seed.go |
modify — assign_riders description stops saying REVIEW ONLY |
Check before starting: endpoints.js:547-553 warns that Doormile-native batch
assign commits with no preview/reconcile step and leaves multi-stop riders
unsequenced. An agent-proposed assignment firing straight to commit is a
behaviour decision, not just a wiring one — confirm that is wanted.
This takes the console from 1 working verb to 2 and makes the highest-severity SLA finding actionable.
D2. alert_low_battery_rider — nearly free
seed.go already targets POST /admin/milers/:id/notify, which is the endpoint
notify_riders already uses successfully. This is a message-text change and an
EXECUTORS entry, not a new capability.
Files
| File | Change |
|---|---|
krow_talent_app/src/lib/assistant/agent/actions.js:122 |
modify — add the alertLowBatteryRider executor |
krow_talent_app/src/lib/assistant/skills/definitions/RiderBatterySafetySkill.js |
check — confirm the proposal carries milerId |
krow_talent_app/tests/lib/agentActions.test.js |
modify — same assertion as D1 |
doormile_backend/internal/ai/registry/seed.go |
modify — drop REVIEW ONLY from the description |
No backend change. notifyMiler already exists at
src/api/doormile/endpoints.js:389 → POST /admin/milers/:id/notify, which is
the endpoint seed.go already names as the target.
D3. The rest need endpoints that do not exist
enforce_otp_verification, dispatch_hub_idle_parcels, trigger_auto_dispatch,
enforce_cash_handoff, rebalance_riders — all marked Target: "none yet".
Each is a product decision first. Not in this phase.
Two more loose ends
src/lib/agentNetwork.jsis a hand-maintained snapshot dated 16–20 Sep, and the file says so honestly. Once C1 gives the engine an HTTP surface, addGET /agents/statusand swap the source. The file is deliberately shaped like that response, so it is a change of source, not a rewrite.src/lib/assistant/ragRouter.jspoints at aservices/aisidecar that does not exist in any repo.VITE_AI_URLappears nowhere, soisRagEnabled()is permanently false and the module is dead code. Do not conflate this with Track B — ragRouter is semantic intent routing for the console assistant, not decision memory. Lower value. Leave it dormant (it is correctly fail-open) or decide to build the sidecar as its own piece of work.
Consolidated file manifest
14 new files, 46 modified, across 4 repos. Per-item detail is in the tracks above.
doormile_backend — 4 new, 17 modified
| File | New? | Items |
|---|---|---|
internal/ai/outcomes/sweeper.go |
new | B3, B5 |
internal/ai/outcomes/rules.go |
new | B3 |
internal/ai/outcomes/sweeper_test.go |
new | B3 |
migrations/migrate.go |
B0.1 (:92), B5 (:98) | |
routes/routes.go |
B0.2 (:594), D1 (:413) | |
controllers/agentDecisionController.go |
B0.2 (:74), B6 (:18, :104) | |
controllers/hubController.go |
D1 (:1963 — extract helper) | |
controllers/adminController.go |
D1 (add AdminBatchAssign) |
|
models/agentdecision.go |
B6 | |
internal/ai/registry/seed.go |
B4 (:114), D1, D2 | |
main.go |
B3 (:244) | |
routes/routes_ai_registry_pg_test.go |
B0.2, B6 |
AI_engine — 5 new, 20 modified
| File | New? | Items |
|---|---|---|
core/embeddings.py |
new | B2 |
core/memory.py |
new | B4 |
core/health.py |
new | C1 |
tests/test_embeddings.py |
new | B2 |
tests/test_memory.py · tests/test_health.py |
new | B2, C1 |
core/decisions.py |
B2 (:33, :53), B6 | |
core/llm.py |
A4 (:11, :22), B4 (:157, :227) | |
core/registry.py |
— read-only (loaded consumed by C1) |
|
agents/exception_agent.py |
B4 (:456), B6, C6 (:339) | |
agents/dispatch_agent.py |
B4 (:335), B6 | |
main.py |
C1 (:130, print_help) |
|
config/system_config.py |
B2 | |
requirements.txt · Dockerfile · docker-compose.yml |
A3, A4, C1 | |
evals/stall_eval.py · assignment_eval.py · both .jsonl |
B4 | |
tests/test_registry_phase5.py · test_llm.py · test_stall_dedup.py |
B2, A4, C6 |
kubernetes — 5 new, 7 modified
| File | New? | Items |
|---|---|---|
manifests/doormile/doormile-config.yaml |
new | A1 |
manifests/doormile/ai-engine.yaml |
new | C1, C6 |
manifests/doormile/kustomization.yaml |
new | C2 |
manifests/doormile/doormile-pdb.yaml |
new | C3 |
deploy-doormile.sh |
new | C2 |
manifests/doormile/miletruth.yaml |
A1, A2, C3 (:23), C4 (:91) | |
manifests/core/ingress-unified.yaml |
C4 | |
conf/nginx-doormile.conf |
C4 (last) | |
scripts/sync_manifests.py |
C2 | |
docs/DEPLOY.md · DEPLOY_CHECKLIST.md · ARCHITECTURE.md |
A2, C2, C5 |
krow_talent_app — 0 new, 4 modified
The console barely changes. Everything it needs already exists.
| File | Items |
|---|---|
src/api/doormile/endpoints.js |
D1 (:554 — one URL) |
src/lib/assistant/agent/actions.js |
D1 (:96-107, :122), D2 (:122) |
tests/lib/agentActions.test.js |
D1, D2 |
src/lib/assistant/skills/definitions/RiderBatterySafetySkill.js |
D2 — check only |
Untouched on purpose: src/lib/agentNetwork.js (until C1 ships an endpoint) and
src/lib/assistant/ragRouter.js (dormant, out of scope — see loose ends).
Files deliberately not touched
src/lib/src/utilsduplicate trees — both have live importers, do not consolidate.AI_engine/core/tool_registry.py— the 2-vs-22 tool gap is a separate decision, not this phase.AI_engine/customer_portal/,dashboard/— not on any path this phase touches.
Decisions needed before coding
- Embedding provider — OpenAI 1536 (no column change), Voyage 1024, or local 384? (B1)
- Is
pgvectorinstalled onlogistics? If creating extensions needs an admin, that is a prerequisite, not a migration. (B0.1) - Outcome definitions — are the two in B3 right, and what is N for
assignment_failure? tenantidonagent_decisions— add it now, or accept cross-tenant precedent? (B6)assign_riders— route (a) admin auto-assign, or (b) console picks the rider? (D1)- Decision retention window — 180 days, or keep-resolved-purge-unresolved? (B5)
Sequencing
A1 ──> A3, A4 ──┬──> B0 ──> B2 ──> B3 ──> B4 ──> B5, B6
│
└──> C1(health) ──> C1(manifest) ──> C2 ──> C3 ──> C4
A2 (independent, before A1 adds more secrets to git)
D1, D2 (independent of everything above)
A1 first and alone. Until the internal key is set, the engine is not talking
to the backend, so every Track B acceptance check would fail for the wrong
reason. Run the kubectl check in A1 before writing any code — if git does not
match the cluster, that changes the shape of Track C.
Cheapest real step up a level, once A1 is in: D1 + D2. Two executors, no new LLM work, and it moves autonomy off the floor.
Standing constraints
- Nothing in this plan is committed, pushed, or deployed without being asked.
- No migrations run against a real database without being asked.
- No files deleted or untracked without per-file approval.
src/libandsrc/utilsin the console both have live importers — neither tree is dead, do not consolidate them as part of this work.