# Doormile Agent Platform — Phase 7: Decision Memory, Deployment, Executors Status: **mostly implemented 2026-10-08 (uncommitted, not deployed, no migration run)** Written 2026-10-08 · Scope: `krow_talent_app`, `doormile_backend`, `AI_engine`, `kubernetes` --- ## Status (2026-10-08) Verified: backend `go build`/`go vet` clean, 20 packages pass · engine 148 tests (3 pre-existing import errors) · console 1397 pass (13 pre-existing `agentStudio.test.jsx` failures) · `kubectl kustomize manifests/doormile` renders clean. **Nothing committed. No migration has run. No SQL in this work has ever executed against a database.** | Item | State | |---|---| | A1 config/secrets manifest, `envFrom`, probes, resources | **built** — placeholders unreplaced | | A2 secrets out of git | **not done** — documented in `manifests/doormile/SECRETS.md`, including the kustomize-clobbers-hand-created-secrets trap. Five more placeholder keys were added. | | A3 `anthropic>=0.49.0` · A4 `claude-opus-5-5` | **done** | | B0.1 `CREATE EXTENSION vector` · B0.2 `/similar` → POST | **done** | | B1 provider (OpenAI `text-embedding-3-small`, 1536) | **done** — my call, flagged | | B2 `core/embeddings.py` + decision wiring | **done** | | B3 outcome sweeper + rules + 14 tests | **done** | | B4 `core/memory.py` + both agents | **done** | | B5 retention/prune | **done** — ivfflat `lists` re-tune still needs a row count | | B6 `tenantid` scoping | **done** | | B2.5 historical backfill | **done** — `internal/ai/outcomes/backfill.go` | | B7 persist console findings | **done** — `aiskillfindings`, 4 routes, `findingReport.js`, 15 tests | | C1 health surface + `ai-engine.yaml` + Dockerfile/compose port | **done** — image not built or pushed | | C2 kustomization + `deploy-doormile.sh` | **done** | | C3 probes/resources/PDB | **done** — StatefulSet→Deployment not done | | C4 ingress | **built** — apply order matters, see the file header | | C5 NATS port note | **done** — `docs/ARCHITECTURE.md` | | C6 engine replica safety | **see correction 1 — was already safe for a different reason** | | D1 admin batch-assign + console executor | **done** | | D2 `alert_low_battery_rider` | **done** | | D3 five review-only tools | **1 of 5 done** — `trigger_auto_dispatch` turned out to BE batch-assign under another name. The other four need endpoints that do not exist. | | Agents-page snapshot → live endpoint | **backend done** — engine serves `GET /agents/status`; the console still reads the snapshot, so the swap is one fetch away | ## Two things this plan got wrong **Correction 1 — C6. The engine's stall dedup was never per-pod.** This plan claimed `ExceptionAgent` kept its dedup in an in-process TTL set and that moving it to Redis was the prerequisite for scaling out. Wrong: `_claim_stall` / `_claim_stall_handled` (`agents/exception_agent.py:336`) have been Redis `SET NX` with a TTL all along — cross-pod, self-expiring, durable across restarts. The comment there saying "replaces the old unbounded in-memory set" describes what it REPLACED; I read it as current state. The real blocker is different and still real: `core/message_bus.py:284` and `:312` bind **durable push consumers with fixed names**, and a durable push consumer admits one active subscriber unless created with a deliver group. A second replica does not duplicate work — it fails to bind. Scaling out means giving those subscriptions a deliver group (or moving to pull consumers). `replicas: 1` stands, for the corrected reason. **Correction 2 — the registry was already honest about the simulated agents.** This plan said the three fictional-data agents were "presented as live" and should be marked. They already were: `seed.go` has `StatusSimulation` for `HUB_AGENT`, `FLEET_AGENT` and `ROUTE_OPTIMIZER`, with purposes reading "an in-memory fleet of 19 fake vehicles" and "8 hard-coded fictional hubs". The console snapshot (`src/lib/agentNetwork.js`) was also honest about activity ("Zero tasks received", `publishes: 'Nothing on the bus.'`, `api: 'No.'`) but silent on the data being invented — so the three entries now say so, and the two surfaces agree. I overstated that surface as "telling anyone something untrue". It was incomplete, not false. --- ## Track A — Unblock (do this first; nothing else matters until it lands) ### A1. `INTERNAL_API_KEY` is missing from the cluster — **everything engine↔backend is 401** `middlewares/internal_auth.go:13` reads `INTERNAL_API_KEY` and **fails closed when unset**: ```go expected := os.Getenv("INTERNAL_API_KEY") if expected == "" || c.Get("X-Internal-Key") != expected { ``` `kubernetes/manifests/doormile/miletruth.yaml` never sets it. The engine's `docker-compose.yml` *does* pass it. So the engine sends a key the backend rejects, and in-cluster **every** `/api/v1/internal/*` call 401s: - `GET /internal/ai/registry` — the registry poll (this is why "engine reads registry in prod" was never confirmable) - `POST /internal/agent-decisions` — the decision log, so Insights sees nothing - `GET /internal/express/riders`, `/express/bookings`, `POST /express/assign` - `POST /internal/bookings/:id/reassign`, `POST /internal/notify` **Step 1 — find out whether git matches reality.** The `kubernetes` repo's log is full of `Restore…` / `Recreate…` commits, so the manifest may already be fiction: ```bash kubectl -n doormile get statefulset doormile -o jsonpath='{range .spec.template.spec.containers[0].env[*]}{.name}{"\n"}{end}' | sort ``` `config/config.go` reads 29 env vars. If that list is much shorter, the live cluster has hand-applied drift and **the next `kubectl apply -f` wipes it.** **Step 2 — close the gap in git, not by hand.** Add to `doormile-secrets` (`stringData`) and reference from the StatefulSet. Absent from the manifest today and read by `config.go`: `INTERNAL_API_KEY`, `JWT_SECRET_KEY`, `APP_PORT`, `ENV`, `TRUSTED_PROXIES`, `AI_LAYER_BASE_URL`, `ROUTE_OPTIMIZER_URL`, `GEOCODER_URL`, `GEOCODER_EMAIL`, `SMTP_HOST`, `SMTP_PORT`, `SMTP_USER`, `SMTP_PASSWORD`, `SMTP_FROM`, `CLIENT_ONBOARDING_OWNERS`, `PLAYGROUND_LLM_API_KEY`, `PLAYGROUND_LLM_BASE_URL`, `PLAYGROUND_LLM_MODEL`, `REDIS_USER`. Secrets vs ConfigMap: `INTERNAL_API_KEY`, `JWT_SECRET_KEY`, `SMTP_PASSWORD`, `PLAYGROUND_LLM_API_KEY` are secrets. The rest belong in a `doormile-config` ConfigMap so a URL change is not a secret edit. **Files** | File | Change | |---|---| | `kubernetes/manifests/doormile/miletruth.yaml` | modify — add the 19 env vars to the StatefulSet + `doormile-secrets` | | `kubernetes/manifests/doormile/doormile-config.yaml` | **new** — ConfigMap for the non-secret vars | | `AI_engine/.env` (server-local, untracked content) | modify — same `INTERNAL_API_KEY` value | Read-only, no change: `doormile_backend/middlewares/internal_auth.go:13`, `doormile_backend/config/config.go:60-102` (the contract being satisfied). > `INTERNAL_API_KEY` must be **byte-identical** in the backend Secret and the > engine's env. Generate once (`openssl rand -hex 32`), set both. **Acceptance:** from inside the cluster, `curl -H "X-Internal-Key: $KEY" http://doormile-service.doormile:8081/api/v1/internal/ai/registry` returns 200 with a body and an `ETag`; a second call with `If-None-Match` returns 304. Backend logs show `ai playground: enabled`. `GET /admin/ai/status` reports the engine as following the registry (`aiRegistryController.go:208` counts the poll). ### A2. Plaintext secrets are committed `miletruth.yaml:19-21` has `DB_PASSWORD`, `REDIS_PASSWORD`, `NATS_PASSWORD` as literal in git (the same password reused for all three). Same class as the committed Firebase keys from the CX audit. Pick one and apply it before A1 adds *more* secrets to that file: Sealed Secrets, SOPS, or an out-of-git `kubectl create secret` with the manifest holding only a reference. Do **not** delete or untrack anything here without per-file approval — see the standing rule. Rotating that shared password is a separate decision; this item is only about stopping new secrets entering git. **Files** | File | Change | |---|---| | `kubernetes/manifests/doormile/miletruth.yaml` | modify — `stringData` block becomes a reference | | `kubernetes/manifests/doormile/sealed-secrets.yaml` | **new** — if Sealed Secrets is the chosen route | | `kubernetes/docs/DEPLOY.md` | modify — document how secrets are supplied now | Same pattern already in git at `kubernetes/manifests/core/core-secrets.yaml` and `manifests/nearle/nearle-secrets.yaml` — whatever is chosen should cover those too, but that is outside this phase. ### A3. `anthropic>=0.21.0` floor is wrong `AI_engine/requirements.txt` pins `anthropic>=0.21.0`, but `core/llm.py:117` sends `thinking: {"type": "adaptive"}` and `output_config.effort` — parameters that old SDK does not know. There is no lockfile, so a fresh `pip install` in a rebuilt image can resolve to something that 400s on every LLM call. Bump to a current floor and pin the image build. `openai>=1.0.0` is already present (unused today — Track B uses it). **Files** | File | Change | |---|---| | `AI_engine/requirements.txt` | modify — raise the `anthropic` floor | | `AI_engine/Dockerfile` | modify — optional: `pip install` from a lockfile instead | | `AI_engine/requirements-dev.txt` | check — may pin the same packages | ### A4. `LLM_MODEL` default is a generation behind `core/llm.py:22` defaults to `claude-opus-4-8` ($5/$25 per MTok). `claude-opus-5-5` is both cheaper ($4/$20) and more capable. One-line change; the registry's per-agent `model` pin already overrides it (`registry.model(agent_id)`), so this only moves the floor. **Files** | File | Change | |---|---| | `AI_engine/core/llm.py:22` | modify — default `LLM_MODEL` | | `AI_engine/core/llm.py:11` | modify — the docstring naming the default | | `AI_engine/docker-compose.yml:45` | modify — `LLM_MODEL` fallback | | `AI_engine/tests/test_llm.py` | check — may assert the old default | Note `core/llm.py:109` special-cases Haiku 4.5 (rejects adaptive thinking and `effort`). Leave that branch alone; it is correct. --- ## Track B — RAG decision memory The skeleton exists and is wired to nothing. Four pieces are already built: | Piece | Where | |---|---| | `context_embedding vector(1536)` on `agent_decisions` | `migrations/migrate.go:92` | | ivfflat cosine index, `lists = 100` | `migrations/migrate.go:98` | | `POST /internal/agent-decisions` accepts `context_embedding` | `controllers/agentDecisionController.go:23` | | cosine kNN query + `PATCH /:id/outcome` | `agentDecisionController.go:74,120` | Nothing produces an embedding, nothing calls `/similar`, nothing records an outcome. There are exactly **two** decision types to cover — `assignment_failure` (`dispatch_agent.py:335`) and `stall_response` (`exception_agent.py:456`). ### B0. Three defects to fix before writing any new code **B0.1 — `CREATE EXTENSION vector` appears nowhere in the repo.** `migrate.go:92` runs `ALTER TABLE agent_decisions ADD COLUMN … vector(1536)` and logs failure *non-fatally*. If the extension is not installed on `logistics`, both the column and the index silently fail and every retrieval 500s. ```sql SELECT extname, extversion FROM pg_extension WHERE extname = 'vector'; ``` If absent, add **before** line 92 in `Migrate()`: ```go if res := db.Exec(`CREATE EXTENSION IF NOT EXISTS vector`); res.Error != nil { utils.Error("❌ pgvector extension unavailable — decision memory disabled", "error", res.Error) } ``` Needs the `pgvector` extension available on the server and a role with rights to create it. On a managed Postgres this may be an admin action, not a migration — check before assuming. **B0.2 — `/similar` is a `GET` that requires a JSON body.** `routes.go:594` registers it as `GET`, and `agentDecisionController.go:84` calls `BodyParser` demanding an `embedding` array. nginx and most HTTP clients drop GET bodies — and doormile is served *through* host nginx (see C4). Change to: ```go internal.Post("/agent-decisions/similar", controllers.FindSimilarDecisions) ``` No caller exists yet, so this breaks nothing. **B0.3 — `/similar` filters `WHERE outcome IS NOT NULL`, and nothing writes outcomes.** Even with embeddings flowing it returns zero rows forever. **B2 is not optional** — it is the half that makes retrieval worth anything. **Files** | File | Change | |---|---| | `doormile_backend/migrations/migrate.go:92` | modify — add `CREATE EXTENSION` above the `ALTER TABLE` | | `doormile_backend/routes/routes.go:594` | modify — `internal.Get` → `internal.Post` for `/agent-decisions/similar` | | `doormile_backend/controllers/agentDecisionController.go:74` | modify — comment the method change; body parsing already correct | | `doormile_backend/routes/routes_ai_registry_pg_test.go` | modify — add coverage for the POST shape | ### B1. Embedding provider — decision required Anthropic has no embeddings endpoint, so this needs a second provider. | Option | Dims | Column change | Notes | |---|---|---|---| | **OpenAI `text-embedding-3-small`** | 1536 | **none** | ~$0.02/MTok. `openai>=1.0.0` already a dependency. Column was sized for it. | | Voyage `voyage-3` | 1024 | yes | Anthropic-recommended; new key, new vendor. | | Local `bge-small` / `MiniLM` | 384 | yes | Free, no egress, no key. +~400MB RAM, model in image, slower cold start. | **Recommendation: OpenAI `text-embedding-3-small`.** The schema already matches, the dependency is already there, and the second-provider line is already crossed (the playground runs on Groq). Revisit if data egress is a constraint — then take the local model and migrate the column to `vector(384)`. ### B2. Write path — `AI_engine/core/embeddings.py` One function, modelled on `core/decisions.py`'s fire-and-forget discipline: ```python async def embed(text: str) -> Optional[List[float]]: """None on any failure. An embedding must never delay an agent's reaction.""" ``` Requirements: - **Embed the same dict `build_payload` already stores.** Serialise the `facts` dict deterministically (sorted keys) so the embedded text and the stored `context` cannot drift apart. Add a `_context_text(facts)` helper and test it pure, the way `build_payload` is tested. - Fail open — return `None`, log once, never raise into the agent path. - Small LRU cache: repeated stalls on one booking produce near-identical facts. - Gate on `EMBEDDINGS_ENABLED` **and** a registry skill flag, so it can be switched off from Agent Studio without a redeploy (`registry.skill_enabled(...)` already exists). Then extend `build_payload`/`record_decision` (`core/decisions.py:33,53`) to carry `context_embedding`. Both call sites (`dispatch_agent.py:335`, `exception_agent.py:456`) keep their signatures. **Acceptance:** after one stall, `SELECT count(*) FROM agent_decisions WHERE context_embedding IS NOT NULL` > 0. **Files** | File | Change | |---|---| | `AI_engine/core/embeddings.py` | **new** — `embed()`, `_context_text()`, LRU cache, fail-open | | `AI_engine/core/decisions.py:33` | modify — `build_payload` carries `context_embedding` | | `AI_engine/core/decisions.py:53` | modify — `record_decision` awaits the embed before posting | | `AI_engine/config/system_config.py` | modify — `EMBEDDINGS_ENABLED`, provider key, model name | | `AI_engine/docker-compose.yml` | modify — pass the embedding env through | | `AI_engine/tests/test_embeddings.py` | **new** — `_context_text` determinism, fail-open returns `None` | | `AI_engine/tests/test_registry_phase5.py:194` | modify — asserts the `record_decision` payload shape | Not touched: `agents/dispatch_agent.py:335` and `agents/exception_agent.py:456` keep their call signatures — the embedding is added inside `decisions.py`, so neither agent changes. ### B3. Outcome loop — the part that makes retrieval useful `PATCH /internal/agent-decisions/:id/outcome` exists with no caller. Define, per decision type, what "it worked" means. Starting proposal: | Type | Outcome = `success` when | `failure` when | |---|---|---| | `stall_response` | booking reaches `Delivered` within its SLA window after the decision | SLA breached, or cancelled | | `assignment_failure` | booking gets an assignment within N minutes of the decision | still unassigned after N, or cancelled | Implement as a backend sweeper following the **established pattern** in `internal/assignment/sweeper.go:88` — ticker, `recover()` per tick, and the Redis lock that keeps one replica sweeping (`sweeper.go:104`). Wire it in `main.go` beside `go assignment.StartPendingSweeper()` (`main.go:244`). Two gotchas from the existing code: - Use `Receivedat`-style true instants, not `utils.DBNow` — `models/ai_runs.go:26` documents that `DBNow` returns IST digits labelled UTC and is 5h30m off for `timestamptz`. The same trap applies to `outcome_recorded_at`. - Leave `outcome` NULL while undecided. `/similar` already treats NULL as "no evidence yet", which is correct. **Acceptance:** rows acquire non-NULL `outcome` within one sweep interval of their window closing, and `/insights` decision-outcome counts stop being all `pending`. **Files** | File | Change | |---|---| | `doormile_backend/internal/ai/outcomes/sweeper.go` | **new** — ticker + Redis lock, modelled on `internal/assignment/sweeper.go:88` | | `doormile_backend/internal/ai/outcomes/rules.go` | **new** — the per-decision-type success/failure predicates | | `doormile_backend/internal/ai/outcomes/sweeper_test.go` | **new** — interval, window, and both predicates | | `doormile_backend/main.go:244` | modify — `go outcomes.StartOutcomeSweeper()` beside the pending sweeper | | `doormile_backend/models/agentdecision.go` | modify — only if B6 adds `Tenantid` | Reference, not modified: `internal/assignment/sweeper.go:88-110` (the ticker + `recover()` + Redis single-replica lock pattern to copy), `models/ai_runs.go:26` (the `utils.DBNow` timezone trap to avoid). ### B4. Read path — precedent in the prompt Before the LLM call in `exception_agent` / `dispatch_agent`, fetch the top-5 similar **resolved** decisions and include them as precedent — "the last 5 comparable situations and whether the action worked." - Feature-flag it on a registry skill so it is switchable from Agent Studio. - Hard timeout (~300ms) with fail-open to today's prompt. The current behaviour is the floor; this can only raise it. (`core/registry.py` already follows this rule; `ragRouter.js:9-14` documents the same discipline on the console side.) - Keep precedent **out** of the structured-output schema. It informs the prompt; it must not become a field the model can invent. **Acceptance:** an eval run shows the decision quality moving. `AI_engine/evals/` already has the harness and cases (`stall_cases.jsonl`, `assignment_cases.jsonl`) — extend those rather than judging by eye. **Files** | File | Change | |---|---| | `AI_engine/core/memory.py` | **new** — `recall(decision_type, embedding, k)` → POST `/internal/agent-decisions/similar`, timeout + fail-open | | `AI_engine/core/llm.py:157` | modify — `decide_stall_response` accepts optional precedent | | `AI_engine/core/llm.py:227` | modify — `decide_assignment_failure` accepts optional precedent | | `AI_engine/agents/exception_agent.py:456` | modify — recall before the decide call | | `AI_engine/agents/dispatch_agent.py:335` | modify — recall before the decide call | | `AI_engine/tests/test_memory.py` | **new** — timeout fails open, empty recall changes nothing | | `AI_engine/evals/stall_eval.py` | modify — run with and without precedent | | `AI_engine/evals/assignment_eval.py` | modify — same | | `AI_engine/evals/stall_cases.jsonl` | modify — cases where precedent should change the answer | | `AI_engine/evals/assignment_cases.jsonl` | modify — same | | `doormile_backend/internal/ai/registry/seed.go:114` | modify — seed a `recall_similar_decisions` read tool + the skill flag that gates B4 | The registry seed row matters: without it the feature cannot be switched off from Agent Studio, which is the whole point of gating it on `registry.skill_enabled`. ### B5. Hygiene - **`agent_decisions` has no retention.** `aiagentruns` purges at 30 days (`telemetry/recorder.go:28,128`); decisions grow forever, and this is the table retrieval scans. Decide a window — longer than 30 days, since old precedent is the point. Suggest 180 days, or keep resolved rows and purge unresolved ones. - **Re-tune `lists = 100`.** That is right for roughly 100k–1M rows. Below ~10k it over-partitions and recall drops. Check `count(*)` once embeddings flow; consider HNSW instead if the pgvector version supports it. **Files** | File | Change | |---|---| | `doormile_backend/internal/ai/outcomes/sweeper.go` | modify — fold the decision purge into the same tick | | `doormile_backend/migrations/migrate.go:98` | modify — index tuning, once row count is known | ### B6. Tenant isolation — decide before B4 ships, not after `agent_decisions` has **no tenant column**. If retrieved precedent crosses tenants, one client's operational history shapes decisions made for another. Given that console logins are already unscoped on `tenantid` NULL (`doormile-console-logins-unscoped`), this needs deciding up front. Recommendation: add `tenantid` to `agent_decisions`, have the engine populate it, and filter in the `/similar` query. Cheap now, expensive after the table fills. **Files** | File | Change | |---|---| | `doormile_backend/models/agentdecision.go` | modify — add `Tenantid *uint64` with an index | | `doormile_backend/controllers/agentDecisionController.go:18` | modify — accept `tenant_id` on create | | `doormile_backend/controllers/agentDecisionController.go:104` | modify — add `AND tenantid = ?` to the kNN query | | `AI_engine/core/decisions.py:33` | modify — `build_payload` carries the tenant | | `AI_engine/agents/exception_agent.py` · `dispatch_agent.py` | modify — source the tenant from the booking facts | | `doormile_backend/routes/routes_ai_registry_pg_test.go` | modify — a cross-tenant recall must return nothing | Nullable, because the engine will not always know the tenant. Decide whether a NULL tenant row is recallable by everyone or by no one — given `doormile-console-logins-unscoped`, **by no one** is the safer default. --- ## Track C — Kubernetes ### C1. `AI_engine` is not in Kubernetes at all No manifest, no kustomization, no deploy script, and `docker-compose.yml` uses `build: .` with no registry push. It runs on a VM by compose. **Prerequisite: the engine has no HTTP server in production mode.** `main.py --production` starts no listener, so there is no liveness/readiness target and no `/metrics`. Without it, Kubernetes can only restart on process exit — a NATS-disconnected engine looks healthy forever. `fastapi` and `uvicorn` are already in `requirements.txt`. Add a small surface in `production_mode()`: - `GET /healthz` — process alive (event loop responsive) - `GET /readyz` — NATS connected **and** `registry.loaded` is true - `GET /metrics` — optional; decisions recorded, tool calls, LLM failures Readiness must include `registry.loaded`, otherwise a pod that cannot reach the backend serves traffic on env defaults while reporting healthy — exactly the A1 failure mode, invisible again. Then: push the image to a registry (compose builds locally), and write `manifests/doormile/ai-engine.yaml` as a `Deployment` (it is stateless; `replicas: 1` to start — the agents are not yet idempotent across replicas, see C6). **Files** | File | Change | |---|---| | `AI_engine/core/health.py` | **new** — the aiohttp/FastAPI surface (`/healthz`, `/readyz`, `/metrics`) | | `AI_engine/main.py:130` | modify — start the health server inside `production_mode()` beside `registry.run(...)` | | `AI_engine/main.py` (`print_help`) | modify — document the health port | | `AI_engine/Dockerfile` | modify — `EXPOSE` the health port | | `AI_engine/docker-compose.yml` | modify — publish the port so compose and k8s behave alike | | `AI_engine/tests/test_health.py` | **new** — `/readyz` is red while `registry.loaded` is false | | `kubernetes/manifests/doormile/ai-engine.yaml` | **new** — Deployment + Service + probes + resources | `fastapi` and `uvicorn` are already in `requirements.txt` — no new dependency. Readiness must check `registry.loaded` (`core/registry.py:45`), not just the process, or A1's failure mode becomes invisible again. ### C2. `doormile/` has no `kustomization.yaml` `alaska/`, `core/` and `nearle/` all have one. `doormile/` does not, and there is no `deploy-doormile.sh` alongside `deploy-core-stack.sh` / `deploy-nearle-stack.sh`. That is *why* it drifts. Add both. **Files** | File | Change | |---|---| | `kubernetes/manifests/doormile/kustomization.yaml` | **new** — list miletruth, config, ai-engine, pdb | | `kubernetes/deploy-doormile.sh` | **new** — copy the shape of `deploy-nearle-stack.sh` | | `kubernetes/scripts/sync_manifests.py` | check — may need the new namespace registering | | `kubernetes/docs/DEPLOY_CHECKLIST.md` | modify — add the doormile stack | Pattern to copy: `manifests/nearle/kustomization.yaml` + `deploy-nearle-stack.sh`. ### C3. The `doormile` StatefulSet has no probes, resources or PDB `replicas: 3` with **no resource requests** means the scheduler can stack all three on one node, and **no readinessProbe** means a pod receives traffic before Postgres/Redis/NATS are connected. `core/` has `worker-pdb.yaml`; doormile has nothing. Add `resources.requests`/`limits`, a readinessProbe and livenessProbe against the backend's health route, and a PodDisruptionBudget (`minAvailable: 2`). Also: a `StatefulSet` for a stateless Go API is the wrong kind — it gives serial rollouts and no benefit. Switching to `Deployment` is low-risk and makes deploys faster. Not urgent; flagging because it is why rollouts feel slow. **Files** | File | Change | |---|---| | `kubernetes/manifests/doormile/miletruth.yaml:23` | modify — `resources`, `readinessProbe`, `livenessProbe` | | `kubernetes/manifests/doormile/doormile-pdb.yaml` | **new** — `minAvailable: 2`, copy `manifests/core/worker-pdb.yaml` | **No backend change needed** — the probe targets already exist and are correct: `GET /api/v1/health` (`routes/routes.go:46`, unauthenticated, always 200) for liveness, and `GET /api/v1/ready` (`routes/routes.go:50`) for readiness, which already returns **503** when Postgres or Redis is unreachable (`routes.go:68-70`). Point the probes at those; do not write new ones. Note `/ready` reports Redis GEO status without gating on it — deliberate, per the comment at `routes.go:73`. A readinessProbe on `/ready` therefore will not pull a pod out of service for a broken rider search, which is the intended behaviour. ### C4. No ingress for `doormile` `nearle` and `alaska` are on `manifests/core/ingress-unified.yaml`. `doormile` is NodePort 30830 plus host nginx (`conf/nginx-doormile.conf`) — half-migrated. This also makes B0.2 (GET-with-body) a certainty rather than a risk. **Files** | File | Change | |---|---| | `kubernetes/manifests/core/ingress-unified.yaml` | modify — add a `doormile` rule (needs a ReferenceGrant if it stays cross-namespace, cf. `manifests/nearle/nearle-reference-grant.yaml`) | | `kubernetes/manifests/doormile/miletruth.yaml:91` | modify — `NodePort` → `ClusterIP` once the ingress serves it | | `kubernetes/conf/nginx-doormile.conf` | modify — retire or repoint, **only after** the ingress is verified | Do these in that order. Flipping the Service type before the ingress works takes the API offline. ### C5. NATS is outside the cluster on two ports Backend uses `nats://66.116.226.161:4223`; `core-config.yaml:10` uses `:4222`. Worth a line in `docs/ARCHITECTURE.md` on which port is which and why, before the engine joins and needs to pick one. ### C6. Decide replica safety before scaling the engine The telemetry recorder is already replica-safe (NATS queue group + `uq_aiagentruns_agent_task`). The **agents** are not obviously so: `exception_agent` has an in-process TTL dedup set (`exception_agent.py:339`), which is per-pod. Two engine replicas would each decide on the same stall. Keep `replicas: 1` until dedup moves to Redis. **Files** (only if scaling past 1 replica) | File | Change | |---|---| | `AI_engine/agents/exception_agent.py:339` | modify — TTL set → Redis `SET NX EX` | | `AI_engine/tests/test_stall_dedup.py` | modify — covers the current in-process behaviour | | `kubernetes/manifests/doormile/ai-engine.yaml` | modify — raise `replicas` | C5 (the NATS port note) is documentation only: `kubernetes/docs/ARCHITECTURE.md`. --- ## Track D — Executor backlog 9 of 22 seeded tools are marked `REVIEW ONLY … No executor` in `internal/ai/registry/seed.go`. The registry is honest about it and the console renders them disabled. This is the feature list, in value order. ### D1. `assign_riders` — needs an admin auto-assign route (not just auth) **Correcting an earlier assumption:** this is *not* a free auth fix. `actions.js:96-107` already explains why — `POST /admin/bookings/:id/assign-miler` (`routes.go:413` → `adminController.go:2965`) requires a **chosen rider per booking** (`{mileruserid}`), and a finding does not pick one. The hub route that *does* pick (`POST /hub/bookings/:id/auto-assign`, `routes.go:520`) is behind `HubStaffAuth` and 403s for every console login. **The cleanest route is an admin batch-assign**, better than the per-booking auto-assign first considered. `HubBatchAssign` (`controllers/hubController.go:1963`) already does exactly what a finding needs: it takes `bookingids[]`, picks riders via Redis GEO + scoring, and commits server-side in one call. Its only hub-specific parts — `c.Locals("hubid")` and `hubPincodePrefix(hubID)` (`hubController.go:1964-1966`) — are used **solely as a fallback when `bookingids` is empty** (`hubController.go:1981-1985`). The console always passes explicit ids, so that branch never runs. So: extract the body into a shared helper taking `(bookingIDs, capPerRider, actorID, scopeFn)` and have both the hub route and a new admin route call it. `scopeBookingsToOwnTenant` (`hubController.go:1986`) already works for admin logins. The console side is then nearly free — `batchAssignBookings` already exists at `src/api/doormile/endpoints.js:554` and already sends `{bookingids, max_per_rider}`. It just points at the hub URL that 403s. One URL change. Option (b), having the skill pick a rider via `nearby_milers` and call the existing `assign-miler`, is worse: more console work and it puts solver logic in the browser. **Files** | File | Change | |---|---| | `doormile_backend/controllers/hubController.go:1963` | modify — extract the shared assign helper out of `HubBatchAssign` | | `doormile_backend/controllers/adminController.go` | **add** `AdminBatchAssign` calling that helper | | `doormile_backend/routes/routes.go:413` | modify — register `adminAuth.Post("/bookings/batch-assign", …)` | | `krow_talent_app/src/api/doormile/endpoints.js:554` | modify — `/hub/bookings/batch-assign` → `/admin/bookings/batch-assign` | | `krow_talent_app/src/lib/assistant/agent/actions.js:122` | modify — add the `assignMiler` executor | | `krow_talent_app/src/lib/assistant/agent/actions.js:96-107` | modify — delete the "deliberately NOT an executor" note | | `krow_talent_app/tests/lib/agentActions.test.js` | modify — asserts the current executor set | | `doormile_backend/internal/ai/registry/seed.go` | modify — `assign_riders` description stops saying REVIEW ONLY | Check before starting: `endpoints.js:547-553` warns that Doormile-native batch assign commits with no preview/reconcile step and leaves multi-stop riders unsequenced. An agent-proposed assignment firing straight to commit is a behaviour decision, not just a wiring one — confirm that is wanted. This takes the console from 1 working verb to 2 and makes the highest-severity SLA finding actionable. ### D2. `alert_low_battery_rider` — nearly free `seed.go` already targets `POST /admin/milers/:id/notify`, which is the endpoint `notify_riders` already uses successfully. This is a message-text change and an `EXECUTORS` entry, not a new capability. **Files** | File | Change | |---|---| | `krow_talent_app/src/lib/assistant/agent/actions.js:122` | modify — add the `alertLowBatteryRider` executor | | `krow_talent_app/src/lib/assistant/skills/definitions/RiderBatterySafetySkill.js` | check — confirm the proposal carries `milerId` | | `krow_talent_app/tests/lib/agentActions.test.js` | modify — same assertion as D1 | | `doormile_backend/internal/ai/registry/seed.go` | modify — drop REVIEW ONLY from the description | No backend change. `notifyMiler` already exists at `src/api/doormile/endpoints.js:389` → `POST /admin/milers/:id/notify`, which is the endpoint `seed.go` already names as the target. ### D3. The rest need endpoints that do not exist `enforce_otp_verification`, `dispatch_hub_idle_parcels`, `trigger_auto_dispatch`, `enforce_cash_handoff`, `rebalance_riders` — all marked `Target: "none yet"`. Each is a product decision first. Not in this phase. --- ## Two more loose ends - **`src/lib/agentNetwork.js` is a hand-maintained snapshot** dated 16–20 Sep, and the file says so honestly. Once C1 gives the engine an HTTP surface, add `GET /agents/status` and swap the source. The file is deliberately shaped like that response, so it is a change of source, not a rewrite. - **`src/lib/assistant/ragRouter.js` points at a `services/ai` sidecar that does not exist in any repo.** `VITE_AI_URL` appears nowhere, so `isRagEnabled()` is permanently false and the module is dead code. **Do not conflate this with Track B** — ragRouter is semantic *intent routing* for the console assistant, not decision memory. Lower value. Leave it dormant (it is correctly fail-open) or decide to build the sidecar as its own piece of work. --- ## Consolidated file manifest **14 new files, 46 modified, across 4 repos.** Per-item detail is in the tracks above. ### `doormile_backend` — 4 new, 17 modified | File | New? | Items | |---|---|---| | `internal/ai/outcomes/sweeper.go` | **new** | B3, B5 | | `internal/ai/outcomes/rules.go` | **new** | B3 | | `internal/ai/outcomes/sweeper_test.go` | **new** | B3 | | `migrations/migrate.go` | | B0.1 (:92), B5 (:98) | | `routes/routes.go` | | B0.2 (:594), D1 (:413) | | `controllers/agentDecisionController.go` | | B0.2 (:74), B6 (:18, :104) | | `controllers/hubController.go` | | D1 (:1963 — extract helper) | | `controllers/adminController.go` | | D1 (**add** `AdminBatchAssign`) | | `models/agentdecision.go` | | B6 | | `internal/ai/registry/seed.go` | | B4 (:114), D1, D2 | | `main.go` | | B3 (:244) | | `routes/routes_ai_registry_pg_test.go` | | B0.2, B6 | ### `AI_engine` — 5 new, 20 modified | File | New? | Items | |---|---|---| | `core/embeddings.py` | **new** | B2 | | `core/memory.py` | **new** | B4 | | `core/health.py` | **new** | C1 | | `tests/test_embeddings.py` | **new** | B2 | | `tests/test_memory.py` · `tests/test_health.py` | **new** | B2, C1 | | `core/decisions.py` | | B2 (:33, :53), B6 | | `core/llm.py` | | A4 (:11, :22), B4 (:157, :227) | | `core/registry.py` | | — read-only (`loaded` consumed by C1) | | `agents/exception_agent.py` | | B4 (:456), B6, C6 (:339) | | `agents/dispatch_agent.py` | | B4 (:335), B6 | | `main.py` | | C1 (:130, `print_help`) | | `config/system_config.py` | | B2 | | `requirements.txt` · `Dockerfile` · `docker-compose.yml` | | A3, A4, C1 | | `evals/stall_eval.py` · `assignment_eval.py` · both `.jsonl` | | B4 | | `tests/test_registry_phase5.py` · `test_llm.py` · `test_stall_dedup.py` | | B2, A4, C6 | ### `kubernetes` — 5 new, 7 modified | File | New? | Items | |---|---|---| | `manifests/doormile/doormile-config.yaml` | **new** | A1 | | `manifests/doormile/ai-engine.yaml` | **new** | C1, C6 | | `manifests/doormile/kustomization.yaml` | **new** | C2 | | `manifests/doormile/doormile-pdb.yaml` | **new** | C3 | | `deploy-doormile.sh` | **new** | C2 | | `manifests/doormile/miletruth.yaml` | | A1, A2, C3 (:23), C4 (:91) | | `manifests/core/ingress-unified.yaml` | | C4 | | `conf/nginx-doormile.conf` | | C4 (last) | | `scripts/sync_manifests.py` | | C2 | | `docs/DEPLOY.md` · `DEPLOY_CHECKLIST.md` · `ARCHITECTURE.md` | | A2, C2, C5 | ### `krow_talent_app` — 0 new, 4 modified The console barely changes. Everything it needs already exists. | File | Items | |---|---| | `src/api/doormile/endpoints.js` | D1 (:554 — one URL) | | `src/lib/assistant/agent/actions.js` | D1 (:96-107, :122), D2 (:122) | | `tests/lib/agentActions.test.js` | D1, D2 | | `src/lib/assistant/skills/definitions/RiderBatterySafetySkill.js` | D2 — check only | Untouched on purpose: `src/lib/agentNetwork.js` (until C1 ships an endpoint) and `src/lib/assistant/ragRouter.js` (dormant, out of scope — see loose ends). ### Files deliberately not touched - `src/lib` / `src/utils` duplicate trees — both have live importers, do not consolidate. - `AI_engine/core/tool_registry.py` — the 2-vs-22 tool gap is a separate decision, not this phase. - `AI_engine/customer_portal/`, `dashboard/` — not on any path this phase touches. ## Decisions needed before coding 1. **Embedding provider** — OpenAI 1536 (no column change), Voyage 1024, or local 384? (B1) 2. **Is `pgvector` installed on `logistics`?** If creating extensions needs an admin, that is a prerequisite, not a migration. (B0.1) 3. **Outcome definitions** — are the two in B3 right, and what is N for `assignment_failure`? 4. **`tenantid` on `agent_decisions`** — add it now, or accept cross-tenant precedent? (B6) 5. **`assign_riders`** — route (a) admin auto-assign, or (b) console picks the rider? (D1) 6. **Decision retention window** — 180 days, or keep-resolved-purge-unresolved? (B5) ## Sequencing ``` A1 ──> A3, A4 ──┬──> B0 ──> B2 ──> B3 ──> B4 ──> B5, B6 │ └──> C1(health) ──> C1(manifest) ──> C2 ──> C3 ──> C4 A2 (independent, before A1 adds more secrets to git) D1, D2 (independent of everything above) ``` **A1 first and alone.** Until the internal key is set, the engine is not talking to the backend, so every Track B acceptance check would fail for the wrong reason. Run the `kubectl` check in A1 before writing any code — if git does not match the cluster, that changes the shape of Track C. Cheapest real step up a level, once A1 is in: **D1 + D2.** Two executors, no new LLM work, and it moves autonomy off the floor. --- ## Standing constraints - Nothing in this plan is committed, pushed, or deployed without being asked. - No migrations run against a real database without being asked. - No files deleted or untracked without per-file approval. - `src/lib` and `src/utils` in the console **both** have live importers — neither tree is dead, do not consolidate them as part of this work.