Files
doormilxpress_astryx/docs/agent-platform-phase7-plan.md

778 lines
38 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Doormile Agent Platform — Phase 7: Decision Memory, Deployment, Executors
Status: **mostly implemented 2026-10-08 (uncommitted, not deployed, no migration run)**
Written 2026-10-08 · Scope: `krow_talent_app`, `doormile_backend`, `AI_engine`, `kubernetes`
---
## Status (2026-10-08)
Verified: backend `go build`/`go vet` clean, 20 packages pass · engine 148 tests
(3 pre-existing import errors) · console 1397 pass (13 pre-existing
`agentStudio.test.jsx` failures) · `kubectl kustomize manifests/doormile` renders
clean. **Nothing committed. No migration has run. No SQL in this work has ever
executed against a database.**
| Item | State |
|---|---|
| A1 config/secrets manifest, `envFrom`, probes, resources | **built** — placeholders unreplaced |
| A2 secrets out of git | **not done** — documented in `manifests/doormile/SECRETS.md`, including the kustomize-clobbers-hand-created-secrets trap. Five more placeholder keys were added. |
| A3 `anthropic>=0.49.0` · A4 `claude-opus-5-5` | **done** |
| B0.1 `CREATE EXTENSION vector` · B0.2 `/similar` → POST | **done** |
| B1 provider (OpenAI `text-embedding-3-small`, 1536) | **done** — my call, flagged |
| B2 `core/embeddings.py` + decision wiring | **done** |
| B3 outcome sweeper + rules + 14 tests | **done** |
| B4 `core/memory.py` + both agents | **done** |
| B5 retention/prune | **done** — ivfflat `lists` re-tune still needs a row count |
| B6 `tenantid` scoping | **done** |
| B2.5 historical backfill | **done** — `internal/ai/outcomes/backfill.go` |
| B7 persist console findings | **done** — `aiskillfindings`, 4 routes, `findingReport.js`, 15 tests |
| C1 health surface + `ai-engine.yaml` + Dockerfile/compose port | **done** — image not built or pushed |
| C2 kustomization + `deploy-doormile.sh` | **done** |
| C3 probes/resources/PDB | **done** — StatefulSet→Deployment not done |
| C4 ingress | **built** — apply order matters, see the file header |
| C5 NATS port note | **done** — `docs/ARCHITECTURE.md` |
| C6 engine replica safety | **see correction 1 — was already safe for a different reason** |
| D1 admin batch-assign + console executor | **done** |
| D2 `alert_low_battery_rider` | **done** |
| D3 five review-only tools | **1 of 5 done** — `trigger_auto_dispatch` turned out to BE batch-assign under another name. The other four need endpoints that do not exist. |
| Agents-page snapshot → live endpoint | **backend done** — engine serves `GET /agents/status`; the console still reads the snapshot, so the swap is one fetch away |
## Two things this plan got wrong
**Correction 1 — C6. The engine's stall dedup was never per-pod.**
This plan claimed `ExceptionAgent` kept its dedup in an in-process TTL set and
that moving it to Redis was the prerequisite for scaling out. Wrong:
`_claim_stall` / `_claim_stall_handled` (`agents/exception_agent.py:336`) have
been Redis `SET NX` with a TTL all along — cross-pod, self-expiring, durable
across restarts. The comment there saying "replaces the old unbounded in-memory
set" describes what it REPLACED; I read it as current state.
The real blocker is different and still real: `core/message_bus.py:284` and
`:312` bind **durable push consumers with fixed names**, and a durable push
consumer admits one active subscriber unless created with a deliver group. A
second replica does not duplicate work — it fails to bind. Scaling out means
giving those subscriptions a deliver group (or moving to pull consumers).
`replicas: 1` stands, for the corrected reason.
**Correction 2 — the registry was already honest about the simulated agents.**
This plan said the three fictional-data agents were "presented as live" and
should be marked. They already were: `seed.go` has `StatusSimulation` for
`HUB_AGENT`, `FLEET_AGENT` and `ROUTE_OPTIMIZER`, with purposes reading "an
in-memory fleet of 19 fake vehicles" and "8 hard-coded fictional hubs". The
console snapshot (`src/lib/agentNetwork.js`) was also honest about activity
("Zero tasks received", `publishes: 'Nothing on the bus.'`, `api: 'No.'`) but
silent on the data being invented — so the three entries now say so, and the two
surfaces agree.
I overstated that surface as "telling anyone something untrue". It was
incomplete, not false.
---
## Track A — Unblock (do this first; nothing else matters until it lands)
### A1. `INTERNAL_API_KEY` is missing from the cluster — **everything engine↔backend is 401**
`middlewares/internal_auth.go:13` reads `INTERNAL_API_KEY` and **fails closed when
unset**:
```go
expected := os.Getenv("INTERNAL_API_KEY")
if expected == "" || c.Get("X-Internal-Key") != expected {
```
`kubernetes/manifests/doormile/miletruth.yaml` never sets it. The engine's
`docker-compose.yml` *does* pass it. So the engine sends a key the backend
rejects, and in-cluster **every** `/api/v1/internal/*` call 401s:
- `GET /internal/ai/registry` — the registry poll (this is why "engine reads
registry in prod" was never confirmable)
- `POST /internal/agent-decisions` — the decision log, so Insights sees nothing
- `GET /internal/express/riders`, `/express/bookings`, `POST /express/assign`
- `POST /internal/bookings/:id/reassign`, `POST /internal/notify`
**Step 1 — find out whether git matches reality.** The `kubernetes` repo's log is
full of `Restore…` / `Recreate…` commits, so the manifest may already be fiction:
```bash
kubectl -n doormile get statefulset doormile -o jsonpath='{range .spec.template.spec.containers[0].env[*]}{.name}{"\n"}{end}' | sort
```
`config/config.go` reads 29 env vars. If that list is much shorter, the live
cluster has hand-applied drift and **the next `kubectl apply -f` wipes it.**
**Step 2 — close the gap in git, not by hand.** Add to `doormile-secrets`
(`stringData`) and reference from the StatefulSet. Absent from the manifest today
and read by `config.go`:
`INTERNAL_API_KEY`, `JWT_SECRET_KEY`, `APP_PORT`, `ENV`, `TRUSTED_PROXIES`,
`AI_LAYER_BASE_URL`, `ROUTE_OPTIMIZER_URL`, `GEOCODER_URL`, `GEOCODER_EMAIL`,
`SMTP_HOST`, `SMTP_PORT`, `SMTP_USER`, `SMTP_PASSWORD`, `SMTP_FROM`,
`CLIENT_ONBOARDING_OWNERS`, `PLAYGROUND_LLM_API_KEY`, `PLAYGROUND_LLM_BASE_URL`,
`PLAYGROUND_LLM_MODEL`, `REDIS_USER`.
Secrets vs ConfigMap: `INTERNAL_API_KEY`, `JWT_SECRET_KEY`, `SMTP_PASSWORD`,
`PLAYGROUND_LLM_API_KEY` are secrets. The rest belong in a `doormile-config`
ConfigMap so a URL change is not a secret edit.
**Files**
| File | Change |
|---|---|
| `kubernetes/manifests/doormile/miletruth.yaml` | modify — add the 19 env vars to the StatefulSet + `doormile-secrets` |
| `kubernetes/manifests/doormile/doormile-config.yaml` | **new** — ConfigMap for the non-secret vars |
| `AI_engine/.env` (server-local, untracked content) | modify — same `INTERNAL_API_KEY` value |
Read-only, no change: `doormile_backend/middlewares/internal_auth.go:13`,
`doormile_backend/config/config.go:60-102` (the contract being satisfied).
> `INTERNAL_API_KEY` must be **byte-identical** in the backend Secret and the
> engine's env. Generate once (`openssl rand -hex 32`), set both.
**Acceptance:** from inside the cluster,
`curl -H "X-Internal-Key: $KEY" http://doormile-service.doormile:8081/api/v1/internal/ai/registry`
returns 200 with a body and an `ETag`; a second call with `If-None-Match` returns
304. Backend logs show `ai playground: enabled`. `GET /admin/ai/status` reports
the engine as following the registry (`aiRegistryController.go:208` counts the
poll).
### A2. Plaintext secrets are committed
`miletruth.yaml:19-21` has `DB_PASSWORD`, `REDIS_PASSWORD`, `NATS_PASSWORD` as
literal in git (the same password reused for all three). Same class as the committed Firebase keys from the
CX audit. Pick one and apply it before A1 adds *more* secrets to that file:
Sealed Secrets, SOPS, or an out-of-git `kubectl create secret` with the manifest
holding only a reference.
Do **not** delete or untrack anything here without per-file approval — see the
standing rule. Rotating that shared password is a separate decision; this item is only
about stopping new secrets entering git.
**Files**
| File | Change |
|---|---|
| `kubernetes/manifests/doormile/miletruth.yaml` | modify — `stringData` block becomes a reference |
| `kubernetes/manifests/doormile/sealed-secrets.yaml` | **new** — if Sealed Secrets is the chosen route |
| `kubernetes/docs/DEPLOY.md` | modify — document how secrets are supplied now |
Same pattern already in git at `kubernetes/manifests/core/core-secrets.yaml` and
`manifests/nearle/nearle-secrets.yaml` — whatever is chosen should cover those too,
but that is outside this phase.
### A3. `anthropic>=0.21.0` floor is wrong
`AI_engine/requirements.txt` pins `anthropic>=0.21.0`, but `core/llm.py:117`
sends `thinking: {"type": "adaptive"}` and `output_config.effort` — parameters
that old SDK does not know. There is no lockfile, so a fresh `pip install` in a
rebuilt image can resolve to something that 400s on every LLM call.
Bump to a current floor and pin the image build. `openai>=1.0.0` is already
present (unused today — Track B uses it).
**Files**
| File | Change |
|---|---|
| `AI_engine/requirements.txt` | modify — raise the `anthropic` floor |
| `AI_engine/Dockerfile` | modify — optional: `pip install` from a lockfile instead |
| `AI_engine/requirements-dev.txt` | check — may pin the same packages |
### A4. `LLM_MODEL` default is a generation behind
`core/llm.py:22` defaults to `claude-opus-4-8` ($5/$25 per MTok). `claude-opus-5-5`
is both cheaper ($4/$20) and more capable. One-line change; the registry's
per-agent `model` pin already overrides it (`registry.model(agent_id)`), so this
only moves the floor.
**Files**
| File | Change |
|---|---|
| `AI_engine/core/llm.py:22` | modify — default `LLM_MODEL` |
| `AI_engine/core/llm.py:11` | modify — the docstring naming the default |
| `AI_engine/docker-compose.yml:45` | modify — `LLM_MODEL` fallback |
| `AI_engine/tests/test_llm.py` | check — may assert the old default |
Note `core/llm.py:109` special-cases Haiku 4.5 (rejects adaptive thinking and
`effort`). Leave that branch alone; it is correct.
---
## Track B — RAG decision memory
The skeleton exists and is wired to nothing. Four pieces are already built:
| Piece | Where |
|---|---|
| `context_embedding vector(1536)` on `agent_decisions` | `migrations/migrate.go:92` |
| ivfflat cosine index, `lists = 100` | `migrations/migrate.go:98` |
| `POST /internal/agent-decisions` accepts `context_embedding` | `controllers/agentDecisionController.go:23` |
| cosine kNN query + `PATCH /:id/outcome` | `agentDecisionController.go:74,120` |
Nothing produces an embedding, nothing calls `/similar`, nothing records an
outcome. There are exactly **two** decision types to cover —
`assignment_failure` (`dispatch_agent.py:335`) and `stall_response`
(`exception_agent.py:456`).
### B0. Three defects to fix before writing any new code
**B0.1 — `CREATE EXTENSION vector` appears nowhere in the repo.**
`migrate.go:92` runs `ALTER TABLE agent_decisions ADD COLUMN … vector(1536)` and
logs failure *non-fatally*. If the extension is not installed on `logistics`,
both the column and the index silently fail and every retrieval 500s.
```sql
SELECT extname, extversion FROM pg_extension WHERE extname = 'vector';
```
If absent, add **before** line 92 in `Migrate()`:
```go
if res := db.Exec(`CREATE EXTENSION IF NOT EXISTS vector`); res.Error != nil {
utils.Error("❌ pgvector extension unavailable — decision memory disabled", "error", res.Error)
}
```
Needs the `pgvector` extension available on the server and a role with rights to
create it. On a managed Postgres this may be an admin action, not a migration —
check before assuming.
**B0.2 — `/similar` is a `GET` that requires a JSON body.**
`routes.go:594` registers it as `GET`, and `agentDecisionController.go:84` calls
`BodyParser` demanding an `embedding` array. nginx and most HTTP clients drop GET
bodies — and doormile is served *through* host nginx (see C4). Change to:
```go
internal.Post("/agent-decisions/similar", controllers.FindSimilarDecisions)
```
No caller exists yet, so this breaks nothing.
**B0.3 — `/similar` filters `WHERE outcome IS NOT NULL`, and nothing writes
outcomes.** Even with embeddings flowing it returns zero rows forever. **B2 is
not optional** — it is the half that makes retrieval worth anything.
**Files**
| File | Change |
|---|---|
| `doormile_backend/migrations/migrate.go:92` | modify — add `CREATE EXTENSION` above the `ALTER TABLE` |
| `doormile_backend/routes/routes.go:594` | modify — `internal.Get` → `internal.Post` for `/agent-decisions/similar` |
| `doormile_backend/controllers/agentDecisionController.go:74` | modify — comment the method change; body parsing already correct |
| `doormile_backend/routes/routes_ai_registry_pg_test.go` | modify — add coverage for the POST shape |
### B1. Embedding provider — decision required
Anthropic has no embeddings endpoint, so this needs a second provider.
| Option | Dims | Column change | Notes |
|---|---|---|---|
| **OpenAI `text-embedding-3-small`** | 1536 | **none** | ~$0.02/MTok. `openai>=1.0.0` already a dependency. Column was sized for it. |
| Voyage `voyage-3` | 1024 | yes | Anthropic-recommended; new key, new vendor. |
| Local `bge-small` / `MiniLM` | 384 | yes | Free, no egress, no key. +~400MB RAM, model in image, slower cold start. |
**Recommendation: OpenAI `text-embedding-3-small`.** The schema already matches,
the dependency is already there, and the second-provider line is already crossed
(the playground runs on Groq). Revisit if data egress is a constraint — then take
the local model and migrate the column to `vector(384)`.
### B2. Write path — `AI_engine/core/embeddings.py`
One function, modelled on `core/decisions.py`'s fire-and-forget discipline:
```python
async def embed(text: str) -> Optional[List[float]]:
"""None on any failure. An embedding must never delay an agent's reaction."""
```
Requirements:
- **Embed the same dict `build_payload` already stores.** Serialise the `facts`
dict deterministically (sorted keys) so the embedded text and the stored
`context` cannot drift apart. Add a `_context_text(facts)` helper and test it
pure, the way `build_payload` is tested.
- Fail open — return `None`, log once, never raise into the agent path.
- Small LRU cache: repeated stalls on one booking produce near-identical facts.
- Gate on `EMBEDDINGS_ENABLED` **and** a registry skill flag, so it can be
switched off from Agent Studio without a redeploy
(`registry.skill_enabled(...)` already exists).
Then extend `build_payload`/`record_decision` (`core/decisions.py:33,53`) to carry
`context_embedding`. Both call sites (`dispatch_agent.py:335`,
`exception_agent.py:456`) keep their signatures.
**Acceptance:** after one stall,
`SELECT count(*) FROM agent_decisions WHERE context_embedding IS NOT NULL` > 0.
**Files**
| File | Change |
|---|---|
| `AI_engine/core/embeddings.py` | **new** — `embed()`, `_context_text()`, LRU cache, fail-open |
| `AI_engine/core/decisions.py:33` | modify — `build_payload` carries `context_embedding` |
| `AI_engine/core/decisions.py:53` | modify — `record_decision` awaits the embed before posting |
| `AI_engine/config/system_config.py` | modify — `EMBEDDINGS_ENABLED`, provider key, model name |
| `AI_engine/docker-compose.yml` | modify — pass the embedding env through |
| `AI_engine/tests/test_embeddings.py` | **new** — `_context_text` determinism, fail-open returns `None` |
| `AI_engine/tests/test_registry_phase5.py:194` | modify — asserts the `record_decision` payload shape |
Not touched: `agents/dispatch_agent.py:335` and `agents/exception_agent.py:456`
keep their call signatures — the embedding is added inside `decisions.py`, so
neither agent changes.
### B3. Outcome loop — the part that makes retrieval useful
`PATCH /internal/agent-decisions/:id/outcome` exists with no caller. Define, per
decision type, what "it worked" means. Starting proposal:
| Type | Outcome = `success` when | `failure` when |
|---|---|---|
| `stall_response` | booking reaches `Delivered` within its SLA window after the decision | SLA breached, or cancelled |
| `assignment_failure` | booking gets an assignment within N minutes of the decision | still unassigned after N, or cancelled |
Implement as a backend sweeper following the **established pattern** in
`internal/assignment/sweeper.go:88` — ticker, `recover()` per tick, and the Redis
lock that keeps one replica sweeping (`sweeper.go:104`). Wire it in `main.go`
beside `go assignment.StartPendingSweeper()` (`main.go:244`).
Two gotchas from the existing code:
- Use `Receivedat`-style true instants, not `utils.DBNow` — `models/ai_runs.go:26`
documents that `DBNow` returns IST digits labelled UTC and is 5h30m off for
`timestamptz`. The same trap applies to `outcome_recorded_at`.
- Leave `outcome` NULL while undecided. `/similar` already treats NULL as
"no evidence yet", which is correct.
**Acceptance:** rows acquire non-NULL `outcome` within one sweep interval of
their window closing, and `/insights` decision-outcome counts stop being all
`pending`.
**Files**
| File | Change |
|---|---|
| `doormile_backend/internal/ai/outcomes/sweeper.go` | **new** — ticker + Redis lock, modelled on `internal/assignment/sweeper.go:88` |
| `doormile_backend/internal/ai/outcomes/rules.go` | **new** — the per-decision-type success/failure predicates |
| `doormile_backend/internal/ai/outcomes/sweeper_test.go` | **new** — interval, window, and both predicates |
| `doormile_backend/main.go:244` | modify — `go outcomes.StartOutcomeSweeper()` beside the pending sweeper |
| `doormile_backend/models/agentdecision.go` | modify — only if B6 adds `Tenantid` |
Reference, not modified: `internal/assignment/sweeper.go:88-110` (the ticker +
`recover()` + Redis single-replica lock pattern to copy),
`models/ai_runs.go:26` (the `utils.DBNow` timezone trap to avoid).
### B4. Read path — precedent in the prompt
Before the LLM call in `exception_agent` / `dispatch_agent`, fetch the top-5
similar **resolved** decisions and include them as precedent — "the last 5
comparable situations and whether the action worked."
- Feature-flag it on a registry skill so it is switchable from Agent Studio.
- Hard timeout (~300ms) with fail-open to today's prompt. The current behaviour
is the floor; this can only raise it. (`core/registry.py` already follows this
rule; `ragRouter.js:9-14` documents the same discipline on the console side.)
- Keep precedent **out** of the structured-output schema. It informs the prompt;
it must not become a field the model can invent.
**Acceptance:** an eval run shows the decision quality moving. `AI_engine/evals/`
already has the harness and cases (`stall_cases.jsonl`,
`assignment_cases.jsonl`) — extend those rather than judging by eye.
**Files**
| File | Change |
|---|---|
| `AI_engine/core/memory.py` | **new** — `recall(decision_type, embedding, k)` → POST `/internal/agent-decisions/similar`, timeout + fail-open |
| `AI_engine/core/llm.py:157` | modify — `decide_stall_response` accepts optional precedent |
| `AI_engine/core/llm.py:227` | modify — `decide_assignment_failure` accepts optional precedent |
| `AI_engine/agents/exception_agent.py:456` | modify — recall before the decide call |
| `AI_engine/agents/dispatch_agent.py:335` | modify — recall before the decide call |
| `AI_engine/tests/test_memory.py` | **new** — timeout fails open, empty recall changes nothing |
| `AI_engine/evals/stall_eval.py` | modify — run with and without precedent |
| `AI_engine/evals/assignment_eval.py` | modify — same |
| `AI_engine/evals/stall_cases.jsonl` | modify — cases where precedent should change the answer |
| `AI_engine/evals/assignment_cases.jsonl` | modify — same |
| `doormile_backend/internal/ai/registry/seed.go:114` | modify — seed a `recall_similar_decisions` read tool + the skill flag that gates B4 |
The registry seed row matters: without it the feature cannot be switched off from
Agent Studio, which is the whole point of gating it on `registry.skill_enabled`.
### B5. Hygiene
- **`agent_decisions` has no retention.** `aiagentruns` purges at 30 days
(`telemetry/recorder.go:28,128`); decisions grow forever, and this is the table
retrieval scans. Decide a window — longer than 30 days, since old precedent is
the point. Suggest 180 days, or keep resolved rows and purge unresolved ones.
- **Re-tune `lists = 100`.** That is right for roughly 100k–1M rows. Below ~10k
it over-partitions and recall drops. Check `count(*)` once embeddings flow;
consider HNSW instead if the pgvector version supports it.
**Files**
| File | Change |
|---|---|
| `doormile_backend/internal/ai/outcomes/sweeper.go` | modify — fold the decision purge into the same tick |
| `doormile_backend/migrations/migrate.go:98` | modify — index tuning, once row count is known |
### B6. Tenant isolation — decide before B4 ships, not after
`agent_decisions` has **no tenant column**. If retrieved precedent crosses
tenants, one client's operational history shapes decisions made for another.
Given that console logins are already unscoped on `tenantid` NULL
(`doormile-console-logins-unscoped`), this needs deciding up front.
Recommendation: add `tenantid` to `agent_decisions`, have the engine populate it,
and filter in the `/similar` query. Cheap now, expensive after the table fills.
**Files**
| File | Change |
|---|---|
| `doormile_backend/models/agentdecision.go` | modify — add `Tenantid *uint64` with an index |
| `doormile_backend/controllers/agentDecisionController.go:18` | modify — accept `tenant_id` on create |
| `doormile_backend/controllers/agentDecisionController.go:104` | modify — add `AND tenantid = ?` to the kNN query |
| `AI_engine/core/decisions.py:33` | modify — `build_payload` carries the tenant |
| `AI_engine/agents/exception_agent.py` · `dispatch_agent.py` | modify — source the tenant from the booking facts |
| `doormile_backend/routes/routes_ai_registry_pg_test.go` | modify — a cross-tenant recall must return nothing |
Nullable, because the engine will not always know the tenant. Decide whether a
NULL tenant row is recallable by everyone or by no one — given
`doormile-console-logins-unscoped`, **by no one** is the safer default.
---
## Track C — Kubernetes
### C1. `AI_engine` is not in Kubernetes at all
No manifest, no kustomization, no deploy script, and `docker-compose.yml` uses
`build: .` with no registry push. It runs on a VM by compose.
**Prerequisite: the engine has no HTTP server in production mode.**
`main.py --production` starts no listener, so there is no liveness/readiness
target and no `/metrics`. Without it, Kubernetes can only restart on process
exit — a NATS-disconnected engine looks healthy forever.
`fastapi` and `uvicorn` are already in `requirements.txt`. Add a small surface in
`production_mode()`:
- `GET /healthz` — process alive (event loop responsive)
- `GET /readyz` — NATS connected **and** `registry.loaded` is true
- `GET /metrics` — optional; decisions recorded, tool calls, LLM failures
Readiness must include `registry.loaded`, otherwise a pod that cannot reach the
backend serves traffic on env defaults while reporting healthy — exactly the A1
failure mode, invisible again.
Then: push the image to a registry (compose builds locally), and write
`manifests/doormile/ai-engine.yaml` as a `Deployment` (it is stateless;
`replicas: 1` to start — the agents are not yet idempotent across replicas, see
C6).
**Files**
| File | Change |
|---|---|
| `AI_engine/core/health.py` | **new** — the aiohttp/FastAPI surface (`/healthz`, `/readyz`, `/metrics`) |
| `AI_engine/main.py:130` | modify — start the health server inside `production_mode()` beside `registry.run(...)` |
| `AI_engine/main.py` (`print_help`) | modify — document the health port |
| `AI_engine/Dockerfile` | modify — `EXPOSE` the health port |
| `AI_engine/docker-compose.yml` | modify — publish the port so compose and k8s behave alike |
| `AI_engine/tests/test_health.py` | **new** — `/readyz` is red while `registry.loaded` is false |
| `kubernetes/manifests/doormile/ai-engine.yaml` | **new** — Deployment + Service + probes + resources |
`fastapi` and `uvicorn` are already in `requirements.txt` — no new dependency.
Readiness must check `registry.loaded` (`core/registry.py:45`), not just the
process, or A1's failure mode becomes invisible again.
### C2. `doormile/` has no `kustomization.yaml`
`alaska/`, `core/` and `nearle/` all have one. `doormile/` does not, and there is
no `deploy-doormile.sh` alongside `deploy-core-stack.sh` / `deploy-nearle-stack.sh`.
That is *why* it drifts. Add both.
**Files**
| File | Change |
|---|---|
| `kubernetes/manifests/doormile/kustomization.yaml` | **new** — list miletruth, config, ai-engine, pdb |
| `kubernetes/deploy-doormile.sh` | **new** — copy the shape of `deploy-nearle-stack.sh` |
| `kubernetes/scripts/sync_manifests.py` | check — may need the new namespace registering |
| `kubernetes/docs/DEPLOY_CHECKLIST.md` | modify — add the doormile stack |
Pattern to copy: `manifests/nearle/kustomization.yaml` + `deploy-nearle-stack.sh`.
### C3. The `doormile` StatefulSet has no probes, resources or PDB
`replicas: 3` with **no resource requests** means the scheduler can stack all
three on one node, and **no readinessProbe** means a pod receives traffic before
Postgres/Redis/NATS are connected. `core/` has `worker-pdb.yaml`; doormile has
nothing.
Add `resources.requests`/`limits`, a readinessProbe and livenessProbe against the
backend's health route, and a PodDisruptionBudget (`minAvailable: 2`).
Also: a `StatefulSet` for a stateless Go API is the wrong kind — it gives serial
rollouts and no benefit. Switching to `Deployment` is low-risk and makes deploys
faster. Not urgent; flagging because it is why rollouts feel slow.
**Files**
| File | Change |
|---|---|
| `kubernetes/manifests/doormile/miletruth.yaml:23` | modify — `resources`, `readinessProbe`, `livenessProbe` |
| `kubernetes/manifests/doormile/doormile-pdb.yaml` | **new** — `minAvailable: 2`, copy `manifests/core/worker-pdb.yaml` |
**No backend change needed** — the probe targets already exist and are correct:
`GET /api/v1/health` (`routes/routes.go:46`, unauthenticated, always 200) for
liveness, and `GET /api/v1/ready` (`routes/routes.go:50`) for readiness, which
already returns **503** when Postgres or Redis is unreachable
(`routes.go:68-70`). Point the probes at those; do not write new ones.
Note `/ready` reports Redis GEO status without gating on it — deliberate, per the
comment at `routes.go:73`. A readinessProbe on `/ready` therefore will not pull a
pod out of service for a broken rider search, which is the intended behaviour.
### C4. No ingress for `doormile`
`nearle` and `alaska` are on `manifests/core/ingress-unified.yaml`. `doormile` is
NodePort 30830 plus host nginx (`conf/nginx-doormile.conf`) — half-migrated. This
also makes B0.2 (GET-with-body) a certainty rather than a risk.
**Files**
| File | Change |
|---|---|
| `kubernetes/manifests/core/ingress-unified.yaml` | modify — add a `doormile` rule (needs a ReferenceGrant if it stays cross-namespace, cf. `manifests/nearle/nearle-reference-grant.yaml`) |
| `kubernetes/manifests/doormile/miletruth.yaml:91` | modify — `NodePort` → `ClusterIP` once the ingress serves it |
| `kubernetes/conf/nginx-doormile.conf` | modify — retire or repoint, **only after** the ingress is verified |
Do these in that order. Flipping the Service type before the ingress works takes
the API offline.
### C5. NATS is outside the cluster on two ports
Backend uses `nats://66.116.226.161:4223`; `core-config.yaml:10` uses `:4222`.
Worth a line in `docs/ARCHITECTURE.md` on which port is which and why, before the
engine joins and needs to pick one.
### C6. Decide replica safety before scaling the engine
The telemetry recorder is already replica-safe (NATS queue group +
`uq_aiagentruns_agent_task`). The **agents** are not obviously so: `exception_agent`
has an in-process TTL dedup set (`exception_agent.py:339`), which is per-pod. Two
engine replicas would each decide on the same stall. Keep `replicas: 1` until
dedup moves to Redis.
**Files** (only if scaling past 1 replica)
| File | Change |
|---|---|
| `AI_engine/agents/exception_agent.py:339` | modify — TTL set → Redis `SET NX EX` |
| `AI_engine/tests/test_stall_dedup.py` | modify — covers the current in-process behaviour |
| `kubernetes/manifests/doormile/ai-engine.yaml` | modify — raise `replicas` |
C5 (the NATS port note) is documentation only: `kubernetes/docs/ARCHITECTURE.md`.
---
## Track D — Executor backlog
9 of 22 seeded tools are marked `REVIEW ONLY … No executor` in
`internal/ai/registry/seed.go`. The registry is honest about it and the console
renders them disabled. This is the feature list, in value order.
### D1. `assign_riders` — needs an admin auto-assign route (not just auth)
**Correcting an earlier assumption:** this is *not* a free auth fix.
`actions.js:96-107` already explains why — `POST /admin/bookings/:id/assign-miler`
(`routes.go:413` → `adminController.go:2965`) requires a **chosen rider per
booking** (`{mileruserid}`), and a finding does not pick one. The hub route that
*does* pick (`POST /hub/bookings/:id/auto-assign`, `routes.go:520`) is behind
`HubStaffAuth` and 403s for every console login.
**The cleanest route is an admin batch-assign**, better than the per-booking
auto-assign first considered. `HubBatchAssign` (`controllers/hubController.go:1963`)
already does exactly what a finding needs: it takes `bookingids[]`, picks riders
via Redis GEO + scoring, and commits server-side in one call. Its only hub-specific
parts — `c.Locals("hubid")` and `hubPincodePrefix(hubID)` (`hubController.go:1964-1966`)
— are used **solely as a fallback when `bookingids` is empty** (`hubController.go:1981-1985`).
The console always passes explicit ids, so that branch never runs.
So: extract the body into a shared helper taking `(bookingIDs, capPerRider, actorID, scopeFn)`
and have both the hub route and a new admin route call it. `scopeBookingsToOwnTenant`
(`hubController.go:1986`) already works for admin logins.
The console side is then nearly free — `batchAssignBookings` already exists at
`src/api/doormile/endpoints.js:554` and already sends `{bookingids, max_per_rider}`.
It just points at the hub URL that 403s. One URL change.
Option (b), having the skill pick a rider via `nearby_milers` and call the
existing `assign-miler`, is worse: more console work and it puts solver logic in
the browser.
**Files**
| File | Change |
|---|---|
| `doormile_backend/controllers/hubController.go:1963` | modify — extract the shared assign helper out of `HubBatchAssign` |
| `doormile_backend/controllers/adminController.go` | **add** `AdminBatchAssign` calling that helper |
| `doormile_backend/routes/routes.go:413` | modify — register `adminAuth.Post("/bookings/batch-assign", …)` |
| `krow_talent_app/src/api/doormile/endpoints.js:554` | modify — `/hub/bookings/batch-assign` → `/admin/bookings/batch-assign` |
| `krow_talent_app/src/lib/assistant/agent/actions.js:122` | modify — add the `assignMiler` executor |
| `krow_talent_app/src/lib/assistant/agent/actions.js:96-107` | modify — delete the "deliberately NOT an executor" note |
| `krow_talent_app/tests/lib/agentActions.test.js` | modify — asserts the current executor set |
| `doormile_backend/internal/ai/registry/seed.go` | modify — `assign_riders` description stops saying REVIEW ONLY |
Check before starting: `endpoints.js:547-553` warns that Doormile-native batch
assign commits with no preview/reconcile step and leaves multi-stop riders
unsequenced. An agent-proposed assignment firing straight to commit is a
behaviour decision, not just a wiring one — confirm that is wanted.
This takes the console from 1 working verb to 2 and makes the highest-severity
SLA finding actionable.
### D2. `alert_low_battery_rider` — nearly free
`seed.go` already targets `POST /admin/milers/:id/notify`, which is the endpoint
`notify_riders` already uses successfully. This is a message-text change and an
`EXECUTORS` entry, not a new capability.
**Files**
| File | Change |
|---|---|
| `krow_talent_app/src/lib/assistant/agent/actions.js:122` | modify — add the `alertLowBatteryRider` executor |
| `krow_talent_app/src/lib/assistant/skills/definitions/RiderBatterySafetySkill.js` | check — confirm the proposal carries `milerId` |
| `krow_talent_app/tests/lib/agentActions.test.js` | modify — same assertion as D1 |
| `doormile_backend/internal/ai/registry/seed.go` | modify — drop REVIEW ONLY from the description |
No backend change. `notifyMiler` already exists at
`src/api/doormile/endpoints.js:389` → `POST /admin/milers/:id/notify`, which is
the endpoint `seed.go` already names as the target.
### D3. The rest need endpoints that do not exist
`enforce_otp_verification`, `dispatch_hub_idle_parcels`, `trigger_auto_dispatch`,
`enforce_cash_handoff`, `rebalance_riders` — all marked `Target: "none yet"`.
Each is a product decision first. Not in this phase.
---
## Two more loose ends
- **`src/lib/agentNetwork.js` is a hand-maintained snapshot** dated 16–20 Sep, and
the file says so honestly. Once C1 gives the engine an HTTP surface, add
`GET /agents/status` and swap the source. The file is deliberately shaped like
that response, so it is a change of source, not a rewrite.
- **`src/lib/assistant/ragRouter.js` points at a `services/ai` sidecar that does
not exist in any repo.** `VITE_AI_URL` appears nowhere, so `isRagEnabled()` is
permanently false and the module is dead code. **Do not conflate this with
Track B** — ragRouter is semantic *intent routing* for the console assistant,
not decision memory. Lower value. Leave it dormant (it is correctly fail-open)
or decide to build the sidecar as its own piece of work.
---
## Consolidated file manifest
**14 new files, 46 modified, across 4 repos.** Per-item detail is in the tracks above.
### `doormile_backend` — 4 new, 17 modified
| File | New? | Items |
|---|---|---|
| `internal/ai/outcomes/sweeper.go` | **new** | B3, B5 |
| `internal/ai/outcomes/rules.go` | **new** | B3 |
| `internal/ai/outcomes/sweeper_test.go` | **new** | B3 |
| `migrations/migrate.go` | | B0.1 (:92), B5 (:98) |
| `routes/routes.go` | | B0.2 (:594), D1 (:413) |
| `controllers/agentDecisionController.go` | | B0.2 (:74), B6 (:18, :104) |
| `controllers/hubController.go` | | D1 (:1963 — extract helper) |
| `controllers/adminController.go` | | D1 (**add** `AdminBatchAssign`) |
| `models/agentdecision.go` | | B6 |
| `internal/ai/registry/seed.go` | | B4 (:114), D1, D2 |
| `main.go` | | B3 (:244) |
| `routes/routes_ai_registry_pg_test.go` | | B0.2, B6 |
### `AI_engine` — 5 new, 20 modified
| File | New? | Items |
|---|---|---|
| `core/embeddings.py` | **new** | B2 |
| `core/memory.py` | **new** | B4 |
| `core/health.py` | **new** | C1 |
| `tests/test_embeddings.py` | **new** | B2 |
| `tests/test_memory.py` · `tests/test_health.py` | **new** | B2, C1 |
| `core/decisions.py` | | B2 (:33, :53), B6 |
| `core/llm.py` | | A4 (:11, :22), B4 (:157, :227) |
| `core/registry.py` | | — read-only (`loaded` consumed by C1) |
| `agents/exception_agent.py` | | B4 (:456), B6, C6 (:339) |
| `agents/dispatch_agent.py` | | B4 (:335), B6 |
| `main.py` | | C1 (:130, `print_help`) |
| `config/system_config.py` | | B2 |
| `requirements.txt` · `Dockerfile` · `docker-compose.yml` | | A3, A4, C1 |
| `evals/stall_eval.py` · `assignment_eval.py` · both `.jsonl` | | B4 |
| `tests/test_registry_phase5.py` · `test_llm.py` · `test_stall_dedup.py` | | B2, A4, C6 |
### `kubernetes` — 5 new, 7 modified
| File | New? | Items |
|---|---|---|
| `manifests/doormile/doormile-config.yaml` | **new** | A1 |
| `manifests/doormile/ai-engine.yaml` | **new** | C1, C6 |
| `manifests/doormile/kustomization.yaml` | **new** | C2 |
| `manifests/doormile/doormile-pdb.yaml` | **new** | C3 |
| `deploy-doormile.sh` | **new** | C2 |
| `manifests/doormile/miletruth.yaml` | | A1, A2, C3 (:23), C4 (:91) |
| `manifests/core/ingress-unified.yaml` | | C4 |
| `conf/nginx-doormile.conf` | | C4 (last) |
| `scripts/sync_manifests.py` | | C2 |
| `docs/DEPLOY.md` · `DEPLOY_CHECKLIST.md` · `ARCHITECTURE.md` | | A2, C2, C5 |
### `krow_talent_app` — 0 new, 4 modified
The console barely changes. Everything it needs already exists.
| File | Items |
|---|---|
| `src/api/doormile/endpoints.js` | D1 (:554 — one URL) |
| `src/lib/assistant/agent/actions.js` | D1 (:96-107, :122), D2 (:122) |
| `tests/lib/agentActions.test.js` | D1, D2 |
| `src/lib/assistant/skills/definitions/RiderBatterySafetySkill.js` | D2 — check only |
Untouched on purpose: `src/lib/agentNetwork.js` (until C1 ships an endpoint) and
`src/lib/assistant/ragRouter.js` (dormant, out of scope — see loose ends).
### Files deliberately not touched
- `src/lib` / `src/utils` duplicate trees — both have live importers, do not consolidate.
- `AI_engine/core/tool_registry.py` — the 2-vs-22 tool gap is a separate decision, not this phase.
- `AI_engine/customer_portal/`, `dashboard/` — not on any path this phase touches.
## Decisions needed before coding
1. **Embedding provider** — OpenAI 1536 (no column change), Voyage 1024, or local
384? (B1)
2. **Is `pgvector` installed on `logistics`?** If creating extensions needs an
admin, that is a prerequisite, not a migration. (B0.1)
3. **Outcome definitions** — are the two in B3 right, and what is N for
`assignment_failure`?
4. **`tenantid` on `agent_decisions`** — add it now, or accept cross-tenant
precedent? (B6)
5. **`assign_riders`** — route (a) admin auto-assign, or (b) console picks the
rider? (D1)
6. **Decision retention window** — 180 days, or keep-resolved-purge-unresolved? (B5)
## Sequencing
```
A1 ──> A3, A4 ──┬──> B0 ──> B2 ──> B3 ──> B4 ──> B5, B6
│
└──> C1(health) ──> C1(manifest) ──> C2 ──> C3 ──> C4
A2 (independent, before A1 adds more secrets to git)
D1, D2 (independent of everything above)
```
**A1 first and alone.** Until the internal key is set, the engine is not talking
to the backend, so every Track B acceptance check would fail for the wrong
reason. Run the `kubectl` check in A1 before writing any code — if git does not
match the cluster, that changes the shape of Track C.
Cheapest real step up a level, once A1 is in: **D1 + D2.** Two executors, no new
LLM work, and it moves autonomy off the floor.
---
## Standing constraints
- Nothing in this plan is committed, pushed, or deployed without being asked.
- No migrations run against a real database without being asked.
- No files deleted or untracked without per-file approval.
- `src/lib` and `src/utils` in the console **both** have live importers — neither
tree is dead, do not consolidate them as part of this work.