# Doormile — ETA & Demand Prediction: Implementation Plan Status: **rungs 1.0–1.1 and Phase 2 built and VERIFIED against a real Postgres; not committed, not deployed** Written 2026-10-08 · Reviewed 2026-10-08 · Implemented and verified 2026-10-09 Scope: `doormile_backend`, `AI_engine`, `kubernetes` Related: `krow_talent_app/docs/agent-platform-phase7-plan.md` — Track A1 is a hard prerequisite for rung 1.1 and Track C1 for Phase 2.3 (see §9) > **Review pass (2026-10-08).** Three corrections to the first draft, all from > reading the code rather than reasoning about it: > 1. **The Phase 1.0 SQL was wrong** — it joined `so.bookingid = c.bookingid`, > and `consignments` has no `bookingid` column. Fixed via the > `consignment_booking` view; see hazard **H4**. > 2. **Routed durations are already captured** — `bookingassignments.etaminutes` > / `cumulativeeta` / `previouskms` / `cumulativekms` / `sequencedat` exist and > are written by `internal/routing`. Rung 1.1 needs no new capture plumbing. > 3. **But they are almost certainly all zeros in production** — sequencing is > gated on `ROUTE_OPTIMIZER_URL` (absent from the k8s manifest) *and* on a > rider having ≥2 stops. Rung 1.1 therefore starts with Track A1 and a > data-accumulation wait. See §3 rung 1.1. --- ## 0c. Verification log (2026-10-09) — what was actually RUN Everything below was executed, not reasoned about. Postgres 16 + pgvector 0.8.7 in Docker, the real `migrations.Migrate()`, the real Go server over HTTP. ### The migration Clean on a fresh schema, zero errors. Objects confirmed present: `consignment_booking` view · `etacalibration` · `aiskillfindings` · `demandforecast` · `agent_decisions.tenantid` · `context_embedding` as a vector type · the ivfflat index · `idx_consignmenthistory_status_consignment`. ### Every SQL statement, against REAL DATA 300 bookings over 60 days across 3 zones, 300 consignments, 277 Delivered events, 300 assignments carrying `etaminutes`, 25 resolved decisions with 1536-dim embeddings. | Statement | Result | |---|---| | `refreshSQL` (calibration) | **87 cells**, p80 factors **3.11 / 3.20** — inside the 0.5–5.0 bounds, so accepted and used | | `historicalSQL` (backfill) | **100 rows** | | `refineSQL` (ETA refine) | **23 rows** — exactly 300/13, matching the seeded `Out_for_Delivery` count. The filter is correct, not merely valid | | `pendingSQL` (outcomes) | 7 pending, delivery + SLA correctly joined | | `consignment_booking` | 300 of 300 resolvable | | similarity search | real neighbours with cosine distances, tenant-scoped | | demand series (`job.py`) | runs | | both prune deletes | run | An empty table proves syntax. These numbers prove the joins and filters. ### Every endpoint, over real HTTP `POST /internal/demand-forecast` → 2 stored · `GET /admin/ai/forecast/demand` → zone 641 expects 42, 6 riders × 5 = 30, **gap 12**; zone 500 with no hub still appears (the LEFT JOIN decision, proven) · findings upsert → written 2, re-POST upserts rather than duplicates, 1 cleared · `/acted` → recorded `partial` · findings stats → `clearedunacted: 1`, `stillopen: 1` · `POST /internal/agent-decisions` with a 1536-dim embedding → stored · `POST /admin/bookings/batch-assign` → 4 assigned, `max_per_rider: 2` respected exactly. ### The index hazard — measured, not estimated `consignmenthistory` grown to **1,000,277 rows / 69 MB**: `CREATE INDEX CONCURRENTLY` completed in **1.3 seconds**, index **valid**, and **all five concurrent INSERTs succeeded during the build**. No write blocking. Count query at 1M rows: 40ms. ### Bugs this verification found Three, none visible to `go build`, `go vet`, or the test suite: 1. **`bd.destinationid` does not exist** (it is `bookingdestinationid`). The `consignment_booking` view was never created — and `Migrate()` logs that non-fatally, so it printed "migration completed successfully" while every query joining through the view would have failed at runtime. 2. **`ORDER BY f.gap`** — `gap` is a computed alias, and qualifying it is a runtime error. `GET /admin/ai/forecast/demand` would have 500'd on every call. 3. **`POST /admin/ai/findings/{fingerprint}/acted` matched nothing.** Fiber returns the raw percent-encoded path param and the console sends `encodeURIComponent`, so `:` and `,` arrived as `%3A`/`%2C`. Because the console treats that call as fire-and-forget it would have failed **silently forever**, leaving the "did acting clear it" measurement permanently empty. `cmd/migratecheck` exists so this gap does not recur: the test suite never calls `Migrate()`, and `Migrate()` returns OK on a logged failure — so its OUTPUT must be read, not its exit code. ### Still not verified - The **Prophet path has never executed** — the library is not installed. - Rung 1.2 is not built (correctly gated on 1.0/1.1 having run). - `AI_engine`'s agents have not run against real NATS + Postgres. - Phase 0 and rung 1.0 have **not been run against production data**, so whether any of this is worth having is still unanswered. ## 0b. Phase 2 — demand forecasting, built 2026-10-08 `AI_engine/prediction/`, 23 tests passing: | File | | |---|---| | `series.py` | the dense daily series and `seasonal_naive`, the baseline every model must beat. Pure stdlib, so it works in an image without the forecasting extras | | `backtest.py` | rolling-origin validation. A tie goes to the baseline | | `demand_model.py` | Prophet, used **only** when it beats the baseline on that zone's own history | | `holidays_in.py` | regional calendar — Pongal and Onam carry as much signal here as Diwali | | `requirements-forecast.txt` | prophet/pandas/numpy, deliberately NOT in `requirements.txt` | **The design decision worth keeping:** `forecast()` backtests Prophet against `seasonal_naive` per zone and uses it only if it wins. Prophet is better on some series and worse on others, and which is which is a property of the data, not of the library. A zone where the baseline wins gets the baseline, and the returned `reason` says so, so the choice is auditable rather than implicit. **Still missing its consumer.** §2.4 of this plan says demand forecasting with no consumer is a dashboard nobody opens, and that remains true: `rebalance_riders` still has no executor, and no backend endpoint serves the forecast. The module is correct and unconsumed. **A test caught a real fixture bug worth recording:** the first synthetic series was perfectly periodic, which makes `seasonal_naive` exact (MAE 0) — so nothing could beat it and two tests failed for that reason rather than any defect. Real demand is never exactly periodic; the fixture now carries deterministic noise. ## 0a. What was built (2026-10-08) Rung 1.1's estimator and calibration, plus the Phase 0 queries. `go build ./...`, `go vet ./...` clean; `go test ./...` 19 packages pass, 0 failures. Uncommitted, not deployed, and **nothing has been run against a database**. | File | | |---|---| | `internal/prediction/eta.go` | `ETAMinutes` / `ETAAt` — the floor rule and the zone→weekday→global fallback ladder | | `internal/prediction/calibration.go` | the p80 refresh query, the store, and the in-memory snapshot | | `internal/prediction/sweeper.go` | 6h ticker + Redis lock, same shape as `internal/assignment/sweeper.go` | | `internal/prediction/eta_test.go` | 17 tests, all passing | | `migrations/migrate.go` | `consignment_booking` view (H4), `etacalibration` table, two indexes | | `main.go` | `go prediction.StartCalibrationSweeper()` | | `scratch/prediction_readiness.sql` | Phase 0 + rung 1.0, read-only | **Three bugs were found and fixed while building, two of them mine:** 1. **The refresh query multiplied every delivery by its assignment history.** A booking holds several `bookingassignments` rows (Assigned, Rejected, Reassigned), so a plain join overstated sample counts and pulled the p80 toward whatever got reassigned most. Fixed with `DISTINCT ON (bookingid)` taking the most recently sequenced row. The same bug was in readiness query Q10; fixed there too so the check predicts production behaviour. 2. **Hazard H1, in this package's own code.** `Refreshedat` is written through `utils.DBNow` (IST digits labelled UTC), so reading it as a raw instant makes a calibration look 5h30m *newer* than it is — a 73h-old table measures as 67.5h and slips under the 72h staleness bound. `Load` now corrects through `utils.IST`. Confirmed by mutation: removing the call fails `TestStaleCalibrationFallsBack`. 3. **Untyped `NULL` in a `UNION ALL`.** Postgres resolves column types across branches and can infer an untyped NULL as text, clashing with the integer from the first branch — a failure that only appears at runtime against the real database. Now `NULL::int` explicitly. **Not done, deliberately — see §3a for why:** the two existing ETA call sites (`adminController.go:2738`, `cxPickupFanout.go:224`) are untouched. Wiring them would have been dead code, and the real integration point needs a product decision. ## 0. Confirmed not implemented Verified by grep across `AI_engine`, `doormile_backend`, `krow_talent_app/src`: | | State | |---|---| | LSTM | not present | | ARIMA / SARIMA / Prophet | not present | | Time-series prediction | not present | | ETA prediction (learned) | not present | | Demand prediction | not present | No ML dependency exists anywhere — no `torch`, `tensorflow`, `sklearn`, `statsmodels`, `xgboost`, `prophet`, not even `numpy`/`pandas` in `AI_engine/requirements.txt`. Zero source matches for `lstm`, `arima`, `forecast`, `time series`. **What does exist, and is the baseline any model must beat:** - **Express/admin ETA** — flat constants, `controllers/adminController.go:2740`: Standard 36h SLA, Fast 12h ETA / 18h SLA, Superfast 6h / 9h. No data input. - **Customer-app ETA** — district promise lookup, `controllers/cxPickupFanout.go:224`: `ServiceableDistrict.promise` → 0/1/2/3 days, delivered-by-8pm. - **`consignments.estimateddeliveryat` and `sladueat`** columns exist and are populated from those two rules. - **Demand** — exists only as text describing absent features. `internal/ai/registry/seed.go:295`: *"Move idle riders into zones with a demand spike. Disabled until an endpoint implements it."* (`rebalance_riders`, `Target: "none yet"`). --- ## 1. The framing: these are two different problems The question "LSTM or ARIMA" assumes one model family covers both. It does not, and picking the wrong family is the most expensive mistake available here. | | ETA prediction | Demand prediction | |---|---|---| | **Problem type** | supervised **regression** — one row per trip | **time series** — counts per zone per interval | | **Input** | distance, hour, zone, rider, weight, attempt count | history of its own past values | | **Right family** | gradient boosting / quantile regression | SARIMA or Prophet | | **ARIMA fit?** | **no** — there is no series, each trip is independent | yes | | **LSTM fit?** | no — tabular, boosting wins | only at volume we almost certainly don't have | **ETA is not a time-series problem.** A delivery's duration depends on its own features, not on the duration of the delivery before it. Fitting ARIMA to trip durations models an ordering that carries no signal. This is the single most common mistake in logistics ML and it is worth stating plainly before any code. **Demand is a time-series problem** — bookings per pincode per day is a real series with weekly seasonality and holiday effects, which is exactly what SARIMA and Prophet are for. **LSTM is almost certainly wrong for both.** It needs tens of thousands of sequences to beat SARIMA on a univariate count series, and it loses to gradient boosting on tabular regression. Section 6 states the condition under which it would become worth revisiting; until that condition is measured and met, building it is cost without benefit. --- ## 2. Phase 0 — Data readiness gate (**this phase can return "don't build it yet"**) Everything downstream depends on history that may not exist. CLAUDE.md §9 lists go-live across the four cities as still ahead; if real delivery volume hasn't accumulated, there is nothing to fit and Phase 0 is the whole project for now. Run these before writing any model code. ### 2.1 Volume and span ```sql -- Delivered consignments and how far back they go. SELECT count(*) AS delivered, min(createdat)::date AS first_day, max(createdat)::date AS last_day, count(DISTINCT createdat::date) AS distinct_days FROM consignmenthistory WHERE eventstatus = 'Delivered'; -- Per-pincode daily series length (demand needs this per series, not in total). SELECT c.deliverypincode, count(DISTINCT h.createdat::date) AS days_with_data, count(*) AS deliveries FROM consignmenthistory h JOIN consignments c ON c.consignmentid = h.consignmentid WHERE h.eventstatus = 'Delivered' GROUP BY 1 ORDER BY 3 DESC LIMIT 30; ``` **Go/no-go thresholds:** | | ETA regression | Demand SARIMA | |---|---|---| | Minimum usable | ~2,000 completed trips | ~90 days per series | | Comfortable | ~10,000+ | ~1 year (two seasonal cycles) | | Below minimum | keep the promise table | aggregate to city level, or wait | If per-pincode series are too short, **aggregate up** — city-level demand on 90 days is forecastable where pincode-level on 90 days is noise. ### 2.2 The three data hazards (all verified in this codebase) **H1 — Timestamps are IST wall-clock digits labelled UTC.** `utils.DBNow` stores IST digits with a UTC tag. CLAUDE.md §8.5 is explicit: `t.UnixMilli()` is off by 5h30m, and this already caused "yesterday's work shown as today" on the miler app. `models/ai_runs.go:26` documents the same trap. This matters more for prediction than anywhere else, because **hour-of-day is a primary ETA feature and day-boundary is the demand bucket.** Fitted on raw `createdat`, hour-of-day is right by accident and the daily bucket is wrong for every event between 18:30 and 00:00 IST. Use `utils.EpochMillis` (`utils/epoch.go`) as the single conversion, exactly as the customer surface does. Write one `trip_features` SQL view that does the correction once, and let every model read the view — never raw columns. **H2 — There is no `deliveredat` column.** `consignments` has `createdat`, `inwardedat`, `estimateddeliveryat`, `sladueat`, `returninitiatedat`, `returndeliveredat` — but **no delivery completion timestamp**. Status reaches `Delivered`; the time only exists as `consignmenthistory.createdat WHERE eventstatus = 'Delivered'`. So the training label is derived, not stored. Verify it is reliably written before trusting it: ```sql SELECT count(*) FILTER (WHERE h.consignmentid IS NULL) AS delivered_without_event FROM consignments c LEFT JOIN consignmenthistory h ON h.consignmentid = c.consignmentid AND h.eventstatus = 'Delivered' WHERE c.status = 'Delivered'; ``` Non-zero means silent label loss — fix the write path before fitting anything. **H4 — `consignments` has no `bookingid`. Every join through it is a two-path resolution.** This is the trap CLAUDE.md §8.5 warns about, and the first draft of this document walked straight into it. The link is `bookingdestinations.consignmentid → bookingdestinations.bookingid` for multi-destination pickups, falling back to `pickupbookings.consignmentid` for console/express bookings and any row written before the fan-out existed — and that legacy column **names only the FIRST order** of a multi-destination pickup. `cxDestinationForConsignment` (`controllers/cxPickupFanout.go:191`) is the canonical resolver in Go. Get this wrong and the service-type and promise features silently attach to the wrong parcel for every multi-destination booking. Resolve it **once**, in the view, and never join through `consignments` directly: ```sql CREATE OR REPLACE VIEW consignment_booking AS SELECT c.consignmentid, COALESCE(bd.bookingid, pb.bookingid) AS bookingid, bd.bookingdestinationid FROM consignments c LEFT JOIN bookingdestinations bd ON bd.consignmentid = c.consignmentid LEFT JOIN pickupbookings pb ON pb.consignmentid = c.consignmentid AND bd.bookingid IS NULL; ``` Verify the resolution covers everything before trusting it: ```sql SELECT count(*) AS unresolvable FROM consignment_booking WHERE bookingid IS NULL; ``` **H3 — Fabricated timestamps are a known pattern in this estate.** DailyGrubs' order data has invented delivery timestamps and a double-labelled timezone. Doormile is a different database, but the same team and the same `DBNow` convention. Spot-check that delivery times aren't clustered on suspiciously round values or identical offsets from `createdat` before fitting. **A model trained on fabricated timestamps predicts confidently and wrongly** — and unlike a broken query, nothing surfaces the error. ### 2.3 Deliverable A one-page readiness note: row counts, series lengths, H1/H2/H3 findings, and a go/no-go per track. If it says no-go, stop here — Phase 1.0 (below) still pays for itself and needs no history. --- ## 3. Phase 1 — ETA A ladder. Each rung ships, is measured against the one below, and is only climbed if the measurement justifies it. ### 1.0 — Measure the current promise table (**no ML, do this regardless**) Before predicting anything, find out how wrong the constants already are: ```sql -- Promise vs actual, per service type. -- Joins through consignment_booking (H4) — NEVER c.bookingid, which does not exist. SELECT so.servicetype, count(*) AS n, avg(EXTRACT(EPOCH FROM (h.createdat - c.createdat))/3600) AS actual_hours_avg, avg(EXTRACT(EPOCH FROM (c.estimateddeliveryat - c.createdat))/3600) AS promised_hours_avg, count(*) FILTER (WHERE h.createdat > c.sladueat) AS sla_breaches FROM consignments c JOIN consignmenthistory h ON h.consignmentid = c.consignmentid AND h.eventstatus = 'Delivered' JOIN consignment_booking cb ON cb.consignmentid = c.consignmentid LEFT JOIN bookingserviceoptions so ON so.bookingid = cb.bookingid GROUP BY 1; ``` Table and column names verified against the models: `bookingserviceoptions` (`models/booking.go:198`), `serviceabledistricts` (`models/customer_app.go:60`), `consignmenthistory` (`models/audit.go:109`). Two outcomes, both useful. If the constants are close, **there is no ETA problem to solve** and the honest answer is to stop. If they are badly off, this query gives the baseline error that every later rung must beat — and it may be fixable by re-tuning the constants per service type and district, which is an afternoon's work rather than a model. ### 1.1 — Routing-based ETA (**the rung most likely to be the right stopping point**) **Better news than the first draft assumed, with a catch.** `internal/routing` is a client for a real Valhalla-backed road-network API (`routes.workolik.com /api/v1/optimization/doormile/sequence`), and its response already carries durations, not just an ordering: `etaminutes`, `cumulativeeta`, `previouskms`, `cumulativekms`, `totaleta` (`internal/routing/optimizer.go:81-95`). And **those values are already persisted**, on `bookingassignments` (`models/booking.go:239-244`): `step`, `previouskms`, `cumulativekms`, `etaminutes`, `cumulativeeta`, `sequencedat`. So a predicted-vs-actual dataset — routed ETA at assignment time paired with the delivery event — is structurally already being collected. No new capture plumbing is needed. **The catch, and it is a real one: those columns are almost certainly all zeros in production.** Two independent gates: 1. **`routing.BaseURL` empty disables sequencing entirely** (deliberately — `optimizer.go:26`). It is set from `cfg.RouteOptimizerURL` at `main.go:246`, and `ROUTE_OPTIMIZER_URL` is one of the 19 env vars missing from `kubernetes/manifests/doormile/miletruth.yaml` — see Track A1 of the Phase 7 plan. CLAUDE.md §8.5 says the same thing from the other side: *"Sequencing … is not deployed yet, so riders with more than one stop come back unsequenced (step: 0)."* 2. **`minStopsToSequence = 2`** (`optimizer.go:44`) — a rider with one stop is never sequenced. In a courier operation a large share of assignments may be single-stop, so even with routing switched on, routed ETAs cover only multi-stop riders. **Two consequences for the plan:** - **Rung 1.1 has a hard prerequisite: Phase 7 Track A1.** Until `ROUTE_OPTIMIZER_URL` is in the manifest, there is no routed duration to calibrate, and `etaminutes` stays 0. Verify before building: ```sql SELECT count(*) AS assignments, count(*) FILTER (WHERE etaminutes > 0) AS with_routed_eta, count(*) FILTER (WHERE step > 0) AS sequenced, min(sequencedat), max(sequencedat) FROM bookingassignments; ``` `with_routed_eta = 0` means this rung starts by turning routing on and waiting for data, not by fitting a calibration. - **Single-stop assignments need their own duration source.** Haversine × a learned road-circuity factor per zone is the pragmatic fallback — `haversineKM` already exists once, in `hubController.go` (do not redefine it; CLAUDE.md §7). Alternatively call the routing API for single stops too, which is a change to `minStopsToSequence`'s rationale and should be decided explicitly rather than assumed. With a routed duration in hand, the calibration is a grouped median over history and not machine learning: ``` eta = routed_duration × calibration[zone, hour_bucket, weekday] + handling_time[hub] ``` One SQL query, refreshed nightly. Interpretable, debuggable, no training infrastructure, no model server — and because `etaminutes` is already stored per assignment, the calibration is fitted on the system's own past predictions against its own actuals, which is the cleanest possible training signal. **Only climb past this rung if 1.1 is measurably insufficient.** For a single-city hub-based courier, it very often isn't. ### 3a. Where the calibrated ETA can actually be applied — **decision needed** Found while implementing, and it changes the integration plan. **At booking-create time there is no routed duration.** The sequence is: booking created → `estimateddeliveryat` written from the promise table → rider assigned → `internal/routing` sequences and writes `etaminutes`. The routed number arrives *after* the promise has been set and shown. So wiring `prediction.ETAMinutes` into `adminController.go:2738` or `cxPickupFanout.go:224` would return `false` on every call, forever — not because the calibration is cold, but because `RoutedMinutes` is structurally 0 at that moment. That is dead code, so those call sites were left alone. The routed ETA first exists at **sequencing time**, which means applying it is revising a promise the customer has already been given. That is a product decision, not a wiring one: | | | |---|---| | **Option A — refine `estimateddeliveryat`, never touch `sladueat`** | The customer's "arrives by" sharpens as the system learns more; the commitment they were given does not move. **Recommended.** | | Option B — leave both, expose the calibrated ETA only on tracking | Nothing stored changes; the sharper number is display-only. Safest, least useful. | | Option C — revise both | The SLA stops being a commitment. Not recommended. | Option A needs one call in `internal/routing` after the ETA columns are written, plus a decision on whether `cxstage` should emit an event when a promise moves — a customer watching the tracking page will see the time change, and silently is probably the wrong way for that to happen. **This is decision 7 in §8.** Nothing should be wired until it is made. ### 1.2 — Gradient-boosted regression (only if 1.1 is insufficient) Features, all already in the schema: | Feature | Source | |---|---| | haversine + routed distance | `pickuplatitude/longitude`, `deliverylatitude/longitude` | | hour of day, weekday | `createdat` **via `EpochMillis`** (H1) | | pickup / delivery pincode | `pickuppincode`, `deliverypincode` | | origin / destination hub | `originhubid`, `destinationhubid` | | chargeable weight | `chargeableweight` | | service type | `bookingserviceoptions.servicetype` | | attempt count | `consignments.attemptcount` | | rider | `assignedmileruserid` (hash, not identity — see §8) | | hub inbound load at assignment | derived from `consignmenthistory` | **Predict a quantile, not a mean.** An ETA shown to a customer should be the p80 — "arrives by" — not the average, which is late half the time. Use quantile regression or a boosted model with a quantile objective. This single choice matters more to perceived accuracy than the model family. Label: `delivered_event.createdat − consignment.createdat`, both corrected for H1. ### 1.3 — Attempt-aware ETA `attemptcount` exists and `MilerSkipDelivery` increments it, so failed attempts are recorded. An ETA that ignores re-attempts is wrong for exactly the parcels customers complain about. Worth a separate model only once 1.2 is in production and its residuals show re-attempts as the dominant error mode. --- ## 4. Phase 2 — Demand ### 2.1 — The series ```sql CREATE OR REPLACE VIEW demand_daily AS SELECT (createdat)::date AS day, -- H1 correction applied in the real view pickuppincode, count(*) AS bookings FROM pickupbookings WHERE status <> 'Cancelled' GROUP BY 1, 2; ``` Decide the grain deliberately: pincode × day is what `rebalance_riders` wants, but it is also the sparsest. Start at **city × day**, prove the pipeline, then descend to zone only where series length supports it. ### 2.2 — Baseline first Seasonal naïve — "same weekday last week" — is the baseline. It is one line of SQL and it beats badly-configured SARIMA routinely. Any model that cannot beat it on held-out data does not ship. ### 2.3 — SARIMA or Prophet | | Choose when | |---|---| | **SARIMA** (`statsmodels`) | few series, weekly seasonality, want interpretable orders | | **Prophet** | many series, holidays matter (Indian festival calendar is a real effect on courier volume), need it to work without per-series tuning | **Recommendation: Prophet**, for the holiday regressors. Diwali, Pongal and regional festivals move courier volume substantially, Prophet takes a holiday calendar as a first-class input, and it does not need per-series order selection across dozens of pincodes. Validate with rolling-origin backtesting (expanding window), **never** a random split — a random train/test split on time series leaks the future and reports an accuracy you will not see in production. ### 2.4 — What consumes the forecast Demand prediction with no consumer is a dashboard nobody opens. The honest consumer is `rebalance_riders` (`seed.go:295`), which is itself unimplemented — so Phase 2 should be scoped **with** that tool's executor or not at all. Minimum useful output: tomorrow's expected bookings per zone, plus a staffing-gap signal against rostered riders. That is actionable; a forecast number alone is not. --- ## 5. Where this runs A prediction service is a **third** runtime next to the Go API and the Python agents. Options, cheapest first: | Option | Shape | Cost | |---|---|---| | **A — SQL + nightly job** | Calibration tables computed by a Go sweeper; serving is a table lookup | no new runtime, no new image | | **B — module inside `AI_engine`** | New package; adds `pandas`/`statsmodels`/`prophet` to the image | one runtime, image grows ~300MB | | **C — separate service** | Own repo/image/deploy/probes | full operational cost | **Recommendation: A for Phase 1.1, B for Phase 2.** Rung 1.1 needs no model server at all — it is a calibration table and a multiply, so it belongs in the Go backend as a sweeper beside `StartPendingSweeper` (`main.go:244`). Prophet genuinely needs Python, so Phase 2 lands in `AI_engine`. Option C only becomes right if Phase 1.2 happens and model serving needs independent scaling. **Prerequisite from the Phase 7 plan:** `AI_engine` is not in Kubernetes and has no HTTP health surface (Track C1 there). Phase 2 inherits that work — it cannot deploy before it. --- ## 6. LSTM — the condition for revisiting Not recommended now. The condition under which it becomes worth measuring: - Phase 2.3 is in production, backtested, and **losing to its own residual structure** — i.e. Prophet's errors are autocorrelated in a way a sequence model could capture; and - ≥ 2 years of daily data across ≥ 50 series (≈ 36,000 observations), and - a measured business cost to the remaining forecast error that exceeds the cost of training infrastructure, GPU or CPU-hours, and the ongoing retraining a neural model needs to not rot. All three, not any one. Until then an LSTM here would be a more expensive way to get a worse number, and the honest recommendation is to say so rather than build it. --- ## 7. File manifest ### Phase 0 — data readiness (no application code) | File | New? | |---|---| | `doormile_backend/docs/prediction-data-readiness.md` | **new** — the findings note | | `doormile_backend/scratch/readiness_queries.sql` | **new** — the queries above, kept for re-running | ### Phase 1.0 / 1.1 — measurement and routing ETA | File | New? | Change | |---|---|---| | `doormile_backend/migrations/migrate.go` | | add the `consignment_booking` view (H4), the `trip_features` view (H1 correction in one place), and the `eta_calibration` table | | `doormile_backend/internal/prediction/calibration.go` | **new** | nightly grouped-median refresh | | `doormile_backend/internal/prediction/eta.go` | **new** | `EstimateETA(booking) (time.Time, confidence)` | | `doormile_backend/internal/prediction/eta_test.go` | **new** | falls back to the promise table when calibration is missing | | `doormile_backend/internal/prediction/sweeper.go` | **new** | ticker + Redis lock, pattern from `internal/assignment/sweeper.go:88` | | `doormile_backend/main.go:244` | | `go prediction.StartCalibrationSweeper()` | | `doormile_backend/controllers/adminController.go:2740` | | call `prediction.EstimateETA`, keep constants as fallback | | `doormile_backend/controllers/cxPickupFanout.go:224` | | same, keep the promise table as fallback | | `doormile_backend/utils/epoch.go` | | read-only — the H1 conversion to reuse | ### Phase 1.2 — learned ETA (only if 1.1 insufficient) | File | New? | |---|---| | `AI_engine/prediction/__init__.py` · `eta_model.py` · `features.py` | **new** | | `AI_engine/prediction/train_eta.py` | **new** — offline training, writes a versioned artifact | | `AI_engine/tests/test_eta_features.py` | **new** — H1 correction asserted on both timestamp taggings | | `AI_engine/requirements.txt` | modify — `pandas`, `scikit-learn` or `lightgbm` | | `doormile_backend/internal/prediction/eta.go` | modify — call the service, fall back to 1.1 | ### Phase 2 — demand | File | New? | |---|---| | `AI_engine/prediction/demand_model.py` | **new** | | `AI_engine/prediction/holidays_in.py` | **new** — the festival calendar | | `AI_engine/prediction/backtest.py` | **new** — rolling-origin, never a random split | | `AI_engine/tests/test_demand_backtest.py` | **new** — must beat seasonal-naïve to pass | | `AI_engine/requirements.txt` | modify — `prophet` or `statsmodels` | | `AI_engine/Dockerfile` | modify — Prophet needs a compiler toolchain | | `doormile_backend/migrations/migrate.go` | modify — `demand_daily` view, `demand_forecast` table | | `doormile_backend/routes/routes.go` | modify — `GET /admin/forecast/demand` (staff-only) | | `doormile_backend/controllers/forecastController.go` | **new** | | `doormile_backend/internal/ai/registry/seed.go:295` | modify — `rebalance_riders` once it has a consumer | | `kubernetes/manifests/doormile/ai-engine.yaml` | modify — resources for Prophet | ### Console (Phase 2 only, optional) | File | Change | |---|---| | `krow_talent_app/src/api/doormile/endpoints.js` | add `getDemandForecast` | | a new forecast panel | render it — **after** a consumer exists, not before | --- ## 8. Decisions needed 0. **Is `ROUTE_OPTIMIZER_URL` set in the cluster?** If not, `etaminutes` is 0 everywhere and rung 1.1 begins with Phase 7 Track A1 plus a data-accumulation wait, not with a calibration. This gates more than anything else here. 1. **Is there enough history?** Phase 0 answers it. Everything else is blocked on that number. 1b. **Do single-stop assignments get routed too?** `minStopsToSequence = 2` excludes them today. Either lower it, or accept a haversine-based fallback for single-stop ETAs. An explicit call, not an assumption. 2. **Does ETA need to be learned at all, or do the constants just need re-tuning?** Rung 1.0 answers it, cheaply. 3. **p80 or mean ETA?** Recommend p80 — "arrives by" is the promise customers hear, and a mean is late half the time. 4. **Demand grain** — city × day to start, or straight to pincode × day? Recommend city first. 5. **Does `rebalance_riders` get an executor in the same phase?** If no, Phase 2 produces a number nobody acts on. 6. **Rider as a feature (1.2)** — per-rider ETA adjustment is a performance signal about a named person. Hash it, use it only in aggregate, and decide deliberately whether it may ever surface in the console. This is a people decision, not a modelling one. 7. **May a promise be revised after the customer has seen it?** §3a. Blocks the last wiring step of rung 1.1 — everything else is built. Recommend Option A: refine `estimateddeliveryat`, never move `sladueat`. --- ## 9. Sequencing ``` Phase 7 A1 (env vars) ──> routing actually runs ──> etaminutes accumulates │ │ ▼ ▼ Phase 0 (readiness) ──> [go / no-go] 1.1 calibration possible │ no-go ─────────────┴──> stop; re-tune constants (1.0) and revisit after go-live │ go ──> 1.0 measure ──> 1.1 routing ETA ──> [measure] ──> 1.2 only if needed └──> 2.1/2.2 baseline ──> 2.3 Prophet (needs Phase 7 C1 first) ``` **Both tracks now depend on the Phase 7 plan**, for different reasons: rung 1.1 needs `ROUTE_OPTIMIZER_URL` (Track A1) before routed ETAs exist at all, and Phase 2.3 needs `AI_engine` deployable with a health surface (Track C1). Track A1 is one manifest edit and unblocks both — do it first regardless of which track you want. **Start with Phase 0 and rung 1.0.** Both are SQL, neither needs a model, and together they either justify the rest of this plan or retire it. Rung 1.1 is the highest-leverage item in the document and is not machine learning at all. The likeliest honest outcome: **1.0 + 1.1 ship, 1.2 is never needed, Phase 2 waits for volume, and LSTM never happens.** That is a success, not a shortfall. --- ## Standing constraints - Nothing committed, pushed or deployed without being asked. - No migrations run against a real database without being asked — the views and tables here are additive, but "additive" is not "has run". - Phase 0's queries are read-only; they are safe to run, and should be run before anything else in this document.