Files
doormile_backend/docs/prediction-plan.md

737 lines
35 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Doormile — ETA & Demand Prediction: Implementation Plan
Status: **rungs 1.0–1.1 and Phase 2 built and VERIFIED against a real Postgres; not committed, not deployed**
Written 2026-10-08 · Reviewed 2026-10-08 · Implemented and verified 2026-10-09
Scope: `doormile_backend`, `AI_engine`, `kubernetes`
Related: `krow_talent_app/docs/agent-platform-phase7-plan.md` — Track A1 is a hard
prerequisite for rung 1.1 and Track C1 for Phase 2.3 (see §9)
> **Review pass (2026-10-08).** Three corrections to the first draft, all from
> reading the code rather than reasoning about it:
> 1. **The Phase 1.0 SQL was wrong** — it joined `so.bookingid = c.bookingid`,
> and `consignments` has no `bookingid` column. Fixed via the
> `consignment_booking` view; see hazard **H4**.
> 2. **Routed durations are already captured** — `bookingassignments.etaminutes`
> / `cumulativeeta` / `previouskms` / `cumulativekms` / `sequencedat` exist and
> are written by `internal/routing`. Rung 1.1 needs no new capture plumbing.
> 3. **But they are almost certainly all zeros in production** — sequencing is
> gated on `ROUTE_OPTIMIZER_URL` (absent from the k8s manifest) *and* on a
> rider having ≥2 stops. Rung 1.1 therefore starts with Track A1 and a
> data-accumulation wait. See §3 rung 1.1.
---
## 0c. Verification log (2026-10-09) — what was actually RUN
Everything below was executed, not reasoned about. Postgres 16 + pgvector 0.8.7
in Docker, the real `migrations.Migrate()`, the real Go server over HTTP.
### The migration
Clean on a fresh schema, zero errors. Objects confirmed present:
`consignment_booking` view · `etacalibration` · `aiskillfindings` ·
`demandforecast` · `agent_decisions.tenantid` · `context_embedding` as a vector
type · the ivfflat index · `idx_consignmenthistory_status_consignment`.
### Every SQL statement, against REAL DATA
300 bookings over 60 days across 3 zones, 300 consignments, 277 Delivered
events, 300 assignments carrying `etaminutes`, 25 resolved decisions with
1536-dim embeddings.
| Statement | Result |
|---|---|
| `refreshSQL` (calibration) | **87 cells**, p80 factors **3.11 / 3.20** — inside the 0.5–5.0 bounds, so accepted and used |
| `historicalSQL` (backfill) | **100 rows** |
| `refineSQL` (ETA refine) | **23 rows** — exactly 300/13, matching the seeded `Out_for_Delivery` count. The filter is correct, not merely valid |
| `pendingSQL` (outcomes) | 7 pending, delivery + SLA correctly joined |
| `consignment_booking` | 300 of 300 resolvable |
| similarity search | real neighbours with cosine distances, tenant-scoped |
| demand series (`job.py`) | runs |
| both prune deletes | run |
An empty table proves syntax. These numbers prove the joins and filters.
### Every endpoint, over real HTTP
`POST /internal/demand-forecast` → 2 stored · `GET /admin/ai/forecast/demand`
→ zone 641 expects 42, 6 riders × 5 = 30, **gap 12**; zone 500 with no hub
still appears (the LEFT JOIN decision, proven) · findings upsert → written 2,
re-POST upserts rather than duplicates, 1 cleared · `/acted` → recorded
`partial` · findings stats → `clearedunacted: 1`, `stillopen: 1` ·
`POST /internal/agent-decisions` with a 1536-dim embedding → stored ·
`POST /admin/bookings/batch-assign` → 4 assigned, `max_per_rider: 2` respected
exactly.
### The index hazard — measured, not estimated
`consignmenthistory` grown to **1,000,277 rows / 69 MB**:
`CREATE INDEX CONCURRENTLY` completed in **1.3 seconds**, index **valid**, and
**all five concurrent INSERTs succeeded during the build**. No write blocking.
Count query at 1M rows: 40ms.
### Bugs this verification found
Three, none visible to `go build`, `go vet`, or the test suite:
1. **`bd.destinationid` does not exist** (it is `bookingdestinationid`). The
`consignment_booking` view was never created — and `Migrate()` logs that
non-fatally, so it printed "migration completed successfully" while every
query joining through the view would have failed at runtime.
2. **`ORDER BY f.gap`** — `gap` is a computed alias, and qualifying it is a
runtime error. `GET /admin/ai/forecast/demand` would have 500'd on every
call.
3. **`POST /admin/ai/findings/{fingerprint}/acted` matched nothing.** Fiber
returns the raw percent-encoded path param and the console sends
`encodeURIComponent`, so `:` and `,` arrived as `%3A`/`%2C`. Because the
console treats that call as fire-and-forget it would have failed **silently
forever**, leaving the "did acting clear it" measurement permanently empty.
`cmd/migratecheck` exists so this gap does not recur: the test suite never
calls `Migrate()`, and `Migrate()` returns OK on a logged failure — so its
OUTPUT must be read, not its exit code.
### Still not verified
- The **Prophet path has never executed** — the library is not installed.
- Rung 1.2 is not built (correctly gated on 1.0/1.1 having run).
- `AI_engine`'s agents have not run against real NATS + Postgres.
- Phase 0 and rung 1.0 have **not been run against production data**, so
whether any of this is worth having is still unanswered.
## 0b. Phase 2 — demand forecasting, built 2026-10-08
`AI_engine/prediction/`, 23 tests passing:
| File | |
|---|---|
| `series.py` | the dense daily series and `seasonal_naive`, the baseline every model must beat. Pure stdlib, so it works in an image without the forecasting extras |
| `backtest.py` | rolling-origin validation. A tie goes to the baseline |
| `demand_model.py` | Prophet, used **only** when it beats the baseline on that zone's own history |
| `holidays_in.py` | regional calendar — Pongal and Onam carry as much signal here as Diwali |
| `requirements-forecast.txt` | prophet/pandas/numpy, deliberately NOT in `requirements.txt` |
**The design decision worth keeping:** `forecast()` backtests Prophet against
`seasonal_naive` per zone and uses it only if it wins. Prophet is better on some
series and worse on others, and which is which is a property of the data, not of
the library. A zone where the baseline wins gets the baseline, and the returned
`reason` says so, so the choice is auditable rather than implicit.
**Still missing its consumer.** §2.4 of this plan says demand forecasting with
no consumer is a dashboard nobody opens, and that remains true: `rebalance_riders`
still has no executor, and no backend endpoint serves the forecast. The module
is correct and unconsumed.
**A test caught a real fixture bug worth recording:** the first synthetic series
was perfectly periodic, which makes `seasonal_naive` exact (MAE 0) — so nothing
could beat it and two tests failed for that reason rather than any defect. Real
demand is never exactly periodic; the fixture now carries deterministic noise.
## 0a. What was built (2026-10-08)
Rung 1.1's estimator and calibration, plus the Phase 0 queries. `go build ./...`,
`go vet ./...` clean; `go test ./...` 19 packages pass, 0 failures. Uncommitted,
not deployed, and **nothing has been run against a database**.
| File | |
|---|---|
| `internal/prediction/eta.go` | `ETAMinutes` / `ETAAt` — the floor rule and the zone→weekday→global fallback ladder |
| `internal/prediction/calibration.go` | the p80 refresh query, the store, and the in-memory snapshot |
| `internal/prediction/sweeper.go` | 6h ticker + Redis lock, same shape as `internal/assignment/sweeper.go` |
| `internal/prediction/eta_test.go` | 17 tests, all passing |
| `migrations/migrate.go` | `consignment_booking` view (H4), `etacalibration` table, two indexes |
| `main.go` | `go prediction.StartCalibrationSweeper()` |
| `scratch/prediction_readiness.sql` | Phase 0 + rung 1.0, read-only |
**Three bugs were found and fixed while building, two of them mine:**
1. **The refresh query multiplied every delivery by its assignment history.** A
booking holds several `bookingassignments` rows (Assigned, Rejected,
Reassigned), so a plain join overstated sample counts and pulled the p80
toward whatever got reassigned most. Fixed with `DISTINCT ON (bookingid)`
taking the most recently sequenced row. The same bug was in readiness query
Q10; fixed there too so the check predicts production behaviour.
2. **Hazard H1, in this package's own code.** `Refreshedat` is written through
`utils.DBNow` (IST digits labelled UTC), so reading it as a raw instant makes
a calibration look 5h30m *newer* than it is — a 73h-old table measures as
67.5h and slips under the 72h staleness bound. `Load` now corrects through
`utils.IST`. Confirmed by mutation: removing the call fails
`TestStaleCalibrationFallsBack`.
3. **Untyped `NULL` in a `UNION ALL`.** Postgres resolves column types across
branches and can infer an untyped NULL as text, clashing with the integer
from the first branch — a failure that only appears at runtime against the
real database. Now `NULL::int` explicitly.
**Not done, deliberately — see §3a for why:** the two existing ETA call sites
(`adminController.go:2738`, `cxPickupFanout.go:224`) are untouched. Wiring them
would have been dead code, and the real integration point needs a product
decision.
## 0. Confirmed not implemented
Verified by grep across `AI_engine`, `doormile_backend`, `krow_talent_app/src`:
| | State |
|---|---|
| LSTM | not present |
| ARIMA / SARIMA / Prophet | not present |
| Time-series prediction | not present |
| ETA prediction (learned) | not present |
| Demand prediction | not present |
No ML dependency exists anywhere — no `torch`, `tensorflow`, `sklearn`,
`statsmodels`, `xgboost`, `prophet`, not even `numpy`/`pandas` in
`AI_engine/requirements.txt`. Zero source matches for `lstm`, `arima`,
`forecast`, `time series`.
**What does exist, and is the baseline any model must beat:**
- **Express/admin ETA** — flat constants, `controllers/adminController.go:2740`:
Standard 36h SLA, Fast 12h ETA / 18h SLA, Superfast 6h / 9h. No data input.
- **Customer-app ETA** — district promise lookup, `controllers/cxPickupFanout.go:224`:
`ServiceableDistrict.promise` → 0/1/2/3 days, delivered-by-8pm.
- **`consignments.estimateddeliveryat` and `sladueat`** columns exist and are
populated from those two rules.
- **Demand** — exists only as text describing absent features.
`internal/ai/registry/seed.go:295`: *"Move idle riders into zones with a
demand spike. Disabled until an endpoint implements it."* (`rebalance_riders`,
`Target: "none yet"`).
---
## 1. The framing: these are two different problems
The question "LSTM or ARIMA" assumes one model family covers both. It does not,
and picking the wrong family is the most expensive mistake available here.
| | ETA prediction | Demand prediction |
|---|---|---|
| **Problem type** | supervised **regression** — one row per trip | **time series** — counts per zone per interval |
| **Input** | distance, hour, zone, rider, weight, attempt count | history of its own past values |
| **Right family** | gradient boosting / quantile regression | SARIMA or Prophet |
| **ARIMA fit?** | **no** — there is no series, each trip is independent | yes |
| **LSTM fit?** | no — tabular, boosting wins | only at volume we almost certainly don't have |
**ETA is not a time-series problem.** A delivery's duration depends on its own
features, not on the duration of the delivery before it. Fitting ARIMA to trip
durations models an ordering that carries no signal. This is the single most
common mistake in logistics ML and it is worth stating plainly before any code.
**Demand is a time-series problem** — bookings per pincode per day is a real
series with weekly seasonality and holiday effects, which is exactly what SARIMA
and Prophet are for.
**LSTM is almost certainly wrong for both.** It needs tens of thousands of
sequences to beat SARIMA on a univariate count series, and it loses to gradient
boosting on tabular regression. Section 6 states the condition under which it
would become worth revisiting; until that condition is measured and met,
building it is cost without benefit.
---
## 2. Phase 0 — Data readiness gate (**this phase can return "don't build it yet"**)
Everything downstream depends on history that may not exist. CLAUDE.md §9 lists
go-live across the four cities as still ahead; if real delivery volume hasn't
accumulated, there is nothing to fit and Phase 0 is the whole project for now.
Run these before writing any model code.
### 2.1 Volume and span
```sql
-- Delivered consignments and how far back they go.
SELECT count(*) AS delivered,
min(createdat)::date AS first_day,
max(createdat)::date AS last_day,
count(DISTINCT createdat::date) AS distinct_days
FROM consignmenthistory
WHERE eventstatus = 'Delivered';
-- Per-pincode daily series length (demand needs this per series, not in total).
SELECT c.deliverypincode,
count(DISTINCT h.createdat::date) AS days_with_data,
count(*) AS deliveries
FROM consignmenthistory h
JOIN consignments c ON c.consignmentid = h.consignmentid
WHERE h.eventstatus = 'Delivered'
GROUP BY 1 ORDER BY 3 DESC LIMIT 30;
```
**Go/no-go thresholds:**
| | ETA regression | Demand SARIMA |
|---|---|---|
| Minimum usable | ~2,000 completed trips | ~90 days per series |
| Comfortable | ~10,000+ | ~1 year (two seasonal cycles) |
| Below minimum | keep the promise table | aggregate to city level, or wait |
If per-pincode series are too short, **aggregate up** — city-level demand on 90
days is forecastable where pincode-level on 90 days is noise.
### 2.2 The three data hazards (all verified in this codebase)
**H1 — Timestamps are IST wall-clock digits labelled UTC.**
`utils.DBNow` stores IST digits with a UTC tag. CLAUDE.md §8.5 is explicit:
`t.UnixMilli()` is off by 5h30m, and this already caused "yesterday's work shown
as today" on the miler app. `models/ai_runs.go:26` documents the same trap.
This matters more for prediction than anywhere else, because **hour-of-day is a
primary ETA feature and day-boundary is the demand bucket.** Fitted on raw
`createdat`, hour-of-day is right by accident and the daily bucket is wrong for
every event between 18:30 and 00:00 IST.
Use `utils.EpochMillis` (`utils/epoch.go`) as the single conversion, exactly as
the customer surface does. Write one `trip_features` SQL view that does the
correction once, and let every model read the view — never raw columns.
**H2 — There is no `deliveredat` column.**
`consignments` has `createdat`, `inwardedat`, `estimateddeliveryat`, `sladueat`,
`returninitiatedat`, `returndeliveredat` — but **no delivery completion
timestamp**. Status reaches `Delivered`; the time only exists as
`consignmenthistory.createdat WHERE eventstatus = 'Delivered'`.
So the training label is derived, not stored. Verify it is reliably written
before trusting it:
```sql
SELECT count(*) FILTER (WHERE h.consignmentid IS NULL) AS delivered_without_event
FROM consignments c
LEFT JOIN consignmenthistory h
ON h.consignmentid = c.consignmentid AND h.eventstatus = 'Delivered'
WHERE c.status = 'Delivered';
```
Non-zero means silent label loss — fix the write path before fitting anything.
**H4 — `consignments` has no `bookingid`. Every join through it is a two-path
resolution.** This is the trap CLAUDE.md §8.5 warns about, and the first draft of
this document walked straight into it.
The link is `bookingdestinations.consignmentid → bookingdestinations.bookingid`
for multi-destination pickups, falling back to `pickupbookings.consignmentid` for
console/express bookings and any row written before the fan-out existed — and
that legacy column **names only the FIRST order** of a multi-destination pickup.
`cxDestinationForConsignment` (`controllers/cxPickupFanout.go:191`) is the
canonical resolver in Go.
Get this wrong and the service-type and promise features silently attach to the
wrong parcel for every multi-destination booking. Resolve it **once**, in the
view, and never join through `consignments` directly:
```sql
CREATE OR REPLACE VIEW consignment_booking AS
SELECT c.consignmentid,
COALESCE(bd.bookingid, pb.bookingid) AS bookingid,
bd.bookingdestinationid
FROM consignments c
LEFT JOIN bookingdestinations bd ON bd.consignmentid = c.consignmentid
LEFT JOIN pickupbookings pb ON pb.consignmentid = c.consignmentid
AND bd.bookingid IS NULL;
```
Verify the resolution covers everything before trusting it:
```sql
SELECT count(*) AS unresolvable
FROM consignment_booking WHERE bookingid IS NULL;
```
**H3 — Fabricated timestamps are a known pattern in this estate.**
DailyGrubs' order data has invented delivery timestamps and a double-labelled
timezone. Doormile is a different database, but the same team and the same
`DBNow` convention. Spot-check that delivery times aren't clustered on
suspiciously round values or identical offsets from `createdat` before fitting.
**A model trained on fabricated timestamps predicts confidently and wrongly** —
and unlike a broken query, nothing surfaces the error.
### 2.3 Deliverable
A one-page readiness note: row counts, series lengths, H1/H2/H3 findings, and a
go/no-go per track. If it says no-go, stop here — Phase 1.0 (below) still pays
for itself and needs no history.
---
## 3. Phase 1 — ETA
A ladder. Each rung ships, is measured against the one below, and is only
climbed if the measurement justifies it.
### 1.0 — Measure the current promise table (**no ML, do this regardless**)
Before predicting anything, find out how wrong the constants already are:
```sql
-- Promise vs actual, per service type.
-- Joins through consignment_booking (H4) — NEVER c.bookingid, which does not exist.
SELECT so.servicetype,
count(*) AS n,
avg(EXTRACT(EPOCH FROM (h.createdat - c.createdat))/3600) AS actual_hours_avg,
avg(EXTRACT(EPOCH FROM (c.estimateddeliveryat - c.createdat))/3600) AS promised_hours_avg,
count(*) FILTER (WHERE h.createdat > c.sladueat) AS sla_breaches
FROM consignments c
JOIN consignmenthistory h ON h.consignmentid = c.consignmentid AND h.eventstatus = 'Delivered'
JOIN consignment_booking cb ON cb.consignmentid = c.consignmentid
LEFT JOIN bookingserviceoptions so ON so.bookingid = cb.bookingid
GROUP BY 1;
```
Table and column names verified against the models: `bookingserviceoptions`
(`models/booking.go:198`), `serviceabledistricts` (`models/customer_app.go:60`),
`consignmenthistory` (`models/audit.go:109`).
Two outcomes, both useful. If the constants are close, **there is no ETA problem
to solve** and the honest answer is to stop. If they are badly off, this query
gives the baseline error that every later rung must beat — and it may be fixable
by re-tuning the constants per service type and district, which is an afternoon's
work rather than a model.
### 1.1 — Routing-based ETA (**the rung most likely to be the right stopping point**)
**Better news than the first draft assumed, with a catch.** `internal/routing`
is a client for a real Valhalla-backed road-network API
(`routes.workolik.com /api/v1/optimization/doormile/sequence`), and its response
already carries durations, not just an ordering: `etaminutes`, `cumulativeeta`,
`previouskms`, `cumulativekms`, `totaleta` (`internal/routing/optimizer.go:81-95`).
And **those values are already persisted**, on `bookingassignments`
(`models/booking.go:239-244`): `step`, `previouskms`, `cumulativekms`,
`etaminutes`, `cumulativeeta`, `sequencedat`. So a predicted-vs-actual dataset —
routed ETA at assignment time paired with the delivery event — is structurally
already being collected. No new capture plumbing is needed.
**The catch, and it is a real one: those columns are almost certainly all zeros
in production.** Two independent gates:
1. **`routing.BaseURL` empty disables sequencing entirely** (deliberately —
`optimizer.go:26`). It is set from `cfg.RouteOptimizerURL` at `main.go:246`,
and `ROUTE_OPTIMIZER_URL` is one of the 19 env vars missing from
`kubernetes/manifests/doormile/miletruth.yaml` — see Track A1 of the Phase 7
plan. CLAUDE.md §8.5 says the same thing from the other side: *"Sequencing
… is not deployed yet, so riders with more than one stop come back
unsequenced (step: 0)."*
2. **`minStopsToSequence = 2`** (`optimizer.go:44`) — a rider with one stop is
never sequenced. In a courier operation a large share of assignments may be
single-stop, so even with routing switched on, routed ETAs cover only
multi-stop riders.
**Two consequences for the plan:**
- **Rung 1.1 has a hard prerequisite: Phase 7 Track A1.** Until
`ROUTE_OPTIMIZER_URL` is in the manifest, there is no routed duration to
calibrate, and `etaminutes` stays 0. Verify before building:
```sql
SELECT count(*) AS assignments,
count(*) FILTER (WHERE etaminutes > 0) AS with_routed_eta,
count(*) FILTER (WHERE step > 0) AS sequenced,
min(sequencedat), max(sequencedat)
FROM bookingassignments;
```
`with_routed_eta = 0` means this rung starts by turning routing on and waiting
for data, not by fitting a calibration.
- **Single-stop assignments need their own duration source.** Haversine ×
a learned road-circuity factor per zone is the pragmatic fallback —
`haversineKM` already exists once, in `hubController.go` (do not redefine it;
CLAUDE.md §7). Alternatively call the routing API for single stops too, which
is a change to `minStopsToSequence`'s rationale and should be decided
explicitly rather than assumed.
With a routed duration in hand, the calibration is a grouped median over history
and not machine learning:
```
eta = routed_duration × calibration[zone, hour_bucket, weekday] + handling_time[hub]
```
One SQL query, refreshed nightly. Interpretable, debuggable, no training
infrastructure, no model server — and because `etaminutes` is already stored per
assignment, the calibration is fitted on the system's own past predictions
against its own actuals, which is the cleanest possible training signal.
**Only climb past this rung if 1.1 is measurably insufficient.** For a
single-city hub-based courier, it very often isn't.
### 3a. Where the calibrated ETA can actually be applied — **decision needed**
Found while implementing, and it changes the integration plan.
**At booking-create time there is no routed duration.** The sequence is: booking
created → `estimateddeliveryat` written from the promise table → rider assigned
→ `internal/routing` sequences and writes `etaminutes`. The routed number
arrives *after* the promise has been set and shown.
So wiring `prediction.ETAMinutes` into `adminController.go:2738` or
`cxPickupFanout.go:224` would return `false` on every call, forever — not
because the calibration is cold, but because `RoutedMinutes` is structurally 0
at that moment. That is dead code, so those call sites were left alone.
The routed ETA first exists at **sequencing time**, which means applying it is
revising a promise the customer has already been given. That is a product
decision, not a wiring one:
| | |
|---|---|
| **Option A — refine `estimateddeliveryat`, never touch `sladueat`** | The customer's "arrives by" sharpens as the system learns more; the commitment they were given does not move. **Recommended.** |
| Option B — leave both, expose the calibrated ETA only on tracking | Nothing stored changes; the sharper number is display-only. Safest, least useful. |
| Option C — revise both | The SLA stops being a commitment. Not recommended. |
Option A needs one call in `internal/routing` after the ETA columns are written,
plus a decision on whether `cxstage` should emit an event when a promise moves —
a customer watching the tracking page will see the time change, and silently
is probably the wrong way for that to happen.
**This is decision 7 in §8.** Nothing should be wired until it is made.
### 1.2 — Gradient-boosted regression (only if 1.1 is insufficient)
Features, all already in the schema:
| Feature | Source |
|---|---|
| haversine + routed distance | `pickuplatitude/longitude`, `deliverylatitude/longitude` |
| hour of day, weekday | `createdat` **via `EpochMillis`** (H1) |
| pickup / delivery pincode | `pickuppincode`, `deliverypincode` |
| origin / destination hub | `originhubid`, `destinationhubid` |
| chargeable weight | `chargeableweight` |
| service type | `bookingserviceoptions.servicetype` |
| attempt count | `consignments.attemptcount` |
| rider | `assignedmileruserid` (hash, not identity — see §8) |
| hub inbound load at assignment | derived from `consignmenthistory` |
**Predict a quantile, not a mean.** An ETA shown to a customer should be the p80
— "arrives by" — not the average, which is late half the time. Use quantile
regression or a boosted model with a quantile objective. This single choice
matters more to perceived accuracy than the model family.
Label: `delivered_event.createdat − consignment.createdat`, both corrected for H1.
### 1.3 — Attempt-aware ETA
`attemptcount` exists and `MilerSkipDelivery` increments it, so failed attempts
are recorded. An ETA that ignores re-attempts is wrong for exactly the parcels
customers complain about. Worth a separate model only once 1.2 is in production
and its residuals show re-attempts as the dominant error mode.
---
## 4. Phase 2 — Demand
### 2.1 — The series
```sql
CREATE OR REPLACE VIEW demand_daily AS
SELECT (createdat)::date AS day, -- H1 correction applied in the real view
pickuppincode,
count(*) AS bookings
FROM pickupbookings
WHERE status <> 'Cancelled'
GROUP BY 1, 2;
```
Decide the grain deliberately: pincode × day is what `rebalance_riders` wants,
but it is also the sparsest. Start at **city × day**, prove the pipeline, then
descend to zone only where series length supports it.
### 2.2 — Baseline first
Seasonal naïve — "same weekday last week" — is the baseline. It is one line of
SQL and it beats badly-configured SARIMA routinely. Any model that cannot beat it
on held-out data does not ship.
### 2.3 — SARIMA or Prophet
| | Choose when |
|---|---|
| **SARIMA** (`statsmodels`) | few series, weekly seasonality, want interpretable orders |
| **Prophet** | many series, holidays matter (Indian festival calendar is a real effect on courier volume), need it to work without per-series tuning |
**Recommendation: Prophet**, for the holiday regressors. Diwali, Pongal and
regional festivals move courier volume substantially, Prophet takes a holiday
calendar as a first-class input, and it does not need per-series order selection
across dozens of pincodes.
Validate with rolling-origin backtesting (expanding window), **never** a random
split — a random train/test split on time series leaks the future and reports an
accuracy you will not see in production.
### 2.4 — What consumes the forecast
Demand prediction with no consumer is a dashboard nobody opens. The honest
consumer is `rebalance_riders` (`seed.go:295`), which is itself unimplemented —
so Phase 2 should be scoped **with** that tool's executor or not at all.
Minimum useful output: tomorrow's expected bookings per zone, plus a
staffing-gap signal against rostered riders. That is actionable; a forecast
number alone is not.
---
## 5. Where this runs
A prediction service is a **third** runtime next to the Go API and the Python
agents. Options, cheapest first:
| Option | Shape | Cost |
|---|---|---|
| **A — SQL + nightly job** | Calibration tables computed by a Go sweeper; serving is a table lookup | no new runtime, no new image |
| **B — module inside `AI_engine`** | New package; adds `pandas`/`statsmodels`/`prophet` to the image | one runtime, image grows ~300MB |
| **C — separate service** | Own repo/image/deploy/probes | full operational cost |
**Recommendation: A for Phase 1.1, B for Phase 2.** Rung 1.1 needs no model
server at all — it is a calibration table and a multiply, so it belongs in the
Go backend as a sweeper beside `StartPendingSweeper` (`main.go:244`). Prophet
genuinely needs Python, so Phase 2 lands in `AI_engine`. Option C only becomes
right if Phase 1.2 happens and model serving needs independent scaling.
**Prerequisite from the Phase 7 plan:** `AI_engine` is not in Kubernetes and has
no HTTP health surface (Track C1 there). Phase 2 inherits that work — it cannot
deploy before it.
---
## 6. LSTM — the condition for revisiting
Not recommended now. The condition under which it becomes worth measuring:
- Phase 2.3 is in production, backtested, and **losing to its own residual
structure** — i.e. Prophet's errors are autocorrelated in a way a sequence
model could capture; and
- ≥ 2 years of daily data across ≥ 50 series (≈ 36,000 observations), and
- a measured business cost to the remaining forecast error that exceeds the cost
of training infrastructure, GPU or CPU-hours, and the ongoing retraining a
neural model needs to not rot.
All three, not any one. Until then an LSTM here would be a more expensive way to
get a worse number, and the honest recommendation is to say so rather than build
it.
---
## 7. File manifest
### Phase 0 — data readiness (no application code)
| File | New? |
|---|---|
| `doormile_backend/docs/prediction-data-readiness.md` | **new** — the findings note |
| `doormile_backend/scratch/readiness_queries.sql` | **new** — the queries above, kept for re-running |
### Phase 1.0 / 1.1 — measurement and routing ETA
| File | New? | Change |
|---|---|---|
| `doormile_backend/migrations/migrate.go` | | add the `consignment_booking` view (H4), the `trip_features` view (H1 correction in one place), and the `eta_calibration` table |
| `doormile_backend/internal/prediction/calibration.go` | **new** | nightly grouped-median refresh |
| `doormile_backend/internal/prediction/eta.go` | **new** | `EstimateETA(booking) (time.Time, confidence)` |
| `doormile_backend/internal/prediction/eta_test.go` | **new** | falls back to the promise table when calibration is missing |
| `doormile_backend/internal/prediction/sweeper.go` | **new** | ticker + Redis lock, pattern from `internal/assignment/sweeper.go:88` |
| `doormile_backend/main.go:244` | | `go prediction.StartCalibrationSweeper()` |
| `doormile_backend/controllers/adminController.go:2740` | | call `prediction.EstimateETA`, keep constants as fallback |
| `doormile_backend/controllers/cxPickupFanout.go:224` | | same, keep the promise table as fallback |
| `doormile_backend/utils/epoch.go` | | read-only — the H1 conversion to reuse |
### Phase 1.2 — learned ETA (only if 1.1 insufficient)
| File | New? |
|---|---|
| `AI_engine/prediction/__init__.py` · `eta_model.py` · `features.py` | **new** |
| `AI_engine/prediction/train_eta.py` | **new** — offline training, writes a versioned artifact |
| `AI_engine/tests/test_eta_features.py` | **new** — H1 correction asserted on both timestamp taggings |
| `AI_engine/requirements.txt` | modify — `pandas`, `scikit-learn` or `lightgbm` |
| `doormile_backend/internal/prediction/eta.go` | modify — call the service, fall back to 1.1 |
### Phase 2 — demand
| File | New? |
|---|---|
| `AI_engine/prediction/demand_model.py` | **new** |
| `AI_engine/prediction/holidays_in.py` | **new** — the festival calendar |
| `AI_engine/prediction/backtest.py` | **new** — rolling-origin, never a random split |
| `AI_engine/tests/test_demand_backtest.py` | **new** — must beat seasonal-naïve to pass |
| `AI_engine/requirements.txt` | modify — `prophet` or `statsmodels` |
| `AI_engine/Dockerfile` | modify — Prophet needs a compiler toolchain |
| `doormile_backend/migrations/migrate.go` | modify — `demand_daily` view, `demand_forecast` table |
| `doormile_backend/routes/routes.go` | modify — `GET /admin/forecast/demand` (staff-only) |
| `doormile_backend/controllers/forecastController.go` | **new** |
| `doormile_backend/internal/ai/registry/seed.go:295` | modify — `rebalance_riders` once it has a consumer |
| `kubernetes/manifests/doormile/ai-engine.yaml` | modify — resources for Prophet |
### Console (Phase 2 only, optional)
| File | Change |
|---|---|
| `krow_talent_app/src/api/doormile/endpoints.js` | add `getDemandForecast` |
| a new forecast panel | render it — **after** a consumer exists, not before |
---
## 8. Decisions needed
0. **Is `ROUTE_OPTIMIZER_URL` set in the cluster?** If not, `etaminutes` is 0
everywhere and rung 1.1 begins with Phase 7 Track A1 plus a data-accumulation
wait, not with a calibration. This gates more than anything else here.
1. **Is there enough history?** Phase 0 answers it. Everything else is blocked
on that number.
1b. **Do single-stop assignments get routed too?** `minStopsToSequence = 2`
excludes them today. Either lower it, or accept a haversine-based fallback for
single-stop ETAs. An explicit call, not an assumption.
2. **Does ETA need to be learned at all, or do the constants just need
re-tuning?** Rung 1.0 answers it, cheaply.
3. **p80 or mean ETA?** Recommend p80 — "arrives by" is the promise customers
hear, and a mean is late half the time.
4. **Demand grain** — city × day to start, or straight to pincode × day?
Recommend city first.
5. **Does `rebalance_riders` get an executor in the same phase?** If no, Phase 2
produces a number nobody acts on.
6. **Rider as a feature (1.2)** — per-rider ETA adjustment is a performance
signal about a named person. Hash it, use it only in aggregate, and decide
deliberately whether it may ever surface in the console. This is a people
decision, not a modelling one.
7. **May a promise be revised after the customer has seen it?** §3a. Blocks the
last wiring step of rung 1.1 — everything else is built. Recommend Option A:
refine `estimateddeliveryat`, never move `sladueat`.
---
## 9. Sequencing
```
Phase 7 A1 (env vars) ──> routing actually runs ──> etaminutes accumulates
│ │
▼ ▼
Phase 0 (readiness) ──> [go / no-go] 1.1 calibration possible
│
no-go ─────────────┴──> stop; re-tune constants (1.0) and revisit after go-live
│
go ──> 1.0 measure ──> 1.1 routing ETA ──> [measure] ──> 1.2 only if needed
└──> 2.1/2.2 baseline ──> 2.3 Prophet (needs Phase 7 C1 first)
```
**Both tracks now depend on the Phase 7 plan**, for different reasons: rung 1.1
needs `ROUTE_OPTIMIZER_URL` (Track A1) before routed ETAs exist at all, and
Phase 2.3 needs `AI_engine` deployable with a health surface (Track C1). Track A1
is one manifest edit and unblocks both — do it first regardless of which track
you want.
**Start with Phase 0 and rung 1.0.** Both are SQL, neither needs a model, and
together they either justify the rest of this plan or retire it. Rung 1.1 is the
highest-leverage item in the document and is not machine learning at all.
The likeliest honest outcome: **1.0 + 1.1 ship, 1.2 is never needed, Phase 2
waits for volume, and LSTM never happens.** That is a success, not a shortfall.
---
## Standing constraints
- Nothing committed, pushed or deployed without being asked.
- No migrations run against a real database without being asked — the views and
tables here are additive, but "additive" is not "has run".
- Phase 0's queries are read-only; they are safe to run, and should be run before
anything else in this document.