Part I
+Doormile Parcel Flow
+How a parcel actually moves through Doormile — from a sales rep's GPS-stamped doorstep survey to a consignee reading six digits aloud. And where that pipeline, built for on-demand courier, meets a business run as milk rounds.
+The parcel's path, end to end
+Seven stages, two route branches, one convergence. Every arrow below is a real state transition in the Go backend — the labels are the endpoints that cause them and the statuses they write.
+ +pickup-complete is the pivot between them, and the only place the OTP is minted. The teal branch never touches a hub.-
+
- 01Acquisition. A rep registers a business on site; the survey coordinates are the proof they stood there. The lead converts to a Tenant with its own rate card and depot sites. +
- 02Origination. Three doors, one object.
Bookingsourceis the only thing that distinguishes them.
+ - 03Dispatch. Nearby available riders from Redis, scored by Claude with reasoning logged, degrading to greedy-nearest after five seconds. +
- 04First mile. Accept, reach, collect — each an idempotent state transition, because riders work on flaky mobile networks. +
- 05Custody handoff. Booking becomes Consignment. A six-digit code goes to the receiver and to nobody else. +
- 06Route fork. Matching 3-digit pincode prefixes skip the hub entirely. Everything else rides a tripsheet. +
- 07Close. The consignee reads six digits aloud. Nothing else marks a parcel delivered. +
Where the model and the code disagree
A milk round has five properties, and all five are anti-on-demand: a fixed beat, a rider who owns it, standing orders that recur by default, a promised slot, and density economics. Stage 03 above is built on the opposite premise.
+ +backend_jupiter has both. Two lineages, one repo.rider_substitutions — absent rider, substitute rider, per tenant, per date, with a scheduled → active → completed lifecycle. On-demand dispatch has no concept of an absent rider; you simply assign someone else. That table can only exist if a specific person owns a specific round.
| Milk-round property | In Doormile today | Where it lives |
|---|---|---|
| Promised time slot | Present | Preferredpickupfrom / Preferredpickupto on PickupBooking |
| Recurring account shape | Present | CRM captures shipping_frequency, parcel_volume, active_contracts |
| Depot / dairy origin | Present | Tenant → TenantLocation, per-site reporting on Tenantlocationid |
| Dense local drop | Present | 3-digit pincode prefix match skips the hub entirely |
| Day's order book in one drop | Present | POST /admin/expressbooking/bulk |
| Round ownership | Absent | only rider_substitutions, and only in backend_jupiter |
| Beat with ordered stops | Absent | no Route / Beat / Round model — Tripsheet is linehaul only |
| Standing orders | Absent | the only "subscription" in the backend is NATS message subscriptions |
| Waves as delivery rounds | Absent | frontend display filter in utils/batchBucket.js — see below |
Why a wave isn't a round
Morning / Afternoon / Evening look like rounds. They aren't. They bucket on orderdate — when the order was placed — so a wave describes arrival, not departure. The dispatch folder's own CLAUDE.md records that the promised-delivery field was tried here and deliberately rejected.
There is a second constraint recorded in that file worth carrying forward: the filter field and the bucket field must be the same field. An earlier attempt to admit rows on updatedat while bucketing on createdat put six orders in the Evening batch at 10:41 in the morning on a day with no orders at all.
What a native milk round would need
Four additions, roughly in dependency order. None of them fight the existing pipeline — the consignment, OTP and proof-of-delivery machinery downstream of stage 05 stays exactly as it is.
+ +-
+
- 01A
Beatentity — an ordered stop sequence owned by a tenant and a hub, with a slot window. This is the object the whole model is missing.
+ - 02Rider-to-beat assignment that persists — across days, not per parcel.
rider_substitutionsin jupiter is already the covering mechanism; it just needs a beat to point at.
+ - 03A standing-order generator — materialises tomorrow's bookings from a recurrence rule against a
TenantLocation, so the day's round exists before anyone books it.
+ - 04Bucketing moved to the slot field — at which point the original clustered windows become correct again, and a wave finally means a departure. +
Dispatch stops choosing a rider per parcel and starts doing two different jobs: sequencing a beat's stops before it goes out, and handling exceptions — the overflow parcel, the absent rider, the address that doesn't fit today's round. That is a smaller, sharper problem than scoring every candidate on every booking, and it preserves the density that makes the model pay.
+Part II
+Doormile Route Intelligence
+Each leg of the journey drawn on its own — first, mid, last — with its real states, endpoints and failure branches. Then all three as one chain. Then where an AI-optimised design would actually put its intelligence, which is not where Doormile puts it now.
+Shipper to origin hub
Everything from a booking existing to a parcel being in a rider's hands. The happy path runs down the centre; every exception branch on the right is one the backend actually implements.
+ +pickup-complete collecting COD twice.Two details worth holding onto. ALREADY_PICKED_UP and INVALID_STATE are stable machine-readable codes the rider app branches on, so the state machine is enforced server-side rather than trusted from the client. And commitAssignment writes booking, assignment and miler availability in a single transaction — there is no window where a rider is holding a job the booking does not know about.
Hub to hub
The line-haul leg. This is the only one with no automation at all — every decision below is made by a person at a screen.
+ +Dispatch timing is the central mid-mile trade: the marginal cost of running an extra vehicle against the SLA-breach risk of the parcels you hold back. Both sides are computable — you already store sladueat on every consignment. Today it is a dispatcher's instinct.
Destination hub to consignee
The only leg with real road-network routing — and the only one with a retry loop, because a delivery can fail in a way a pickup cannot.
+ +Attemptcount and the skip reason are already being recorded on every failure. That is a labelled training set for first-attempt success prediction, sitting unused.Every trip around that retry loop is a second full delivery run for one parcel. First-attempt delivery rate is the dominant cost line in Indian last-mile, and the loop above is currently entered on discovery rather than predicted in advance.
+The whole flow
Every status a parcel can hold, in the order it holds them, colour-coded by leg. Read it as a snake: left to right, drop, right to left, drop, left to right.
+ +PickupBooking, everything after Converted_To_Consignment is a Consignment. The teal bypass and the gold loop are the two places a parcel's path stops being linear.| Leg | Entity | Ends when | Loops | Optimised today by |
|---|---|---|---|---|
| First | PickupBooking | parcel in rider's hands | reject → reassign | per-parcel geo-search + LLM pick |
| Mid | Consignment | at destination hub | multi-leg re-inward | nothing — hand-built tripsheets |
| Last | Consignment | OTP verified, proof filed | skip → next attempt | Valhalla sequencing, post-assignment |
The ordering mistake
This is the most consequential thing in the codebase, and it is a sequencing question about the software rather than the parcels. Doormile assigns first and sequences second. Those are not two problems — they are one problem, and solving them in series throws away most of the available gain.
+ +Greedy nearest-neighbour assignment is a known-bad heuristic for vehicle routing — on realistic instances it lands well short of what a proper solver finds on the same data, and no amount of smarter per-parcel scoring recovers the gap, because the loss comes from committing to each parcel before seeing the rest. A better model choosing one rider at a time is still choosing one rider at a time.
+Put the intelligence where it belongs
The expert move is to stop treating "AI" as one layer. Logistics decisions live on three clocks, and each clock wants a different technique. Doormile currently has a language model on the second clock, which is the one place it fits worst.
+ +Per leg, what to build
+ +-
+
- First — batch the window, solve once. Accumulate 10–15 minutes of open pickups and solve assignment and sequence together. Make the candidate set adaptive: widen the radius until k genuine candidates are found, instead of a fixed 10 km that returns ten near-identical riders at noon and nothing at 6 AM. Learn dwell time per pickup point — every
arrivedat→pickup-completepair is already a labelled example.
+ - Mid — auto-build the tripsheet. Parcel-to-truck given capacity, destination mix and cut-off is bin-packing with a clean objective. Then price hold-vs-go against
sladueat, and consolidate thin lanes through a transfer hub rather than running half-empty direct trips —Batchkind'slocal/transfersplit already anticipates this.
+ - Last — predict first-attempt success and sequence on it.
Attemptcountand skip reasons are your training labels. Then resolve addresses with the pgvector you already run: embedding delivery addresses and snapping them to confirmedDeliveryProofcoordinates turns your delivered history into a private geocoder that beats any general one inside your own zones.
+
calculateETA is (distance / 20) × 60 + 10 — a flat 20 km/h against straight-line distance, plus a fixed ten-minute buffer. It sets rider expectations, customer-facing ETAs and support load. Road time from the Valhalla matrix you already pay for, multiplied by a time-of-day factor and added to learned dwell, is a same-day change with effect on every leg.
Build order
Sequenced by ratio of effect to effort, not by ambition. The first two need no new infrastructure at all.
+ +| # | Move | Leg | Depends on | Why it's here |
|---|---|---|---|---|
| 1 | Score AgentDecision against outcomes | all | nothing — data exists | You cannot currently tell whether AI dispatch beats greedy-nearest. Until you can, every other change is unmeasurable. |
| 2 | Real ETA from road time + learned dwell | all | Valhalla matrix | Replaces a flat 20 km/h constant. Touches rider trust, customer ETA and support volume at once. |
| 3 | Adaptive candidate set | first | nothing | Removes both failure modes of the fixed 10 km / top 10 window. |
| 4 | Batched joint assign + sequence | first, last | solver, batch window | The ordering fix in §05. Largest routing gain available. |
| 5 | Address resolution on pgvector | last | delivered history | Turns your own proof-of-delivery coordinates into a private geocoder for your zones. |
| 6 | First-attempt success model | last | #5, Attemptcount history | Attacks the dominant cost line in last mile. |
| 7 | Automated tripsheet load planning | mid | #2, volume forecast | The only leg with no optimisation at all today. |
| 8 | Hold-vs-go and dynamic cut-offs | mid | #7, sladueat | Converts a dispatcher's guess into a priced decision. |
| 9 | Beat districting | first, last | #4, a Beat entity | The strategic layer. Only worth it once daily routing is solid. |
Every item above gets cheaper under a round model rather than more expensive. When stops are stable, the heavy optimisation moves to beat-design time — monthly, offline, with as much compute as you like — and the daily solve shrinks to the deltas: today's new stops, today's skips, today's absences. Real-time combinatorial search over the whole city, every wave, is a cost you only pay because the beats do not exist yet.
+Part III
+Doormile Agent Layer
+Where generative AI genuinely earns its place in the platform — auto-assignment included — how it plugs into the Go backend you already have, and the trust ladder that gets it from logging quietly to acting alone without a bad week in between.
+The rule that makes all of this work
The model proposes. The solver decides. The system validates.
+ +Every capability below is an application of that one sentence. A language model is exceptional at turning mess into structure — a dispatcher's sentence into a constraint, a WhatsApp forward into a booking, a photo into a condition code, a solver's output into an explanation a rider will actually read. It is poor and expensive at choosing the best of ten thousand arrangements, which is exactly what dispatch is.
+ +So "auto-assign with GenAI" does not mean a model picking riders. It means a model doing the four things around the pick that were previously impossible to automate, while a solver does the pick itself in milliseconds and a validator refuses anything that breaks an invariant. That is a more capable system than an LLM choosing alone, and a cheaper one.
+ +The payload your backend sends to routemate is entirely structured — miler_id, distance_km, rating, on_time_rate_30d, hub_load, hub_capacity, plus hour, peak flag and zone. No customer-supplied free text goes into it. That closed payload is a security property, and §06 explains why keeping it is the single most important rule as you expand.
Where GenAI actually earns its place
Ten capabilities, ranked by what they return for what they cost. Note how few of them are "the AI decides" and how many are "the AI turns something unstructured into something the system can already handle."
+ +| # | Capability | Leg | What the model does | Value | Effort |
|---|---|---|---|---|---|
| 1 | Address resolution | last | Embeds messy Indian address text, clusters duplicates, snaps to confirmed DeliveryProof coordinates. Builds you a private geocoder from your own delivered history. | highest | low |
| 2 | Decision explanation | all | Turns solver output into a sentence for the dispatcher and a line for the rider. Drives adoption more than accuracy does. | high | low |
| 3 | Unstructured intake | first | A client's WhatsApp message, email or PDF manifest becomes structured bookings. You support Excel; clients don't always send Excel. | high | medium |
| 4 | NL constraint compiler | all | "Don't give Ravi the mall runs" becomes a typed solver constraint a human approves once and the solver honours forever. | high | medium |
| 5 | Exception agent | all | Rider absent, vehicle down, forty parcels stranded. Tool-calling over a bounded action space, proposing a recovery plan. | high | medium |
| 6 | Damage detection at inbound | mid | Vision on the inbound scan photo auto-populates Condition, which is a manual dropdown today and therefore mostly says Good. | medium | low |
| 7 | POD verification | last | Vision checks the delivery photo actually shows a parcel at a door, and that the signature isn't blank. Catches soft fraud. | medium | low |
| 8 | Vernacular voice | first, last | Rider status by speech in Tamil, Telugu, Kannada or Hindi; an outbound call confirming the consignee will be home. Directly attacks first-attempt failure. | high | high |
| 9 | Operator copilot | all | Extends the existing console assistant from read-only queries to proposing actions with confirmation. | medium | medium |
| 10 | Scenario simulation | all | Replays historical days against a proposed policy before it touches production. Your safety net for everything above. | medium | medium |
Four hubs across Coimbatore, Hyderabad, Bengaluru and Chennai means four language regions. Item 8 is not a nicety in that footprint — a rider who can speak a status update is a rider who actually files one.
+Auto-assign, properly
This is the thing you asked about, drawn end to end. The deterministic spine runs down the left and never waits on a model. The three GenAI touch points sit beside it, feeding in out of band.
+ +Why the constraint compiler is the sleeper
+Dispatchers hold dozens of rules in their heads that never reach the system: which rider is slow at a particular gate, which client insists on the same face, which lane floods in monsoon. Today the only way to encode those is a developer ticket, so they stay in the dispatcher's head and the optimiser looks stupid. A model that turns a typed sentence into a reviewable constraint — which a human approves before it ever binds — converts tacit knowledge into solver input at conversational speed. That is a capability the solver cannot have on its own, and it is where generative AI is doing something genuinely irreplaceable.
+ +// what the compiler emits — reviewed by a human, then stored
+{
+ "type": "exclude_miler_from_location_class",
+ "miler_id": 412,
+ "location_class": "mall_service_gate",
+ "reason": "slow at gate handover",
+ "author": "dispatcher:coimbatore:7",
+ "expires_at": null
+}
+Wiring it into what you have
You do not need new infrastructure. Every piece below already exists in the estate; the work is contracts and placement.
+ +| Piece | Already there | What changes |
|---|---|---|
| Async invocation | NATS JetStream, the booking worker, a 60 s budget | Model calls become subscribers, not inline HTTP. Nothing in a request path waits on one. |
| Fallback discipline | AI_LAYER_FALLBACK and the 5 s cap in selectMilerWithAI | Keep the pattern exactly; apply it to every new capability. Best-effort or it does not ship. |
| Audit | AgentDecision with a DecisionType discriminator | One row per model call, whatever the capability. The discriminator was clearly built for this. |
| Idempotency | Redis-backed Idempotency-Key on payment, pickup and deliver | Extend to every agent-initiated write. An agent that retries must not double-act. |
| Vector store | pgvector, in the database already | Address embeddings land here. No new dependency. |
| Road distance | Valhalla behind routes.workolik.com | Becomes the solver's cost matrix rather than a post-hoc sequencer. |
| Object storage | presigned uploads for signatures and photos | The same URLs feed the vision capabilities. Images never transit your API. |
| Console surface | the existing deterministic assistant and intents | Gains an action tier behind a confirmation step. |
One deliberate constraint: the agent layer never writes to the database directly. It calls the same internal endpoints a human operator would, so tenant scoping, idempotency, validation and audit apply identically whether the caller is a person or a model. If an agent can reach a table your dispatcher cannot, you have built two systems and will secure only one.
+The trust ladder
Every capability climbs the same four rungs, and each rung has a numeric gate. Nothing goes to the next rung on a good feeling.
+ +The loop you already built
This is the sharpest finding in the codebase, and it makes everything above cheaper.
+ +AgentDecision.Outcome, OutcomeRecordedAt, the controller and the route are all in place. Nothing writes to them. This is the cheapest high-value change available in the whole system.Call UpdateDecisionOutcome from the deliver, skip and RTO handlers with a simple verdict — delivered first attempt, delivered late, failed, reassigned. From that day forward every dispatch decision is scored, you can replay history against any new policy offline, and every claim anyone makes about the AI dispatcher becomes checkable. Until then, no gate in §04 can be measured at all.
Guardrails
-
+
- Closed-set selection, enforced server-side. A model returning
chosen_miler_idmust have that ID checked against the candidate set before commit. Models hallucinate identifiers; a validator costs three lines and removes the entire failure class.
+ - Keep the payload structured. Your current request to
routematecarries no customer free text. The moment you pass rawNotes, address strings or client emails into a prompt that also drives actions, you have built a prompt-injection channel — anyone who can type into a booking form can try to instruct your dispatcher. If unstructured text must be read, do it in a separate call that only extracts, and never let that call's output reach an action without validation.
+ - Tenant isolation is an invariant, not a preference.
HubStaffAccount.Tenantidalready isolates partner freight. Agents must be scoped identically, and cross-tenant leakage should be a hard validator failure, not a scoring penalty. This is the one mistake that is a breach rather than a bad route.
+ - Every agent action is idempotent. You already have the Redis-backed key infrastructure. Agents retry more than humans do; a duplicated reassignment or a double COD entry is far worse than a slow one. +
- Budget per booking, enforced. Set a token ceiling per decision and per day, with a hard cutoff to the deterministic path. Cost overruns in agentic systems come from retry storms, not from steady state. +
- Mind what leaves the estate. Consignee phone numbers, COD amounts and precise home coordinates are in scope for whatever model endpoint you call. Decide deliberately what is sent, and prefer IDs over identities wherever the model does not need the human detail. +
- Escalate rather than improvise. Give the exception agent an explicit "page a human" tool and reward using it. An agent with no honourable way to stop will invent an action instead. +
Ninety days
| Phase | Ship | Rung | Proves |
|---|---|---|---|
| Wk 1–2 | Wire UpdateDecisionOutcome from deliver / skip / RTO. Build the replay harness over AgentDecision. | — | Every later claim becomes measurable. |
| Wk 3–4 | Real ETA: Valhalla road time × time-of-day factor + learned dwell, replacing calculateETA. | direct | No model needed; immediate accuracy win on all three legs. |
| Wk 3–6 | Address resolution on pgvector. Embed delivered addresses, cluster, snap to DeliveryProof coordinates. | shadow → advisory | Highest-leverage capability, lowest risk. |
| Wk 5–8 | VRP solver behind a flag: batched joint assign + sequence, running in shadow beside today's path. | shadow | Head-to-head on the replay corpus against live decisions. |
| Wk 7–9 | Explainer on solver output; dispatcher and rider-facing, in-language. | advisory | Adoption. Riders follow routes they understand. |
| Wk 9–11 | NL constraint compiler with a human approval queue. | advisory | Tacit dispatcher knowledge starts reaching the solver. |
| Wk 10–12 | Exception agent with a bounded tool set, proposals only. | advisory | The escalation queue starts shrinking. |
| Wk 12 | Promote whichever capabilities have cleared their gates. Not the ones that haven't. | → auto with veto | The ladder does the deciding, not the roadmap. |
Vision on inbound condition and POD, vernacular voice, and the operator copilot's action tier are all worth building — and all of them are easier once the outcome loop, the replay harness and the validator exist. Building those three foundations first is what makes the rest take weeks instead of quarters.
+