Architecture · operations · integration plan

Doormile Platform Dossier

The whole review in one place: how a parcel actually moves through Doormile, each leg drawn on its own, and where generative AI belongs in the platform. Nineteen sections, twelve diagrams, read against the nearle-de source.

Part I

Doormile Parcel Flow

How a parcel actually moves through Doormile — from a sales rep's GPS-stamped doorstep survey to a consignee reading six digits aloud. And where that pipeline, built for on-demand courier, meets a business run as milk rounds.

01

The parcel's path, end to end

Seven stages, two route branches, one convergence. Every arrow below is a real state transition in the Go backend — the labels are the endpoints that cause them and the statuses they write.

01 · ACQUISITION Field rep on site surveylat / surveylong DoormileClient POST /crm/clients converts Tenant billable account TenantPricing + TenantLocation sites 02 · ORIGINATION Customer App CRM Console Express Console bulk Excel upload PickupBooking status: Pending_Pickup · priced vs TenantPricing Preferredpickupfrom / to the time-window field — unused as a round 03 · DISPATCH Redis GEOSEARCH milers:locations · 10 km candidates Claude Sonnet scoring distance · load · rating 30-day on-time record pgvector RAG memory AgentDecision reasoning logged BookingAssignment Miler_Assigned 5 s timeout → greedy nearest miler 04 · FIRST MILE Accept /assignments/:id/accept Reached doorstep arrivedat + arrival GPS Pickup complete COD collected · Idempotency-Key 05 · CUSTODY HANDOFF Consignment created Collected_By_Miler · tracking number issued 6-digit delivery OTP to the receiver only — stripped from rider & console APIs 06 · ROUTE FORK first 3 pincode digits match? YES · hyperlocal NO · hub & spoke /consignments/:id/start-delivery Out_for_Delivery inbound scan → Inwarded_at_Hub tripsheet loaded → In_Transit destination hub arrive → assign-miler OTP verified → DeliveryProof stored signature · photo · geolocation → Delivered
Figure I.1 — the lifecycleThe booking is a promise to collect; the consignment is a parcel in custody. pickup-complete is the pivot between them, and the only place the OTP is minted. The teal branch never touches a hub.
  1. 01Acquisition. A rep registers a business on site; the survey coordinates are the proof they stood there. The lead converts to a Tenant with its own rate card and depot sites.
  2. 02Origination. Three doors, one object. Bookingsource is the only thing that distinguishes them.
  3. 03Dispatch. Nearby available riders from Redis, scored by Claude with reasoning logged, degrading to greedy-nearest after five seconds.
  4. 04First mile. Accept, reach, collect — each an idempotent state transition, because riders work on flaky mobile networks.
  5. 05Custody handoff. Booking becomes Consignment. A six-digit code goes to the receiver and to nobody else.
  6. 06Route fork. Matching 3-digit pincode prefixes skip the hub entirely. Everything else rides a tripsheet.
  7. 07Close. The consignee reads six digits aloud. Nothing else marks a parcel delivered.
02

Where the model and the code disagree

A milk round has five properties, and all five are anti-on-demand: a fixed beat, a rider who owns it, standing orders that recur by default, a promised slot, and density economics. Stage 03 above is built on the opposite premise.

TODAY · ASSIGNED JOB Each parcel scored on its own. The rider set is re-drawn from scratch every single time. P1 P2 P3 P4 GEOSEARCH + scoring R1 R2 R3 R4 No Route, Beat or Round entity exists in doormile_backend — only Tripsheet, which is a hub-to-hub manifest, not a delivery beat. stops per rider: unstable, day to day MILK ROUND · OWNED BEAT One rider, one ordered sequence, unchanged tomorrow. Familiarity is the cost saving. Rider A owns beat 7 1 2 3 4 repeats tomorrow rider_substitutions absent rider → sub rider, per tenant, per date jupiter stops per rider: stable, learned, dense
Figure I.2 — the seamThe right-hand pattern only needs a substitution table because a round has an owner. Doormile has no owner and no round; backend_jupiter has both. Two lineages, one repo.
The tell

rider_substitutions — absent rider, substitute rider, per tenant, per date, with a scheduled → active → completed lifecycle. On-demand dispatch has no concept of an absent rider; you simply assign someone else. That table can only exist if a specific person owns a specific round.

Milk-round propertyIn Doormile todayWhere it lives
Promised time slotPresentPreferredpickupfrom / Preferredpickupto on PickupBooking
Recurring account shapePresentCRM captures shipping_frequency, parcel_volume, active_contracts
Depot / dairy originPresentTenant → TenantLocation, per-site reporting on Tenantlocationid
Dense local dropPresent3-digit pincode prefix match skips the hub entirely
Day's order book in one dropPresentPOST /admin/expressbooking/bulk
Round ownershipAbsentonly rider_substitutions, and only in backend_jupiter
Beat with ordered stopsAbsentno Route / Beat / Round model — Tripsheet is linehaul only
Standing ordersAbsentthe only "subscription" in the backend is NATS message subscriptions
Waves as delivery roundsAbsentfrontend display filter in utils/batchBucket.js — see below
03

Why a wave isn't a round

Morning / Afternoon / Evening look like rounds. They aren't. They bucket on orderdate — when the order was placed — so a wave describes arrival, not departure. The dispatch folder's own CLAUDE.md records that the promised-delivery field was tried here and deliberately rejected.

Morning · 0–9 Afternoon · 9–16 Evening · 16–24 12 AM 9 AM 4 PM 12 AM bucket on orderdate current placed 12:22 → Afternoon bucket on expecteddeliverytime jupiter · rejected here promised 16:30 → Evening
Figure I.3 — the same parcel, two wavesOnly the lower one describes when the parcel goes out. Bucketing on arrival is correct for a courier queue and wrong for a round — which is why the windows inherited from jupiter had to be widened to cover all 24 hours once the field changed.

There is a second constraint recorded in that file worth carrying forward: the filter field and the bucket field must be the same field. An earlier attempt to admit rows on updatedat while bucketing on createdat put six orders in the Evening batch at 10:41 in the morning on a day with no orders at all.

04

What a native milk round would need

Four additions, roughly in dependency order. None of them fight the existing pipeline — the consignment, OTP and proof-of-delivery machinery downstream of stage 05 stays exactly as it is.

  1. 01A Beat entity — an ordered stop sequence owned by a tenant and a hub, with a slot window. This is the object the whole model is missing.
  2. 02Rider-to-beat assignment that persists — across days, not per parcel. rider_substitutions in jupiter is already the covering mechanism; it just needs a beat to point at.
  3. 03A standing-order generator — materialises tomorrow's bookings from a recurrence rule against a TenantLocation, so the day's round exists before anyone books it.
  4. 04Bucketing moved to the slot field — at which point the original clustered windows become correct again, and a wave finally means a departure.
What the AI layer becomes

Dispatch stops choosing a rider per parcel and starts doing two different jobs: sequencing a beat's stops before it goes out, and handling exceptions — the overflow parcel, the absent rider, the address that doesn't fit today's round. That is a smaller, sharper problem than scoring every candidate on every booking, and it preserves the density that makes the model pay.

Part II

Doormile Route Intelligence

Each leg of the journey drawn on its own — first, mid, last — with its real states, endpoints and failure branches. Then all three as one chain. Then where an AI-optimised design would actually put its intelligence, which is not where Doormile puts it now.

01First mile

Shipper to origin hub

Everything from a booking existing to a parcel being in a rider's hands. The happy path runs down the centre; every exception branch on the right is one the backend actually implements.

HAPPY PATH EXCEPTIONS Booking created Created | Pending_Pickup NATS queue → TryAssignOnce 60 s worker budget GEOSEARCH — any candidates? geoRadiusKm 10 · geoMaxCount 10 none in radius Escalated → manual assignment /admin/bookings/:id/assign-miler AI layer scores candidates routemate.workolik.com · 5 s cap timeout Greedy nearest miler deterministic fallback, never blocks commitAssignment — one transaction booking → Miler_Assigned assignment → Assigned · miler → Assigned AgentDecision row written reasoning + candidate count, for audit push notification to rider rider accepts? reject Rejected → Reassigned back into the search Accept Pickup_Scheduled · miler On_Pickup Reached doorstep Arrived_At_Pickup · arrivedat + GPS Pickup complete Picked_Up → Converted_To_Consignment COD collected · Idempotency-Key Consignment issued Collected_By_Miler · trackingno + 6-digit OTP to receiver
Figure II.1 — first mileEvery branch here is real: escalation on an empty radius, the five-second greedy fallback, rejection looping back into the search, and the idempotency key that stops a retried pickup-complete collecting COD twice.

Two details worth holding onto. ALREADY_PICKED_UP and INVALID_STATE are stable machine-readable codes the rider app branches on, so the state machine is enforced server-side rather than trusted from the client. And commitAssignment writes booking, assignment and miler availability in a single transaction — there is no window where a rider is holding a job the booking does not know about.

02Mid mile

Hub to hub

The line-haul leg. This is the only one with no automation at all — every decision below is made by a person at a screen.

LINE-HAUL PATH BRANCHES Parcel arrives at origin hub Inbound barcode scan shelf assigned · Condition recorded → Inwarded_at_Hub Condition ≠ Good Damaged — held at hub exception queue, not loaded destination hub = this hub? currenthubid vs destinationhubid yes mid mile skipped straight to last-mile assignment no Tripsheet created by hand status Draft · source + destination hub Batchkind: local | transfer Items scanned onto the sheet item Pending → Loaded consignment → Tripsheet_Loaded mismatch Discrepancy scanned but not on the manifest, or missing Vehicle + driver attached status → Ready Dispatch Dispatched · dispatchtime · → In_Transit hold-vs-go decided by guess depart at 80% fill, or wait for 95%? Arrival scan & offload Arrived · arrivaltime · items Unloaded currenthubid updated final destination reached? yes → last mile no — re-inward, next leg
Figure II.2 — mid mileNote the loop on the left: a parcel can ride more than one tripsheet. Nothing decides which parcels board which truck, when it should leave, or whether the lane should exist — those are the three open decisions on this leg.
The unpriced decision

Dispatch timing is the central mid-mile trade: the marginal cost of running an extra vehicle against the SLA-breach risk of the parcels you hold back. Both sides are computable — you already store sladueat on every consignment. Today it is a dispatcher's instinct.

03Last mile

Destination hub to consignee

The only leg with real road-network routing — and the only one with a retry loop, because a delivery can fail in a way a pickup cannot.

DELIVERY PATH FAILURE PATH manual assign-miler sync AI auto-assign greedy batch-assign Rider assigned assignment Assigned · miler On_Delivery SequenceMilerStops → Valhalla writes step · previouskms · cumulativeeta async, best-effort, needs ≥ 2 stops unconfigured → stops stay unsequenced assignment never fails because of this Start delivery → Out_for_Delivery rider reaches the consignee receiver present, OTP correct? OTP_REQUIRED | OTP_INVALID no Skip — failed attempt Attemptcount + 1 · reason logged re-queue for next attempt yes Deliver OTP verified · Idempotency-Key signature + photo via presigned URL DeliveryProof row written attempts exhausted RTO_Initiated Returnreason · Parentconsignmentid Returned_to_Sender the return leg is itself a full journey Delivered Codcollected reconciled Billingstatus → Billed other terminal states Missing · Damaged · Cancelled
Figure II.3 — last mileAttemptcount and the skip reason are already being recorded on every failure. That is a labelled training set for first-attempt success prediction, sitting unused.
Where the money leaks

Every trip around that retry loop is a second full delivery run for one parcel. First-attempt delivery rate is the dominant cost line in Indian last-mile, and the loop above is currently entered on discovery rather than predicted in advance.

04All three

The whole flow

Every status a parcel can hold, in the order it holds them, colour-coded by leg. Read it as a snake: left to right, drop, right to left, drop, left to right.

FIRST MILE · booking Created | Pending_Pickup Miler_Assigned Pickup_Scheduled Arrived_At_Pickup Picked_Up Converted_To_ Consignment MID MILE · consignment Collected_By_ Miler Inwarded_at_Hub Tripsheet_Loaded In_Transit next hub · multi-leg hyperlocal — pincode prefix match LAST MILE Out_for_Delivery Delivered attempts exhausted RTO_Initiated Returned_to_ Sender skip → Attemptcount + 1 EXCEPTION TERMINALS · reachable from any leg Missing Damaged Cancelled
Figure II.4 — the complete chainTwo entities, one chain: the first row is a PickupBooking, everything after Converted_To_Consignment is a Consignment. The teal bypass and the gold loop are the two places a parcel's path stops being linear.
LegEntityEnds whenLoopsOptimised today by
FirstPickupBookingparcel in rider's handsreject → reassignper-parcel geo-search + LLM pick
MidConsignmentat destination hubmulti-leg re-inwardnothing — hand-built tripsheets
LastConsignmentOTP verified, proof filedskip → next attemptValhalla sequencing, post-assignment
05

The ordering mistake

This is the most consequential thing in the codebase, and it is a sequencing question about the software rather than the parcels. Doormile assigns first and sequences second. Those are not two problems — they are one problem, and solving them in series throws away most of the available gain.

TODAY · ASSIGN, THEN SEQUENCE one parcel as it arrives GEOSEARCH 10 km · top 10 LLM picks one 5 s cap assignment committed Valhalla sequences that rider's stops, async can only reorder a set it did not get to choose PROPOSED · ONE JOINT SOLVE N parcels · batch window M riders · live capacity constraints slots · capacity · skills VRP solver assignment and order decided together OR-Tools / VROOM · ms complete routes rider · step · ETA LLM exceptions only overflow · absence · bad address · breakdown
Figure II.5 — the fix is an ordering changeThe Valhalla client already exists and already speaks Doormile's vocabulary. Moving it in front of assignment, and feeding it the whole open set rather than one rider's inherited stops, is a smaller change than it sounds.
Why this matters more than model quality

Greedy nearest-neighbour assignment is a known-bad heuristic for vehicle routing — on realistic instances it lands well short of what a proper solver finds on the same data, and no amount of smarter per-parcel scoring recovers the gap, because the loss comes from committing to each parcel before seeing the rest. A better model choosing one rider at a time is still choosing one rider at a time.

06

Put the intelligence where it belongs

The expert move is to stop treating "AI" as one layer. Logistics decisions live on three clocks, and each clock wants a different technique. Doormile currently has a language model on the second clock, which is the one place it fits worst.

HORIZON DECIDES TECHNIQUE CADENCE Strategic the shape of the network beat & territory design · hub siting which lanes exist at all rider headcount per zone districting · facility location demand forecasting monthly Tactical the plan for this wave ← the LLM sits here today joint assignment + sequencing tripsheet load planning hold-vs-go on every departure rider-to-beat for the day MILP / VRP solver OR-Tools, VROOM, Valhalla matrix deterministic · explainable · milliseconds per wave Operational what just went wrong unresolvable address · receiver unreachable rider absent · vehicle down · overflow parcel operator asks a question in plain English LLM + rules judgment over messy, unstructured input per second Completed-route exhaust AgentDecision · breadcrumbs · arrivedat + GPS · Attemptcount · DeliveryProof geolocation → learned dwell, travel time, failure risk trains every layer
Figure II.6 — three clocks, three techniquesLanguage models are strong on unstructured judgment and weak on combinatorial search. Solvers are the reverse. The current design has them swapped.

Per leg, what to build

  • First — batch the window, solve once. Accumulate 10–15 minutes of open pickups and solve assignment and sequence together. Make the candidate set adaptive: widen the radius until k genuine candidates are found, instead of a fixed 10 km that returns ten near-identical riders at noon and nothing at 6 AM. Learn dwell time per pickup point — every arrivedat → pickup-complete pair is already a labelled example.
  • Mid — auto-build the tripsheet. Parcel-to-truck given capacity, destination mix and cut-off is bin-packing with a clean objective. Then price hold-vs-go against sladueat, and consolidate thin lanes through a transfer hub rather than running half-empty direct trips — Batchkind's local/transfer split already anticipates this.
  • Last — predict first-attempt success and sequence on it. Attemptcount and skip reasons are your training labels. Then resolve addresses with the pgvector you already run: embedding delivery addresses and snapping them to confirmed DeliveryProof coordinates turns your delivered history into a private geocoder that beats any general one inside your own zones.
Replace this function first

calculateETA is (distance / 20) × 60 + 10 — a flat 20 km/h against straight-line distance, plus a fixed ten-minute buffer. It sets rider expectations, customer-facing ETAs and support load. Road time from the Valhalla matrix you already pay for, multiplied by a time-of-day factor and added to learned dwell, is a same-day change with effect on every leg.

07

Build order

Sequenced by ratio of effect to effort, not by ambition. The first two need no new infrastructure at all.

#MoveLegDepends onWhy it's here
1Score AgentDecision against outcomesallnothing — data existsYou cannot currently tell whether AI dispatch beats greedy-nearest. Until you can, every other change is unmeasurable.
2Real ETA from road time + learned dwellallValhalla matrixReplaces a flat 20 km/h constant. Touches rider trust, customer ETA and support volume at once.
3Adaptive candidate setfirstnothingRemoves both failure modes of the fixed 10 km / top 10 window.
4Batched joint assign + sequencefirst, lastsolver, batch windowThe ordering fix in §05. Largest routing gain available.
5Address resolution on pgvectorlastdelivered historyTurns your own proof-of-delivery coordinates into a private geocoder for your zones.
6First-attempt success modellast#5, Attemptcount historyAttacks the dominant cost line in last mile.
7Automated tripsheet load planningmid#2, volume forecastThe only leg with no optimisation at all today.
8Hold-vs-go and dynamic cut-offsmid#7, sladueatConverts a dispatcher's guess into a priced decision.
9Beat districtingfirst, last#4, a Beat entityThe strategic layer. Only worth it once daily routing is solid.
The milk-round dividend

Every item above gets cheaper under a round model rather than more expensive. When stops are stable, the heavy optimisation moves to beat-design time — monthly, offline, with as much compute as you like — and the daily solve shrinks to the deltas: today's new stops, today's skips, today's absences. Real-time combinatorial search over the whole city, every wave, is a cost you only pay because the beats do not exist yet.

Part III

Doormile Agent Layer

Where generative AI genuinely earns its place in the platform — auto-assignment included — how it plugs into the Go backend you already have, and the trust ladder that gets it from logging quietly to acting alone without a bad week in between.

00

The rule that makes all of this work

The model proposes. The solver decides. The system validates.

Every capability below is an application of that one sentence. A language model is exceptional at turning mess into structure — a dispatcher's sentence into a constraint, a WhatsApp forward into a booking, a photo into a condition code, a solver's output into an explanation a rider will actually read. It is poor and expensive at choosing the best of ten thousand arrangements, which is exactly what dispatch is.

So "auto-assign with GenAI" does not mean a model picking riders. It means a model doing the four things around the pick that were previously impossible to automate, while a solver does the pick itself in milliseconds and a validator refuses anything that breaks an invariant. That is a more capable system than an LLM choosing alone, and a cheaper one.

You already got the hard part right

The payload your backend sends to routemate is entirely structured — miler_id, distance_km, rating, on_time_rate_30d, hub_load, hub_capacity, plus hour, peak flag and zone. No customer-supplied free text goes into it. That closed payload is a security property, and §06 explains why keeping it is the single most important rule as you expand.

01

Where GenAI actually earns its place

Ten capabilities, ranked by what they return for what they cost. Note how few of them are "the AI decides" and how many are "the AI turns something unstructured into something the system can already handle."

#CapabilityLegWhat the model doesValueEffort
1Address resolutionlastEmbeds messy Indian address text, clusters duplicates, snaps to confirmed DeliveryProof coordinates. Builds you a private geocoder from your own delivered history.highestlow
2Decision explanationallTurns solver output into a sentence for the dispatcher and a line for the rider. Drives adoption more than accuracy does.highlow
3Unstructured intakefirstA client's WhatsApp message, email or PDF manifest becomes structured bookings. You support Excel; clients don't always send Excel.highmedium
4NL constraint compilerall"Don't give Ravi the mall runs" becomes a typed solver constraint a human approves once and the solver honours forever.highmedium
5Exception agentallRider absent, vehicle down, forty parcels stranded. Tool-calling over a bounded action space, proposing a recovery plan.highmedium
6Damage detection at inboundmidVision on the inbound scan photo auto-populates Condition, which is a manual dropdown today and therefore mostly says Good.mediumlow
7POD verificationlastVision checks the delivery photo actually shows a parcel at a door, and that the signature isn't blank. Catches soft fraud.mediumlow
8Vernacular voicefirst, lastRider status by speech in Tamil, Telugu, Kannada or Hindi; an outbound call confirming the consignee will be home. Directly attacks first-attempt failure.highhigh
9Operator copilotallExtends the existing console assistant from read-only queries to proposing actions with confirmation.mediummedium
10Scenario simulationallReplays historical days against a proposed policy before it touches production. Your safety net for everything above.mediummedium

Four hubs across Coimbatore, Hyderabad, Bengaluru and Chennai means four language regions. Item 8 is not a nicety in that footprint — a rider who can speak a status update is a rider who actually files one.

02

Auto-assign, properly

This is the thing you asked about, drawn end to end. The deterministic spine runs down the left and never waits on a model. The three GenAI touch points sit beside it, feeding in out of band.

DETERMINISTIC SPINE · never waits on a model GENERATIVE · out of band booking enqueued · NATS candidate generation Redis GEOSEARCH · adaptive k widen radius until k found constraint set assembled capacity · slots · skills · tenant scope + approved rules from the compiler VRP solver assign + sequence, jointly Valhalla matrix · milliseconds VALIDATOR — hard gate miler_id ∈ candidate set? tenant scope intact? capacity respected? reject → greedy nearest, always commitAssignment one transaction · Idempotency-Key booking + assignment + miler status AgentDecision written Context · Decision · Reasoning Outcome ← filled later, see §05 ① NL constraint compiler dispatcher types: "Ravi shouldn't take mall runs" → typed constraint → human approves once offline · stored · reused every solve ② Exception agent invoked only when the solve is infeasible or an operator escalates tools: reassign · split · defer · widen · page human proposals go back through the validator infeasible proposal — never a direct write ③ Explainer solver output → a sentence for the dispatcher, a line for the rider, in their language async · after commit · failure is cosmetic critical path budget: candidate generation + solve + validate + commit, with no network call to a model in it
Figure III.1 — auto-assign with generative AI in its right placesCompare with today: a model on the critical path with a 5-second cap, choosing one rider for one parcel. Here the model never blocks a booking, and the thing it produces — a constraint, a recovery plan, an explanation — is something no solver could have produced instead.

Why the constraint compiler is the sleeper

Dispatchers hold dozens of rules in their heads that never reach the system: which rider is slow at a particular gate, which client insists on the same face, which lane floods in monsoon. Today the only way to encode those is a developer ticket, so they stay in the dispatcher's head and the optimiser looks stupid. A model that turns a typed sentence into a reviewable constraint — which a human approves before it ever binds — converts tacit knowledge into solver input at conversational speed. That is a capability the solver cannot have on its own, and it is where generative AI is doing something genuinely irreplaceable.

// what the compiler emits — reviewed by a human, then stored
{
  "type": "exclude_miler_from_location_class",
  "miler_id": 412,
  "location_class": "mall_service_gate",
  "reason": "slow at gate handover",
  "author": "dispatcher:coimbatore:7",
  "expires_at": null
}
03

Wiring it into what you have

You do not need new infrastructure. Every piece below already exists in the estate; the work is contracts and placement.

PieceAlready thereWhat changes
Async invocationNATS JetStream, the booking worker, a 60 s budgetModel calls become subscribers, not inline HTTP. Nothing in a request path waits on one.
Fallback disciplineAI_LAYER_FALLBACK and the 5 s cap in selectMilerWithAIKeep the pattern exactly; apply it to every new capability. Best-effort or it does not ship.
AuditAgentDecision with a DecisionType discriminatorOne row per model call, whatever the capability. The discriminator was clearly built for this.
IdempotencyRedis-backed Idempotency-Key on payment, pickup and deliverExtend to every agent-initiated write. An agent that retries must not double-act.
Vector storepgvector, in the database alreadyAddress embeddings land here. No new dependency.
Road distanceValhalla behind routes.workolik.comBecomes the solver's cost matrix rather than a post-hoc sequencer.
Object storagepresigned uploads for signatures and photosThe same URLs feed the vision capabilities. Images never transit your API.
Console surfacethe existing deterministic assistant and intentsGains an action tier behind a confirmation step.

One deliberate constraint: the agent layer never writes to the database directly. It calls the same internal endpoints a human operator would, so tenant scoping, idempotency, validation and audit apply identically whether the caller is a person or a model. If an agent can reach a table your dispatcher cannot, you have built two systems and will secure only one.

04

The trust ladder

Every capability climbs the same four rungs, and each rung has a numeric gate. Nothing goes to the next rung on a good feeling.

RUNG 01 Shadow runs on every booking, logs, changes nothing risk: zero RUNG 02 Advisory recommendation shown in the console; a dispatcher clicks to apply it risk: a wasted click RUNG 03 Auto with veto the system acts; a dispatcher can undo inside a window before the rider is notified risk: bounded & reversible RUNG 04 Autonomous the system acts; humans handle escalations only risk: managed by §06 gate: 2 000 logged + agreement measured gate: acceptance > 80% · 0 violations gate: veto < 5% · outcome parity Any capability can sit on a different rung. Address resolution may be autonomous while assignment is still advisory — that is the point of separating them.
Figure III.2 — four rungs, three numeric gatesShadow mode is free and tells you more than any amount of design discussion. Run every new capability there first — including the ones you are confident about, because that is where the surprises are.
05

The loop you already built

This is the sharpest finding in the codebase, and it makes everything above cheaper.

assignment made selectMilerWithAI AgentDecision row ✓ Context jsonb ✓ Decision jsonb ✓ Reasoning text ✗ Outcome always null PATCH /internal/agent-decisions/:id/outcome routed · implemented · working called by nothing parcel resolves deliver · skip · RTO the one wire that is missing Column, index, endpoint and route all exist. One call from the resolution handlers turns every past assignment into a labelled training example — and gives you the replay corpus every gate in §04 is measured against.
Figure III.3 — 90% built, 0% connectedAgentDecision.Outcome, OutcomeRecordedAt, the controller and the route are all in place. Nothing writes to them. This is the cheapest high-value change available in the whole system.
Do this in week one

Call UpdateDecisionOutcome from the deliver, skip and RTO handlers with a simple verdict — delivered first attempt, delivered late, failed, reassigned. From that day forward every dispatch decision is scored, you can replay history against any new policy offline, and every claim anyone makes about the AI dispatcher becomes checkable. Until then, no gate in §04 can be measured at all.

06

Guardrails

  • Closed-set selection, enforced server-side. A model returning chosen_miler_id must have that ID checked against the candidate set before commit. Models hallucinate identifiers; a validator costs three lines and removes the entire failure class.
  • Keep the payload structured. Your current request to routemate carries no customer free text. The moment you pass raw Notes, address strings or client emails into a prompt that also drives actions, you have built a prompt-injection channel — anyone who can type into a booking form can try to instruct your dispatcher. If unstructured text must be read, do it in a separate call that only extracts, and never let that call's output reach an action without validation.
  • Tenant isolation is an invariant, not a preference. HubStaffAccount.Tenantid already isolates partner freight. Agents must be scoped identically, and cross-tenant leakage should be a hard validator failure, not a scoring penalty. This is the one mistake that is a breach rather than a bad route.
  • Every agent action is idempotent. You already have the Redis-backed key infrastructure. Agents retry more than humans do; a duplicated reassignment or a double COD entry is far worse than a slow one.
  • Budget per booking, enforced. Set a token ceiling per decision and per day, with a hard cutoff to the deterministic path. Cost overruns in agentic systems come from retry storms, not from steady state.
  • Mind what leaves the estate. Consignee phone numbers, COD amounts and precise home coordinates are in scope for whatever model endpoint you call. Decide deliberately what is sent, and prefer IDs over identities wherever the model does not need the human detail.
  • Escalate rather than improvise. Give the exception agent an explicit "page a human" tool and reward using it. An agent with no honourable way to stop will invent an action instead.
07

Ninety days

PhaseShipRungProves
Wk 1–2Wire UpdateDecisionOutcome from deliver / skip / RTO. Build the replay harness over AgentDecision.—Every later claim becomes measurable.
Wk 3–4Real ETA: Valhalla road time × time-of-day factor + learned dwell, replacing calculateETA.directNo model needed; immediate accuracy win on all three legs.
Wk 3–6Address resolution on pgvector. Embed delivered addresses, cluster, snap to DeliveryProof coordinates.shadow → advisoryHighest-leverage capability, lowest risk.
Wk 5–8VRP solver behind a flag: batched joint assign + sequence, running in shadow beside today's path.shadowHead-to-head on the replay corpus against live decisions.
Wk 7–9Explainer on solver output; dispatcher and rider-facing, in-language.advisoryAdoption. Riders follow routes they understand.
Wk 9–11NL constraint compiler with a human approval queue.advisoryTacit dispatcher knowledge starts reaching the solver.
Wk 10–12Exception agent with a bounded tool set, proposals only.advisoryThe escalation queue starts shrinking.
Wk 12Promote whichever capabilities have cleared their gates. Not the ones that haven't.→ auto with vetoThe ladder does the deciding, not the roadmap.
What is deliberately not in the first ninety days

Vision on inbound condition and POD, vernacular voice, and the operator copilot's action tier are all worth building — and all of them are easier once the outcome loop, the replay harness and the validator exist. Building those three foundations first is what makes the rest take weeks instead of quarters.