diff --git a/DOORMILE_AUTOMATION_PLAN.html b/DOORMILE_AUTOMATION_PLAN.html new file mode 100644 index 0000000..1b0dd64 --- /dev/null +++ b/DOORMILE_AUTOMATION_PLAN.html @@ -0,0 +1,654 @@ +Doormile Automation Plan + + + + + + + + +
+ +
+

Implementation plan · fourth in the Doormile series

+

Doormile Automation Plan

+

What it would take to run the parcel pipeline without a dispatcher in it — decision by decision, from the one that is already automated to the leg that has no code in it at all. Grounded in doormile_backend as it stands today.

+
+ + + +
+
00

What "fully automated" has to mean

+ +

Full automation is not one switch, and a plan that treats it as one will fail in a specific way: it will automate the decisions that are easy to automate, leave the hard ones to a dispatcher, and then discover that the dispatcher is still working every day because the hard ones are the ones that happen.

+ +

So the target has to be stated as a property of decisions, not of the system. There are nine decisions between a booking existing and a parcel being delivered. Fully automated means every one of them is made by code, every one is scored against its outcome, and a human is involved only where the system explicitly escalates.

+ +

That last clause is the honest part. The ceiling is not zero humans; it is humans on the exception queue only. §09 says exactly what stays there and why that is the correct design rather than a shortfall.

+ +

The autonomy ladder

+

Each of the nine decisions climbs these rungs independently, and each rung above shadow has a numeric gate. Nothing is promoted on confidence.

+ +
+
L0
Manual

A person decides at a screen. No record exists of what they considered, so the decision cannot be replayed or improved.

+
L1
Shadow

Code decides in parallel and writes what it would have done. Nothing acts on it. Free, and where the surprises actually surface.

+
L2
Advisory

The proposal is shown to the operator as the default. They accept or override, and every override is a labelled training example.

+
L3
Auto with veto

The system acts on a timer. The operator sees it happen and has a window to reverse it. Most decisions should end their life here.

+
L4
Autonomous

Acts and notifies nobody unless it escalates. Reserved for decisions with a measured outcome history and a cheap failure mode.

+
+ +
+ The gate, in one rule +

A decision moves up a rung when its replayed outcomes beat the incumbent on the corpus, and its escalation rate is stable for two weeks. A decision moves down a rung automatically when either regresses. The ladder is a controller, not a roadmap — it should be demoting things without anyone filing a ticket.

+
+
+ +
+
01

The nine decisions

+ +

This is the spine of the plan. Every row is a real decision the business makes hundreds of times a day, with what makes it today and where that sits on the ladder. Read the middle column as the honest current state, not the intended one.

+ +
+
+ + + + + + + + FIRST MILE + MID MILE + LAST MILE + + + + + + + + D1 · Which rider collects + GEOSEARCH 10 km → LLM picks one → 5 s + fallback to greedy nearest + L2 + + + D2 · Order of a rider's stops + Valhalla sequencing — but only after + assignment is already committed + L2 + + + D3 · How wide to search + Hard-coded 10 km, max 3 active jobs. + Same at noon and at 6 AM. + L0 + + + + D4 · Which parcels board + A person builds the tripsheet by hand, + scanning parcels into a vehicle. + L0 + + + D5 · When the truck leaves + Dispatcher instinct. Cost of a half-empty + run vs SLA risk is never priced. + L0 + + + D6 · Whether a lane exists + Standing arrangement. Thin lanes run + direct rather than consolidating. + L0 + + + + D7 · Which rider delivers + Hub staff assign from the console after + offload. Batch-assign is greedy nearest. + L0 + + + D8 · Will this attempt land + Not decided at all. Failure is discovered + at the door, then retried in full. + L0 + + + D9 · Where the address is + General geocoder, corrected by hand + when a rider cannot find the door. + L1 + + + + ACROSS ALL THREE LEGS + + + How long anything will take + calculateETA is a flat 20 km/h against straight-line distance, + plus a fixed ten-minute buffer. Sets every promise made. + L0 + + + What to do when it breaks + Rider absent, vehicle down, forty parcels stranded — every + recovery is improvised by a person, and none is recorded. + L0 + + + + L0 — a person decides, or nothing does + + L1–L2 — code decides, unmeasured + + L3–L4 — acts on its own (nothing is here yet) + +
+
Figure 1 — where the automation actually isTwo of nine decisions have code making them, and neither is scored. The middle column being entirely red is the finding: mid mile has no automation at all, which also makes it the cheapest leg to improve, because there is no incumbent to beat.
+
+ +

Two things fall out of this picture immediately. Nothing sits at L3 or L4, so nothing in Doormile currently acts without a person — the "AI dispatch" is L2 at best because no one can tell whether it beats the greedy fallback. And the cross-cutting pair at the bottom, ETA and exceptions, touch every leg, which makes them worth more than their row count suggests.

+
+ +
+
02

The blocker: nothing is measured

+ +

Before any of the nine can climb a rung, one thing has to be true that is not true today: outcomes have to be recorded. Everything in this plan is gated on a change that takes about a day.

+ +

The audit trail is built. AgentDecision carries Context, Decision, Reasoning, Outcome and OutcomeRecordedAt. UpdateDecisionOutcome exists in controllers/agentDecisionController.go. The route is live at PATCH /internal/agent-decisions/:id/outcome. There is even a context_embedding pgvector column that CreateAgentDecision populates and FindSimilarDecisions queries.

+ +

Nothing in the Go codebase calls it. Not one handler. Every dispatch decision Doormile has ever made is stored with a null outcome.

+ +
+ Week one, and it unblocks everything else +

Call the outcome update from the three terminal handlers — MilerDeliverConsignment, MilerSkipDelivery and MilerSkipPickup — with a verdict of delivered_first_attempt, delivered_late, failed or reassigned. From that day every decision is scored, and every claim about the dispatcher becomes checkable rather than arguable.

+
+ +

The second half of the same change is the replay harness: a command that takes a date range of AgentDecision rows, re-runs a candidate policy against the stored Context, and scores its choices against the recorded outcomes. That harness is what every gate in §00 reads from. Without it, promotion is a vote.

+ +
// the loop, closed — internal/assignment/outcome.go
+func RecordOutcome(bookingID uint64, verdict string) {
+    // best-effort, exactly like AI_LAYER_FALLBACK: never block a delivery
+    // on the audit write. A missing outcome is a gap in the corpus,
+    // not a failed handover.
+}
+
+ +

Write it as best-effort and off the request path, matching the pattern selectMilerWithAI already uses. A rider standing at a door must never wait on an analytics write.

+
+ +
+
03

First mile: stop deciding one parcel at a time

+ +

D1 and D2 are both automated and both compromised by the same thing — they run in series. SequenceMilerStopsAsync is called from five places, and every one of them fires after the assignment has been committed. The solver is therefore always sequencing stops that someone else already chose.

+ +
+
+ + + + + + + + + + + TODAY — ONE PARCEL AT A TIME + + + one booking + NATS + + + + GEOSEARCH 10 km + fixed radius + + + + LLM picks 1 rider + 5 s cap, on path + + + + commit assignment + irreversible + + + + then sequence + Valhalla, too late + + Each parcel is committed to a rider before the next parcel is even looked at. The solver inherits a set it had no say in. + + + + TARGET — SOLVE THE WINDOW + + + accumulate 10–15 min + of open pickups + batch window + + + + adaptive radius + widen until k found + replaces fixed 10 km + + + + joint solve + who carries it AND + in what order — one problem + Valhalla cost matrix + + + + validate + commit + whole plan + one transaction + + + + veto window + then it stands + L3 + + The model is not on this path at all. It compiles constraints into the solver beforehand and explains the output afterwards — §06 and §07. + Superfast bookings bypass the window entirely and take today's immediate path. A batch window is a latency budget, and not every parcel has one. + +
+
Figure 2 — the ordering changeThe Valhalla client already exists and already speaks Doormile's vocabulary. The work is moving it in front of the commit and feeding it the whole open set rather than one rider's inherited stops. Note what leaves the critical path: the five-second model call.
+
+ +

What changes in code

+ + +
+ Already done — do not rebuild it +

ai_layer.go already validates that the model's chosen_miler_id is inside the eligibility set and falls back to legacy scoring when it is not. That is the closed-set guardrail, and it is correctly implemented. The joint solver needs the same check on every rider in its returned plan.

+
+
+ +
+
04

Mid mile: three decisions, no code

+ +

This is the largest single gap and the easiest to close, because there is no incumbent behaviour to beat and no model quality argument to have. All three decisions are classical operations research with objectives you can write down, and the constraints are already columns in the database.

+ +
+ + + + + + + + + + + + + + + + + + + + + + +
DecisionTechniqueConstraints already storedObjective
D4 — which parcels board which truckBin packingVehicle.Maxweight, Vehicle.Maxvolume, consignment dimensions, destination hubFill vehicles by destination mix without breaching capacity or cut-off
D5 — when it leavesPriced thresholdSladueat on every consignment and bookingMarginal cost of an extra run vs expected SLA breach cost of holding
D6 — whether the lane runs directNetwork consolidationTripsheet.Batchkind — the local / transfer split already existsRoute thin lanes through a transfer hub instead of running half-empty
+
+ +

D5 is the one worth doing first, and it is worth doing even before D4. It is a scalar comparison, not a search: for the parcels currently at the hub, the cost of dispatching now is known, and the cost of holding is the sum of breach probabilities against Sladueat. Today that trade is a dispatcher's instinct and is never written down. Turning it into a number, shown on the hub console next to the tripsheet, is an advisory-rung feature that needs no solver at all.

+ +

D4 then automates the tripsheet build: propose a manifest, let hub staff scan against it, and treat every manual deviation as a labelled correction. The scan step stays exactly as it is — TripsheetItem.Scanstatus already carries Pending / Loaded / Unloaded / Discrepancy, which is precisely the reconciliation surface an auto-built manifest needs.

+ +
+ The one thing to add to the schema +

Tripsheets have no cut-off time and no planned capacity fields. Both D4 and D5 need them. That is the only schema change mid-mile automation requires — everything else it needs is already a column.

+
+
+ +
+
05

Last mile: predict instead of discover

+ +

Every failed delivery is a second full run for one parcel, and first-attempt success is the dominant cost line in Indian last mile. Doormile enters that retry loop on discovery — the rider reaches the door and finds nobody there. The information needed to predict it is already being written on every failure.

+ + + +

D9 is the highest-leverage item in the plan and among the lowest risk, because it can run in shadow indefinitely: compute the corrected coordinate, log the delta against what the geocoder said, and change nothing until the deltas are demonstrably better. It also compounds — every leg's routing improves when the destination coordinates are right.

+
+ +
+
06

The exception agent: the last human out of the loop

+ +

Automating the nine decisions gets you a pipeline that runs itself on a good day. What keeps a dispatcher employed is the bad day — a rider goes dark mid-round, a vehicle breaks down with forty parcels aboard, a client calls at 4 PM to move a pickup. These are exactly the decisions a solver is bad at and a language model is good at, because they are unstructured, contextual and rare.

+ +

This is the one place in the plan where generative AI does something irreplaceable, and the design rule from the agent-layer document holds: the model proposes, the solver decides, the system validates. The exception agent's job is to turn a mess into a structured recovery request that the existing solver can price, not to choose the recovery itself.

+ +
+ The boundary that makes it safe +

The agent never writes to the database. It calls the same internal endpoints a human operator would, so tenant scoping, idempotency, validation and audit apply identically whether the caller is a person or a model. If an agent can reach a table your dispatcher cannot, you have built two systems and will only secure one.

+
+ +

Give it a bounded tool set — reassign a stop, rebuild a manifest, extend a window, notify a consignee, page a human — and reward the last one. An agent with no honourable way to stop will invent an action instead. The escalation rate is also the metric that tells you whether the rest of the automation is working: it should fall as the other eight decisions climb, and a rising escalation rate is the earliest signal that something below it has regressed.

+
+ +
+
07

The control plane

+ +

Running unattended is a different engineering problem from deciding well. This is the machinery that makes it safe to let the nine decisions act, and most of it already exists in the estate.

+ +
+
+ + + + + + + + + decide + AgentDecision row + + + + act + idempotent write + + + + record outcome + missing today + + + + replay + policy vs corpus + + + + gate + promote / demote + + + + the rung changes what "decide" may do + + Break the loop at "record outcome" and the two boxes after it cannot run at all — which is the state the system is in today. + Every capability rides the same loop. There is one control plane, not one per feature. + +
+
Figure 3 — one loop, every decisionThe ladder in §00 is the feedback path of this loop rather than a project schedule. When it is closed, autonomy levels move on their own evidence; when it is open, as now, every promotion is someone's opinion.
+
+ +

Invariants that must hold before anything reaches L3

+ +
+ +
+
08

Build order

+ +

Sequenced by ratio of effect to effort and by what unblocks what — not by ambition. The first three need no new infrastructure and no model.

+ +
+ + + + + + + + + + + + + + +
#ShipDecisionsDepends onTarget rung
1Close the outcome loop. Call the outcome update from the deliver, skip-delivery and skip-pickup handlers; build the replay harness over AgentDecision.allnothingenables all
2Real ETA. Valhalla road time × time-of-day factor + learned dwell, replacing calculateETA.ETA#1 to measure itL4 direct
3Adaptive candidate set. Widen the radius until k real candidates exist; move maxActive into the solver as a constraint.D3nothingL4
4Hold-vs-go, priced. Show the dispatch-now vs hold number on the hub console against Sladueat.D5#1L2 advisory
5Address resolution on pgvector. Embed delivered addresses, cluster, snap to DeliveryProof coordinates.D9delivered historyL1 → L3
6Batched joint assign + sequence behind a flag, running in shadow beside today's path.D1, D2, D7#2, #3, #5L1 → L3
7Auto-built tripsheets. Bin-pack the manifest; reconcile against the existing scan flow.D4#4, cut-off fieldsL2 → L3
8First-attempt success model, sequencing stops on predicted success.D8#5, attempt historyL1 → L3
9Exception agent, bounded tool set, proposals only at first.exceptions#1, #6L2 → L3
10Lane consolidation. Route thin lanes via transfer hubs using Batchkind.D6#7, volume forecastL2
+
+ +

Items 1 through 3 are the ones to commit to now. None of them involves a model, all three are measurable within a fortnight, and #1 is the difference between a plan and a hope: until outcomes are recorded, no gate in this document can be evaluated and every later item is unfalsifiable.

+ +
+ The dependency worth respecting +

Item 6 is the largest routing gain available and it is deliberately sixth. It depends on a real cost matrix (#2), a sane candidate set (#3) and trustworthy coordinates (#5). Built before those, it is a better solver fed worse data, and it will underperform the greedy path it replaces — which is the failure mode most likely to kill the whole programme politically.

+
+
+ +
+
09

What stays human, permanently

+ +

A plan for full automation is only credible if it says where automation stops. These are not gaps to close later; they are the designed boundary.

+ + + +

Measured against that boundary, the target is reachable and specific: nine decisions made by code, one loop scoring all of them, and a dispatcher whose day is the exception queue rather than the queue. Doormile is closer to it than the current state suggests — the audit table, the vector store, the road-network client, the idempotency middleware and the async worker are all already built. The thing standing between the estate and full automation is not infrastructure. It is that nobody is writing down what happened.

+
+ + + +
diff --git a/src/api/doormile/queries.js b/src/api/doormile/queries.js index e40f6df..cae2ecd 100644 --- a/src/api/doormile/queries.js +++ b/src/api/doormile/queries.js @@ -83,6 +83,8 @@ const BOOKING_STATUS_TO_DELIVERY_STATUS = { pending_assignment: 'pending', pending: 'pending', miler_assigned: 'pending', + assigned: 'pending', + rider_assigned: 'pending', pickup_scheduled: 'accepted', arrived: 'arrived', converted_to_consignment: 'picked', @@ -797,9 +799,17 @@ export const fetchDeliveries = async ({ pageParam = 1, queryKey }) => { const isDispatched = (b) => { if (!b) return false; - const riderId = b.assignedmileruserid ?? b.assigned_miler_user_id ?? b.mileruserid ?? b.milerid ?? b.assignedmilerid; + const riderId = + b.assignedmileruserid ?? + b.assigned_miler_user_id ?? + b.mileruserid ?? + b.milerid ?? + b.assignedmilerid ?? + b.milerprofileid ?? + b.riderid ?? + b.rider_id; const consignmentId = b.consignmentid ?? b.consignment_id; - const rawStatus = String(b.status || b.orderstatus || '').trim().toLowerCase(); + const rawStatus = String(b.status || b.orderstatus || b.bookingstatus || '').trim().toLowerCase(); if (riderId || consignmentId) return true; if (rawStatus && !['pending_pickup', 'created', 'new', 'booked', 'order_placed', 'unassigned', ''].includes(rawStatus)) { return true; @@ -819,6 +829,9 @@ export const fetchDeliveries = async ({ pageParam = 1, queryKey }) => { b.mileruserid ?? b.milerid ?? b.assignedmilerid ?? + b.milerprofileid ?? + b.riderid ?? + b.rider_id ?? consignment?.assignedmileruserid ?? consignment?.mileruserid ?? consignment?.milerid; @@ -841,6 +854,8 @@ export const fetchDeliveries = async ({ pageParam = 1, queryKey }) => { b.milername || b.ridername || b.assignedmilername || + b.miler_name || + b.rider_name || consignment?.milername || consignment?.ridername || consignment?.assignedmilername || @@ -853,10 +868,14 @@ export const fetchDeliveries = async ({ pageParam = 1, queryKey }) => { bookingno: b.bookingno, orderid: b.bookingno || (b.bookingid ? `#${b.bookingid}` : (b.id ? `#${b.id}` : '')), consignmentid: consignmentId, + hubid: b.hubid, + sourcehubid: b.sourcehubid, tenantid: b.tenantid, tenantname: tenant?.tenantname || '', tenantsuburb: '', applocation: '', + applocationid: b.applocationid, + tenantlocationid: b.tenantlocationid, tenantadress: tenant?.primaryemail || '', locationname: tenant?.tenantname || '', locationsuburb: '', @@ -898,10 +917,10 @@ export const fetchDeliveries = async ({ pageParam = 1, queryKey }) => { collectionamt: undefined, notes: b.notes || '', deliverytype: customer ? 'B' : 'C', - orderdate: b.createdat, + orderdate: b.createdat || b.orderdate || b.updatedat, deliverydate: b.serviceoptions?.[0]?.estimateddeliveryat || b.updatedat, reachedat: reached, - assigntime: b.updatedat, + assigntime: b.updatedat || b.createdat, // The consignment's status WINS when there is one (e.g. Collected_By_Miler -> Picked, // Out_for_Delivery -> Active, Delivered -> Delivered). // If status == Converted_To_Consignment, switches to consignmentstatus. @@ -961,9 +980,17 @@ const getDeliveryStatusCounts = async () => { }); const isDispatched = (b) => { if (!b) return false; - const riderId = b.assignedmileruserid ?? b.assigned_miler_user_id ?? b.mileruserid ?? b.milerid ?? b.assignedmilerid; + const riderId = + b.assignedmileruserid ?? + b.assigned_miler_user_id ?? + b.mileruserid ?? + b.milerid ?? + b.assignedmilerid ?? + b.milerprofileid ?? + b.riderid ?? + b.rider_id; const consignmentId = b.consignmentid ?? b.consignment_id; - const rawStatus = String(b.status || b.orderstatus || '').trim().toLowerCase(); + const rawStatus = String(b.status || b.orderstatus || b.bookingstatus || '').trim().toLowerCase(); if (riderId || consignmentId) return true; if (rawStatus && !['pending_pickup', 'created', 'new', 'booked', 'order_placed', 'unassigned', ''].includes(rawStatus)) { return true; diff --git a/src/lib/ZoneContext.jsx b/src/lib/ZoneContext.jsx index 8f4722a..01ee7a5 100644 --- a/src/lib/ZoneContext.jsx +++ b/src/lib/ZoneContext.jsx @@ -108,11 +108,25 @@ export function ZoneProvider({ children }) { const currentHub = selectedZone; const targetHubId = String(currentHub.hubid); + const isOrderOrDeliveryOrTripsheet = Boolean( + item.bookingid != null || + item.orderheaderid != null || + item.deliveryid != null || + item.bookingno != null || + item.consignmentid != null || + item.tripsheetid != null || + item.pickupaddress != null || + item.deliveryaddress != null + ); + // Rider matching: riders belong to the tenant fleet and serve all tenant kitchen hubs const isRiderItem = - item.milerprofileid != null || - item.authname != null || - (item.userid != null && item.availabilitystatus != null); + !isOrderOrDeliveryOrTripsheet && + Boolean( + item.milerprofileid != null || + item.authname != null || + (item.userid != null && item.availabilitystatus != null) + ); if (isRiderItem) { if (isTenantUser) { diff --git a/tests/lib/ZoneContext.test.jsx b/tests/lib/ZoneContext.test.jsx index cd36603..e2cd983 100644 --- a/tests/lib/ZoneContext.test.jsx +++ b/tests/lib/ZoneContext.test.jsx @@ -298,6 +298,22 @@ describe('ZoneContext', () => { expect(result.current.matchesZone({ pickuplatitude: null, pickuplongitude: null })).toBe(false); expect(result.current.matchesZone({ pickuplatitude: 'abc', pickuplongitude: 'abc' })).toBe(false); }); + + it('should correctly match a delivery row that carries an assigned rider (milerprofileid) by address/city', () => { + const { result } = selectKoramangala(); + // Even though milerprofileid and userid are present, this is a delivery row (has bookingid/pickupaddress) + // and should match Koramangala Hub's city (Bengaluru). + const deliveryRow = { + bookingid: 4021, + bookingno: 'DM-BK-7959CF19-97503', + pickupaddress: '12 MG Road, Bengaluru 560001', + deliveryaddress: '44 Indiranagar, Bengaluru 560038', + milerprofileid: 11, + userid: 91, + ridername: 'Vasanth' + }; + expect(result.current.matchesZone(deliveryRow)).toBe(true); + }); }); describe('matchesZone — riders', () => {