diff --git a/DOORMILE_AUTOMATION_PLAN.html b/DOORMILE_AUTOMATION_PLAN.html new file mode 100644 index 0000000..1b0dd64 --- /dev/null +++ b/DOORMILE_AUTOMATION_PLAN.html @@ -0,0 +1,654 @@ +
Implementation plan · fourth in the Doormile series
+What it would take to run the parcel pipeline without a dispatcher in it — decision by decision, from the one that is already automated to the leg that has no code in it at all. Grounded in doormile_backend as it stands today.
Full automation is not one switch, and a plan that treats it as one will fail in a specific way: it will automate the decisions that are easy to automate, leave the hard ones to a dispatcher, and then discover that the dispatcher is still working every day because the hard ones are the ones that happen.
+ +So the target has to be stated as a property of decisions, not of the system. There are nine decisions between a booking existing and a parcel being delivered. Fully automated means every one of them is made by code, every one is scored against its outcome, and a human is involved only where the system explicitly escalates.
+ +That last clause is the honest part. The ceiling is not zero humans; it is humans on the exception queue only. §09 says exactly what stays there and why that is the correct design rather than a shortfall.
+ +Each of the nine decisions climbs these rungs independently, and each rung above shadow has a numeric gate. Nothing is promoted on confidence.
+ +A person decides at a screen. No record exists of what they considered, so the decision cannot be replayed or improved.
Code decides in parallel and writes what it would have done. Nothing acts on it. Free, and where the surprises actually surface.
The proposal is shown to the operator as the default. They accept or override, and every override is a labelled training example.
The system acts on a timer. The operator sees it happen and has a window to reverse it. Most decisions should end their life here.
Acts and notifies nobody unless it escalates. Reserved for decisions with a measured outcome history and a cheap failure mode.
A decision moves up a rung when its replayed outcomes beat the incumbent on the corpus, and its escalation rate is stable for two weeks. A decision moves down a rung automatically when either regresses. The ladder is a controller, not a roadmap — it should be demoting things without anyone filing a ticket.
+This is the spine of the plan. Every row is a real decision the business makes hundreds of times a day, with what makes it today and where that sits on the ladder. Read the middle column as the honest current state, not the intended one.
+ +Two things fall out of this picture immediately. Nothing sits at L3 or L4, so nothing in Doormile currently acts without a person — the "AI dispatch" is L2 at best because no one can tell whether it beats the greedy fallback. And the cross-cutting pair at the bottom, ETA and exceptions, touch every leg, which makes them worth more than their row count suggests.
+Before any of the nine can climb a rung, one thing has to be true that is not true today: outcomes have to be recorded. Everything in this plan is gated on a change that takes about a day.
+ +The audit trail is built. AgentDecision carries Context, Decision, Reasoning, Outcome and OutcomeRecordedAt. UpdateDecisionOutcome exists in controllers/agentDecisionController.go. The route is live at PATCH /internal/agent-decisions/:id/outcome. There is even a context_embedding pgvector column that CreateAgentDecision populates and FindSimilarDecisions queries.
Nothing in the Go codebase calls it. Not one handler. Every dispatch decision Doormile has ever made is stored with a null outcome.
+ +Call the outcome update from the three terminal handlers — MilerDeliverConsignment, MilerSkipDelivery and MilerSkipPickup — with a verdict of delivered_first_attempt, delivered_late, failed or reassigned. From that day every decision is scored, and every claim about the dispatcher becomes checkable rather than arguable.
The second half of the same change is the replay harness: a command that takes a date range of AgentDecision rows, re-runs a candidate policy against the stored Context, and scores its choices against the recorded outcomes. That harness is what every gate in §00 reads from. Without it, promotion is a vote.
// the loop, closed — internal/assignment/outcome.go +func RecordOutcome(bookingID uint64, verdict string) { + // best-effort, exactly like AI_LAYER_FALLBACK: never block a delivery + // on the audit write. A missing outcome is a gap in the corpus, + // not a failed handover. +} ++ +
Write it as best-effort and off the request path, matching the pattern selectMilerWithAI already uses. A rider standing at a door must never wait on an analytics write.
D1 and D2 are both automated and both compromised by the same thing — they run in series. SequenceMilerStopsAsync is called from five places, and every one of them fires after the assignment has been committed. The solver is therefore always sequencing stops that someone else already chose.
booking.assignment_requested already exists as a NATS subject with a worker behind it. Buffer into a window keyed by hub and service type rather than handling each message on arrival.geoRadiusKm = 10.0. Widen until k genuine candidates are found. The fixed radius returns ten near-identical riders at noon and nothing at 6 AM, and both are failures of the same constant.arrivedat → pickup-complete pair in the database is a labelled example of how long a pickup point actually takes. Per-location dwell plus Valhalla road time plus a time-of-day factor replaces calculateETA outright.maxActive = 3 as a solver constraint, not a filter applied before the solver sees the candidate. It is a capacity limit, and capacity limits belong in the model.ai_layer.go already validates that the model's chosen_miler_id is inside the eligibility set and falls back to legacy scoring when it is not. That is the closed-set guardrail, and it is correctly implemented. The joint solver needs the same check on every rider in its returned plan.
This is the largest single gap and the easiest to close, because there is no incumbent behaviour to beat and no model quality argument to have. All three decisions are classical operations research with objectives you can write down, and the constraints are already columns in the database.
+ +| Decision | Technique | Constraints already stored | Objective |
|---|---|---|---|
| D4 — which parcels board which truck | +Bin packing | +Vehicle.Maxweight, Vehicle.Maxvolume, consignment dimensions, destination hub |
+ Fill vehicles by destination mix without breaching capacity or cut-off | +
| D5 — when it leaves | +Priced threshold | +Sladueat on every consignment and booking |
+ Marginal cost of an extra run vs expected SLA breach cost of holding | +
| D6 — whether the lane runs direct | +Network consolidation | +Tripsheet.Batchkind — the local / transfer split already exists |
+ Route thin lanes through a transfer hub instead of running half-empty | +
D5 is the one worth doing first, and it is worth doing even before D4. It is a scalar comparison, not a search: for the parcels currently at the hub, the cost of dispatching now is known, and the cost of holding is the sum of breach probabilities against Sladueat. Today that trade is a dispatcher's instinct and is never written down. Turning it into a number, shown on the hub console next to the tripsheet, is an advisory-rung feature that needs no solver at all.
D4 then automates the tripsheet build: propose a manifest, let hub staff scan against it, and treat every manual deviation as a labelled correction. The scan step stays exactly as it is — TripsheetItem.Scanstatus already carries Pending / Loaded / Unloaded / Discrepancy, which is precisely the reconciliation surface an auto-built manifest needs.
Tripsheets have no cut-off time and no planned capacity fields. Both D4 and D5 need them. That is the only schema change mid-mile automation requires — everything else it needs is already a column.
+Every failed delivery is a second full run for one parcel, and first-attempt success is the dominant cost line in Indian last mile. Doormile enters that retry loop on discovery — the rider reaches the door and finds nobody there. The information needed to predict it is already being written on every failure.
+ +Attemptcount and the skip reason are recorded on every failed attempt. That is a labelled training set sitting unused. Predict per-stop success, then sequence on it and shift low-probability stops to a window where they are likely to land, rather than burning an attempt to learn what the data already says.DeliveryProof records. Your delivered history becomes a private geocoder that beats any general one inside your own zones, because it is built from the doors your riders actually found.D9 is the highest-leverage item in the plan and among the lowest risk, because it can run in shadow indefinitely: compute the corrected coordinate, log the delta against what the geocoder said, and change nothing until the deltas are demonstrably better. It also compounds — every leg's routing improves when the destination coordinates are right.
+Automating the nine decisions gets you a pipeline that runs itself on a good day. What keeps a dispatcher employed is the bad day — a rider goes dark mid-round, a vehicle breaks down with forty parcels aboard, a client calls at 4 PM to move a pickup. These are exactly the decisions a solver is bad at and a language model is good at, because they are unstructured, contextual and rare.
+ +This is the one place in the plan where generative AI does something irreplaceable, and the design rule from the agent-layer document holds: the model proposes, the solver decides, the system validates. The exception agent's job is to turn a mess into a structured recovery request that the existing solver can price, not to choose the recovery itself.
+ +The agent never writes to the database. It calls the same internal endpoints a human operator would, so tenant scoping, idempotency, validation and audit apply identically whether the caller is a person or a model. If an agent can reach a table your dispatcher cannot, you have built two systems and will only secure one.
+Give it a bounded tool set — reassign a stop, rebuild a manifest, extend a window, notify a consignee, page a human — and reward the last one. An agent with no honourable way to stop will invent an action instead. The escalation rate is also the metric that tells you whether the rest of the automation is working: it should fall as the other eight decisions climb, and a rising escalation rate is the earliest signal that something below it has regressed.
+Running unattended is a different engineering problem from deciding well. This is the machinery that makes it safe to let the nine decisions act, and most of it already exists in the estate.
+ +ai_layer.go; apply the same pattern to solver plans and agent tool calls. Models hallucinate identifiers, and a validator is three lines.Idempotency-Key middleware is already on /consignments/:id/deliver. Agents retry more than humans do, and a duplicated reassignment or double COD entry is far worse than a slow one.HubStaffAccount.Tenantid already isolates partner freight. Cross-tenant leakage must fail the plan, not cost it points. This is the one mistake that is a breach rather than a bad route.routing.BaseURL already models this correctly: empty disables sequencing, and stops simply stay unsequenced rather than assignment failing.Sequenced by ratio of effect to effort and by what unblocks what — not by ambition. The first three need no new infrastructure and no model.
+ +| # | Ship | Decisions | Depends on | Target rung |
|---|---|---|---|---|
| 1 | Close the outcome loop. Call the outcome update from the deliver, skip-delivery and skip-pickup handlers; build the replay harness over AgentDecision. | all | nothing | enables all |
| 2 | Real ETA. Valhalla road time × time-of-day factor + learned dwell, replacing calculateETA. | ETA | #1 to measure it | L4 direct |
| 3 | Adaptive candidate set. Widen the radius until k real candidates exist; move maxActive into the solver as a constraint. | D3 | nothing | L4 |
| 4 | Hold-vs-go, priced. Show the dispatch-now vs hold number on the hub console against Sladueat. | D5 | #1 | L2 advisory |
| 5 | Address resolution on pgvector. Embed delivered addresses, cluster, snap to DeliveryProof coordinates. | D9 | delivered history | L1 → L3 |
| 6 | Batched joint assign + sequence behind a flag, running in shadow beside today's path. | D1, D2, D7 | #2, #3, #5 | L1 → L3 |
| 7 | Auto-built tripsheets. Bin-pack the manifest; reconcile against the existing scan flow. | D4 | #4, cut-off fields | L2 → L3 |
| 8 | First-attempt success model, sequencing stops on predicted success. | D8 | #5, attempt history | L1 → L3 |
| 9 | Exception agent, bounded tool set, proposals only at first. | exceptions | #1, #6 | L2 → L3 |
| 10 | Lane consolidation. Route thin lanes via transfer hubs using Batchkind. | D6 | #7, volume forecast | L2 |
Items 1 through 3 are the ones to commit to now. None of them involves a model, all three are measurable within a fortnight, and #1 is the difference between a plan and a hope: until outcomes are recorded, no gate in this document can be evaluated and every later item is unfalsifiable.
+ +Item 6 is the largest routing gain available and it is deliberately sixth. It depends on a real cost matrix (#2), a sane candidate set (#3) and trustworthy coordinates (#5). Built before those, it is a better solver fed worse data, and it will underperform the greedy path it replaces — which is the failure mode most likely to kill the whole programme politically.
+A plan for full automation is only credible if it says where automation stops. These are not gaps to close later; they are the designed boundary.
+ +Measured against that boundary, the target is reachable and specific: nine decisions made by code, one loop scoring all of them, and a dispatcher whose day is the exception queue rather than the queue. Doormile is closer to it than the current state suggests — the audit table, the vector store, the road-network client, the idempotency middleware and the async worker are all already built. The thing standing between the estate and full automation is not infrastructure. It is that nobody is writing down what happened.
+