Integration plan · third in the Doormile series
Where generative AI genuinely earns its place in the platform — auto-assignment included — how it plugs into the Go backend you already have, and the trust ladder that gets it from logging quietly to acting alone without a bad week in between.
The model proposes. The solver decides. The system validates.
Every capability below is an application of that one sentence. A language model is exceptional at turning mess into structure — a dispatcher's sentence into a constraint, a WhatsApp forward into a booking, a photo into a condition code, a solver's output into an explanation a rider will actually read. It is poor and expensive at choosing the best of ten thousand arrangements, which is exactly what dispatch is.
So "auto-assign with GenAI" does not mean a model picking riders. It means a model doing the four things around the pick that were previously impossible to automate, while a solver does the pick itself in milliseconds and a validator refuses anything that breaks an invariant. That is a more capable system than an LLM choosing alone, and a cheaper one.
The payload your backend sends to routemate is entirely structured — miler_id, distance_km, rating, on_time_rate_30d, hub_load, hub_capacity, plus hour, peak flag and zone. No customer-supplied free text goes into it. That closed payload is a security property, and §06 explains why keeping it is the single most important rule as you expand.
Ten capabilities, ranked by what they return for what they cost. Note how few of them are "the AI decides" and how many are "the AI turns something unstructured into something the system can already handle."
| # | Capability | Leg | What the model does | Value | Effort |
|---|---|---|---|---|---|
| 1 | Address resolution | last | Embeds messy Indian address text, clusters duplicates, snaps to confirmed DeliveryProof coordinates. Builds you a private geocoder from your own delivered history. | highest | low |
| 2 | Decision explanation | all | Turns solver output into a sentence for the dispatcher and a line for the rider. Drives adoption more than accuracy does. | high | low |
| 3 | Unstructured intake | first | A client's WhatsApp message, email or PDF manifest becomes structured bookings. You support Excel; clients don't always send Excel. | high | medium |
| 4 | NL constraint compiler | all | "Don't give Ravi the mall runs" becomes a typed solver constraint a human approves once and the solver honours forever. | high | medium |
| 5 | Exception agent | all | Rider absent, vehicle down, forty parcels stranded. Tool-calling over a bounded action space, proposing a recovery plan. | high | medium |
| 6 | Damage detection at inbound | mid | Vision on the inbound scan photo auto-populates Condition, which is a manual dropdown today and therefore mostly says Good. | medium | low |
| 7 | POD verification | last | Vision checks the delivery photo actually shows a parcel at a door, and that the signature isn't blank. Catches soft fraud. | medium | low |
| 8 | Vernacular voice | first, last | Rider status by speech in Tamil, Telugu, Kannada or Hindi; an outbound call confirming the consignee will be home. Directly attacks first-attempt failure. | high | high |
| 9 | Operator copilot | all | Extends the existing console assistant from read-only queries to proposing actions with confirmation. | medium | medium |
| 10 | Scenario simulation | all | Replays historical days against a proposed policy before it touches production. Your safety net for everything above. | medium | medium |
Four hubs across Coimbatore, Hyderabad, Bengaluru and Chennai means four language regions. Item 8 is not a nicety in that footprint — a rider who can speak a status update is a rider who actually files one.
This is the thing you asked about, drawn end to end. The deterministic spine runs down the left and never waits on a model. The three GenAI touch points sit beside it, feeding in out of band.
Dispatchers hold dozens of rules in their heads that never reach the system: which rider is slow at a particular gate, which client insists on the same face, which lane floods in monsoon. Today the only way to encode those is a developer ticket, so they stay in the dispatcher's head and the optimiser looks stupid. A model that turns a typed sentence into a reviewable constraint — which a human approves before it ever binds — converts tacit knowledge into solver input at conversational speed. That is a capability the solver cannot have on its own, and it is where generative AI is doing something genuinely irreplaceable.
// what the compiler emits — reviewed by a human, then stored
{
"type": "exclude_miler_from_location_class",
"miler_id": 412,
"location_class": "mall_service_gate",
"reason": "slow at gate handover",
"author": "dispatcher:coimbatore:7",
"expires_at": null
}
You do not need new infrastructure. Every piece below already exists in the estate; the work is contracts and placement.
| Piece | Already there | What changes |
|---|---|---|
| Async invocation | NATS JetStream, the booking worker, a 60 s budget | Model calls become subscribers, not inline HTTP. Nothing in a request path waits on one. |
| Fallback discipline | AI_LAYER_FALLBACK and the 5 s cap in selectMilerWithAI | Keep the pattern exactly; apply it to every new capability. Best-effort or it does not ship. |
| Audit | AgentDecision with a DecisionType discriminator | One row per model call, whatever the capability. The discriminator was clearly built for this. |
| Idempotency | Redis-backed Idempotency-Key on payment, pickup and deliver | Extend to every agent-initiated write. An agent that retries must not double-act. |
| Vector store | pgvector, in the database already | Address embeddings land here. No new dependency. |
| Road distance | Valhalla behind routes.workolik.com | Becomes the solver's cost matrix rather than a post-hoc sequencer. |
| Object storage | presigned uploads for signatures and photos | The same URLs feed the vision capabilities. Images never transit your API. |
| Console surface | the existing deterministic assistant and intents | Gains an action tier behind a confirmation step. |
One deliberate constraint: the agent layer never writes to the database directly. It calls the same internal endpoints a human operator would, so tenant scoping, idempotency, validation and audit apply identically whether the caller is a person or a model. If an agent can reach a table your dispatcher cannot, you have built two systems and will secure only one.
Every capability climbs the same four rungs, and each rung has a numeric gate. Nothing goes to the next rung on a good feeling.
This is the sharpest finding in the codebase, and it makes everything above cheaper.
AgentDecision.Outcome, OutcomeRecordedAt, the controller and the route are all in place. Nothing writes to them. This is the cheapest high-value change available in the whole system.Call UpdateDecisionOutcome from the deliver, skip and RTO handlers with a simple verdict — delivered first attempt, delivered late, failed, reassigned. From that day forward every dispatch decision is scored, you can replay history against any new policy offline, and every claim anyone makes about the AI dispatcher becomes checkable. Until then, no gate in §04 can be measured at all.
chosen_miler_id must have that ID checked against the candidate set before commit. Models hallucinate identifiers; a validator costs three lines and removes the entire failure class.routemate carries no customer free text. The moment you pass raw Notes, address strings or client emails into a prompt that also drives actions, you have built a prompt-injection channel — anyone who can type into a booking form can try to instruct your dispatcher. If unstructured text must be read, do it in a separate call that only extracts, and never let that call's output reach an action without validation.HubStaffAccount.Tenantid already isolates partner freight. Agents must be scoped identically, and cross-tenant leakage should be a hard validator failure, not a scoring penalty. This is the one mistake that is a breach rather than a bad route.| Phase | Ship | Rung | Proves |
|---|---|---|---|
| Wk 1–2 | Wire UpdateDecisionOutcome from deliver / skip / RTO. Build the replay harness over AgentDecision. | — | Every later claim becomes measurable. |
| Wk 3–4 | Real ETA: Valhalla road time × time-of-day factor + learned dwell, replacing calculateETA. | direct | No model needed; immediate accuracy win on all three legs. |
| Wk 3–6 | Address resolution on pgvector. Embed delivered addresses, cluster, snap to DeliveryProof coordinates. | shadow → advisory | Highest-leverage capability, lowest risk. |
| Wk 5–8 | VRP solver behind a flag: batched joint assign + sequence, running in shadow beside today's path. | shadow | Head-to-head on the replay corpus against live decisions. |
| Wk 7–9 | Explainer on solver output; dispatcher and rider-facing, in-language. | advisory | Adoption. Riders follow routes they understand. |
| Wk 9–11 | NL constraint compiler with a human approval queue. | advisory | Tacit dispatcher knowledge starts reaching the solver. |
| Wk 10–12 | Exception agent with a bounded tool set, proposals only. | advisory | The escalation queue starts shrinking. |
| Wk 12 | Promote whichever capabilities have cleared their gates. Not the ones that haven't. | → auto with veto | The ladder does the deciding, not the roadmap. |
Vision on inbound condition and POD, vernacular voice, and the operator copilot's action tier are all worth building — and all of them are easier once the outcome loop, the replay harness and the validator exist. Building those three foundations first is what makes the rest take weeks instead of quarters.