Integration plan · third in the Doormile series

Doormile Agent Layer

Where generative AI genuinely earns its place in the platform — auto-assignment included — how it plugs into the Go backend you already have, and the trust ladder that gets it from logging quietly to acting alone without a bad week in between.

00

The rule that makes all of this work

The model proposes. The solver decides. The system validates.

Every capability below is an application of that one sentence. A language model is exceptional at turning mess into structure — a dispatcher's sentence into a constraint, a WhatsApp forward into a booking, a photo into a condition code, a solver's output into an explanation a rider will actually read. It is poor and expensive at choosing the best of ten thousand arrangements, which is exactly what dispatch is.

So "auto-assign with GenAI" does not mean a model picking riders. It means a model doing the four things around the pick that were previously impossible to automate, while a solver does the pick itself in milliseconds and a validator refuses anything that breaks an invariant. That is a more capable system than an LLM choosing alone, and a cheaper one.

You already got the hard part right

The payload your backend sends to routemate is entirely structured — miler_id, distance_km, rating, on_time_rate_30d, hub_load, hub_capacity, plus hour, peak flag and zone. No customer-supplied free text goes into it. That closed payload is a security property, and §06 explains why keeping it is the single most important rule as you expand.

01

Where GenAI actually earns its place

Ten capabilities, ranked by what they return for what they cost. Note how few of them are "the AI decides" and how many are "the AI turns something unstructured into something the system can already handle."

#CapabilityLegWhat the model doesValueEffort
1Address resolutionlastEmbeds messy Indian address text, clusters duplicates, snaps to confirmed DeliveryProof coordinates. Builds you a private geocoder from your own delivered history.highestlow
2Decision explanationallTurns solver output into a sentence for the dispatcher and a line for the rider. Drives adoption more than accuracy does.highlow
3Unstructured intakefirstA client's WhatsApp message, email or PDF manifest becomes structured bookings. You support Excel; clients don't always send Excel.highmedium
4NL constraint compilerall"Don't give Ravi the mall runs" becomes a typed solver constraint a human approves once and the solver honours forever.highmedium
5Exception agentallRider absent, vehicle down, forty parcels stranded. Tool-calling over a bounded action space, proposing a recovery plan.highmedium
6Damage detection at inboundmidVision on the inbound scan photo auto-populates Condition, which is a manual dropdown today and therefore mostly says Good.mediumlow
7POD verificationlastVision checks the delivery photo actually shows a parcel at a door, and that the signature isn't blank. Catches soft fraud.mediumlow
8Vernacular voicefirst, lastRider status by speech in Tamil, Telugu, Kannada or Hindi; an outbound call confirming the consignee will be home. Directly attacks first-attempt failure.highhigh
9Operator copilotallExtends the existing console assistant from read-only queries to proposing actions with confirmation.mediummedium
10Scenario simulationallReplays historical days against a proposed policy before it touches production. Your safety net for everything above.mediummedium

Four hubs across Coimbatore, Hyderabad, Bengaluru and Chennai means four language regions. Item 8 is not a nicety in that footprint — a rider who can speak a status update is a rider who actually files one.

02

Auto-assign, properly

This is the thing you asked about, drawn end to end. The deterministic spine runs down the left and never waits on a model. The three GenAI touch points sit beside it, feeding in out of band.

DETERMINISTIC SPINE · never waits on a model GENERATIVE · out of band booking enqueued · NATS candidate generation Redis GEOSEARCH · adaptive k widen radius until k found constraint set assembled capacity · slots · skills · tenant scope + approved rules from the compiler VRP solver assign + sequence, jointly Valhalla matrix · milliseconds VALIDATOR — hard gate miler_id ∈ candidate set? tenant scope intact? capacity respected? reject → greedy nearest, always commitAssignment one transaction · Idempotency-Key booking + assignment + miler status AgentDecision written Context · Decision · Reasoning Outcome ← filled later, see §05 ① NL constraint compiler dispatcher types: "Ravi shouldn't take mall runs" → typed constraint → human approves once offline · stored · reused every solve ② Exception agent invoked only when the solve is infeasible or an operator escalates tools: reassign · split · defer · widen · page human proposals go back through the validator infeasible proposal — never a direct write ③ Explainer solver output → a sentence for the dispatcher, a line for the rider, in their language async · after commit · failure is cosmetic critical path budget: candidate generation + solve + validate + commit, with no network call to a model in it
Figure 1 — auto-assign with generative AI in its right placesCompare with today: a model on the critical path with a 5-second cap, choosing one rider for one parcel. Here the model never blocks a booking, and the thing it produces — a constraint, a recovery plan, an explanation — is something no solver could have produced instead.

Why the constraint compiler is the sleeper

Dispatchers hold dozens of rules in their heads that never reach the system: which rider is slow at a particular gate, which client insists on the same face, which lane floods in monsoon. Today the only way to encode those is a developer ticket, so they stay in the dispatcher's head and the optimiser looks stupid. A model that turns a typed sentence into a reviewable constraint — which a human approves before it ever binds — converts tacit knowledge into solver input at conversational speed. That is a capability the solver cannot have on its own, and it is where generative AI is doing something genuinely irreplaceable.

// what the compiler emits — reviewed by a human, then stored
{
  "type": "exclude_miler_from_location_class",
  "miler_id": 412,
  "location_class": "mall_service_gate",
  "reason": "slow at gate handover",
  "author": "dispatcher:coimbatore:7",
  "expires_at": null
}
03

Wiring it into what you have

You do not need new infrastructure. Every piece below already exists in the estate; the work is contracts and placement.

PieceAlready thereWhat changes
Async invocationNATS JetStream, the booking worker, a 60 s budgetModel calls become subscribers, not inline HTTP. Nothing in a request path waits on one.
Fallback disciplineAI_LAYER_FALLBACK and the 5 s cap in selectMilerWithAIKeep the pattern exactly; apply it to every new capability. Best-effort or it does not ship.
AuditAgentDecision with a DecisionType discriminatorOne row per model call, whatever the capability. The discriminator was clearly built for this.
IdempotencyRedis-backed Idempotency-Key on payment, pickup and deliverExtend to every agent-initiated write. An agent that retries must not double-act.
Vector storepgvector, in the database alreadyAddress embeddings land here. No new dependency.
Road distanceValhalla behind routes.workolik.comBecomes the solver's cost matrix rather than a post-hoc sequencer.
Object storagepresigned uploads for signatures and photosThe same URLs feed the vision capabilities. Images never transit your API.
Console surfacethe existing deterministic assistant and intentsGains an action tier behind a confirmation step.

One deliberate constraint: the agent layer never writes to the database directly. It calls the same internal endpoints a human operator would, so tenant scoping, idempotency, validation and audit apply identically whether the caller is a person or a model. If an agent can reach a table your dispatcher cannot, you have built two systems and will secure only one.

04

The trust ladder

Every capability climbs the same four rungs, and each rung has a numeric gate. Nothing goes to the next rung on a good feeling.

RUNG 01 Shadow runs on every booking, logs, changes nothing risk: zero RUNG 02 Advisory recommendation shown in the console; a dispatcher clicks to apply it risk: a wasted click RUNG 03 Auto with veto the system acts; a dispatcher can undo inside a window before the rider is notified risk: bounded & reversible RUNG 04 Autonomous the system acts; humans handle escalations only risk: managed by §06 gate: 2 000 logged + agreement measured gate: acceptance > 80% · 0 violations gate: veto < 5% · outcome parity Any capability can sit on a different rung. Address resolution may be autonomous while assignment is still advisory — that is the point of separating them.
Figure 2 — four rungs, three numeric gatesShadow mode is free and tells you more than any amount of design discussion. Run every new capability there first — including the ones you are confident about, because that is where the surprises are.
05

The loop you already built

This is the sharpest finding in the codebase, and it makes everything above cheaper.

assignment made selectMilerWithAI AgentDecision row ✓ Context jsonb ✓ Decision jsonb ✓ Reasoning text ✗ Outcome always null PATCH /internal/agent-decisions/:id/outcome routed · implemented · working called by nothing parcel resolves deliver · skip · RTO the one wire that is missing Column, index, endpoint and route all exist. One call from the resolution handlers turns every past assignment into a labelled training example — and gives you the replay corpus every gate in §04 is measured against.
Figure 3 — 90% built, 0% connectedAgentDecision.Outcome, OutcomeRecordedAt, the controller and the route are all in place. Nothing writes to them. This is the cheapest high-value change available in the whole system.
Do this in week one

Call UpdateDecisionOutcome from the deliver, skip and RTO handlers with a simple verdict — delivered first attempt, delivered late, failed, reassigned. From that day forward every dispatch decision is scored, you can replay history against any new policy offline, and every claim anyone makes about the AI dispatcher becomes checkable. Until then, no gate in §04 can be measured at all.

06

Guardrails

07

Ninety days

PhaseShipRungProves
Wk 1–2Wire UpdateDecisionOutcome from deliver / skip / RTO. Build the replay harness over AgentDecision.—Every later claim becomes measurable.
Wk 3–4Real ETA: Valhalla road time × time-of-day factor + learned dwell, replacing calculateETA.directNo model needed; immediate accuracy win on all three legs.
Wk 3–6Address resolution on pgvector. Embed delivered addresses, cluster, snap to DeliveryProof coordinates.shadow → advisoryHighest-leverage capability, lowest risk.
Wk 5–8VRP solver behind a flag: batched joint assign + sequence, running in shadow beside today's path.shadowHead-to-head on the replay corpus against live decisions.
Wk 7–9Explainer on solver output; dispatcher and rider-facing, in-language.advisoryAdoption. Riders follow routes they understand.
Wk 9–11NL constraint compiler with a human approval queue.advisoryTacit dispatcher knowledge starts reaching the solver.
Wk 10–12Exception agent with a bounded tool set, proposals only.advisoryThe escalation queue starts shrinking.
Wk 12Promote whichever capabilities have cleared their gates. Not the ones that haven't.→ auto with vetoThe ladder does the deciding, not the roadmap.
What is deliberately not in the first ninety days

Vision on inbound condition and POD, vernacular voice, and the operator copilot's action tier are all worth building — and all of them are easier once the outcome loop, the replay harness and the validator exist. Building those three foundations first is what makes the rest take weeks instead of quarters.