Files
AI_engine/docs/handoff-assignment-failed-event.md

6.2 KiB

Handoff: publish a booking.assignment_failed NATS event

To: whoever owns the Go backend (assignment package) From: logistics-ai (Python agents) side Why: the DispatchAgent's assignment-failure decision path is built and tested but inert — it consumes an event the backend does not currently publish.


TL;DR

When miler assignment fails, the backend currently handles it synchronously (an escalate flag in the decision response) and publishes nothing to NATS. Confirmed by inspecting the running binary: booking.assigned is published, but there is no booking.assignment_failed (or any failure) subject anywhere in the binary.

To activate the DispatchAgent coverage-gap / escalation logic, the backend needs to publish one event on assignment failure and add its subject to the existing ASSIGNMENTS stream. No further Python changes are required — the consumer, its decision logic, its eval set, and its tests are already in place.


1. The event to publish

  • Subject: booking.assignment_failed

  • Transport: JetStream (the code already uses PublishAsync — reuse it; publishing must not block the request path).

  • When: in the failure branch of assignment — i.e. where the flow currently decides it cannot assign a miler. Based on the binary's symbols, the natural call sites are in the assignment package:

    • collectEligibleCandidates returns empty (no eligible milers), and/or
    • callDecisionEngine / selectMilerWithAI returns escalate, and/or
    • AutoAssignResult / tryAssign / tryCustomerAssign resolves to "not assigned".

    Publish alongside the existing success publish (publishAssignment / publishCustomerAssignment) so success and failure are emitted from symmetric places.

Payload (JSON)

The consumer reads these fields (extra fields are ignored, so it is forward-compatible):

Field Type Required Notes
booking_id string or int yes The booking that could not be assigned.
zone_id string recommended Used for the per-zone daily failure counter and coverage-gap detection. Defaults to "unknown" if omitted.
lat float recommended Pickup latitude. Enables the nearest-rider coverage sweep (GEORADIUS).
lon float recommended Pickup longitude.
reason string optional e.g. "no_candidates", "ai_escalate", "provider_unavailable". Not required today; useful context for the decision and for future logic.

Without lat/lon the decision still runs but degrades to the safe side (it can't measure coverage, so it leans toward "monitor"/"escalate"). Include them whenever available.

Example

{
  "booking_id": 24,
  "zone_id": "hyderabad",
  "lat": 17.4486,
  "lon": 78.3908,
  "reason": "no_candidates"
}

This mirrors the shape of the existing booking.assigned payload (booking_id, miler_id, hub_id, confidence, reasoning), just for the failure case.


2. Add the subject to the ASSIGNMENTS stream

The stream already exists (Go-owned):

STREAM: ASSIGNMENTS | subjects: ['booking.assigned', 'booking.reassigned']

Add booking.assignment_failed so JetStream captures it. Do this before/at the same time as the first publish — messages published to a subject no stream covers are dropped.

Wherever the stream is provisioned in code (the same place that created ASSIGNMENTS), update its subject list to:

booking.assigned, booking.reassigned, booking.assignment_failed

Or, as a one-off from the NATS host:

nats stream edit ASSIGNMENTS --subjects="booking.assigned,booking.reassigned,booking.assignment_failed"

(Prefer updating the provisioning code so it stays correct on redeploys / fresh environments.)


3. Consumer side — already done, nothing to change

The DispatchAgent (agents/dispatch_agent.py) already:

  • auto-discovers the stream for booking.assignment_failed (find_stream_name_by_subject), so once the subject is on ASSIGNMENTS it binds automatically on next restart — no config needed. (An optional DISPATCH_STREAM env var can pin the stream name if you ever want to bypass discovery.)
  • binds a durable push consumer named dispatch-booking-assignment-failed.
  • on each event: gathers coverage context (nearest-rider sweep + per-zone daily failure counter), asks the model to choose monitor / notify_customer / ops_alert / escalate, and acts (customer notification is gated behind DISPATCH_AGENT_AUTONOMOUS; default off = proposes internally).

Delivery is at-least-once (JetStream), and the handler is safe to re-run, but keep publishes reasonable — one event per failed assignment attempt.


4. How to verify end to end

  1. Deploy the DispatchAgent stream fix (the binding change) and the backend change together.
  2. Confirm the subject is on the stream:
    nats stream info ASSIGNMENTS      # subjects should include booking.assignment_failed
    
  3. Trigger a real (or forced) assignment failure, or publish a test event:
    nats pub booking.assignment_failed '{"booking_id":9999,"zone_id":"testzone","lat":17.44,"lon":78.39,"reason":"no_candidates"}'
    
  4. In the DispatchAgent logs, you should see:
    [DISPATCH] booking.assignment_failed — booking=9999 zone=testzone
    [DISPATCH] booking=9999 context facts: {...}
    [DISPATCH] booking=9999 zone=testzone decision=<action> confidence=<x.xx> reasoning='...'
    
  5. nats consumer info ASSIGNMENTS dispatch-booking-assignment-failed should show delivered/ack'd counts incrementing.

Once you see that decision line on a real failure, the assignment-failure path is live — and it can be graduated to autonomy the same way as the stall path (turn on DISPATCH_AGENT_AUTONOMOUS only after the eval passes on real cases).


Reference — current NATS topology (for context)

ASSIGNMENTS    booking.assigned, booking.reassigned          (add: booking.assignment_failed)
BOOKINGS       api.v1.bookings.create/update/cancel
CHAT           chat.room.created/closed
NOTIFICATIONS  notification.send
STATUS         booking.status.updated
TRACKING       miler.location.updated, miler.stalled          (ExceptionAgent consumes)
logistics      logistics.>                                    (internal agent-to-agent bus)