fix: assignment retries no longer die with the pod

AssignCustomerMiler and AssignCRMMiler retried five times, two minutes
apart, using time.Sleep inside a bare goroutine — about ten minutes of
state held only in one pod's memory. Any restart dropped every retry in
flight, and nothing recorded it: the booking just stayed unassigned
forever with no failure event, because publishAssignmentFailed only runs
at the end of a loop that no longer existed. Deploying during a quiet
patch was enough to lose bookings this way, and it happened during
today's rollout.

Retries now run on the ASSIGNMENTS stream. Each entry point publishes
one booking.assignment_requested message; a durable consumer performs a
single attempt per delivery and NAKs with retryDelay when no miler is
available, so JetStream owns both the waiting and the delivery count.
A pod dying mid-wait costs nothing — the message is still on the server
and another replica takes it.

Behaviour is deliberately unchanged from the caller's side: same five
attempts, same two-minute spacing, same publishAssignmentFailed handoff
to the DispatchAgent. The failure event is fired explicitly on the last
delivery, since JetStream stops redelivering at MaxDeliver and would
otherwise let the booking fail silently again.

runInline keeps the old loop as a fallback for when JetStream is down.
Assignment is how a booking reaches a rider, so it must not become
dependent on the event bus: an outage should cost durability, which is
what we had before, not stop bookings being assigned at all.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Suriya
2026-08-10 16:29:32 +05:30
parent 41f0013751
commit 2158031191
5 changed files with 248 additions and 71 deletions

View File

@@ -20,10 +20,18 @@ type providerResult struct {
reliability float64
}
// AssignCustomerMiler is the B2C goroutine entry point.
// It selects both a miler (first-mile pickup) and a provider (delivery routing),
// then commits the assignment. Retries up to maxRetries times with retryDelay in between.
// Must be called as a goroutine after tx.Commit() in CreateCustomerBooking.
// AssignCustomerMiler is the B2C entry point. It selects both a miler
// (first-mile pickup) and a provider (delivery routing), then commits the
// assignment. Call it after tx.Commit() in CreateCustomerBooking.
//
// The attempt and its retries now run on the ASSIGNMENTS stream rather than in
// this process — see queue.go for why. Publishing is fast enough that the
// caller's `go` is no longer strictly needed, but it is harmless and the call
// sites are left as they are.
//
// Terminal failure still hands off to the AI layer's DispatchAgent via
// publishAssignmentFailed, which the worker fires once the last delivery is
// exhausted.
func AssignCustomerMiler(bookingID int) {
defer func() {
if r := recover(); r != nil {
@@ -31,38 +39,7 @@ func AssignCustomerMiler(bookingID int) {
}
}()
for attempt := 1; attempt <= maxRetries; attempt++ {
if attempt > 1 {
time.Sleep(retryDelay)
}
utils.Info("B2CAssignment: attempting assignment", "booking_id", bookingID, "attempt", attempt)
done, err := tryCustomerAssign(bookingID)
if err != nil {
utils.Error("B2CAssignment: attempt error", "booking_id", bookingID, "attempt", attempt, "error", err)
continue
}
if done {
return
}
utils.Warn("B2CAssignment: no eligible miler on attempt",
"booking_id", bookingID,
"attempt", attempt,
"remaining", maxRetries-attempt,
)
}
utils.Error("B2CAssignment: NO_MILER_AVAILABLE — all retries exhausted",
"booking_id", bookingID,
"max_retries", maxRetries,
)
// Terminal failure — hand off to the AI layer's DispatchAgent, which owns
// what happens next (coverage sweep, escalation). Reached only after every
// retry is exhausted, so it fires at most once per booking.
publishAssignmentFailed(bookingID, reasonNoMilerAvailable)
enqueue(bookingID, kindCustomer)
}
// tryCustomerAssign performs one full attempt: GEO miler search → provider selection → commit.