AssignCustomerMiler and AssignCRMMiler retried five times, two minutes apart, using time.Sleep inside a bare goroutine — about ten minutes of state held only in one pod's memory. Any restart dropped every retry in flight, and nothing recorded it: the booking just stayed unassigned forever with no failure event, because publishAssignmentFailed only runs at the end of a loop that no longer existed. Deploying during a quiet patch was enough to lose bookings this way, and it happened during today's rollout. Retries now run on the ASSIGNMENTS stream. Each entry point publishes one booking.assignment_requested message; a durable consumer performs a single attempt per delivery and NAKs with retryDelay when no miler is available, so JetStream owns both the waiting and the delivery count. A pod dying mid-wait costs nothing — the message is still on the server and another replica takes it. Behaviour is deliberately unchanged from the caller's side: same five attempts, same two-minute spacing, same publishAssignmentFailed handoff to the DispatchAgent. The failure event is fired explicitly on the last delivery, since JetStream stops redelivering at MaxDeliver and would otherwise let the booking fail silently again. runInline keeps the old loop as a fallback for when JetStream is down. Assignment is how a booking reaches a rider, so it must not become dependent on the event bus: an outage should cost durability, which is what we had before, not stop bookings being assigned at all. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
6.7 KiB
6.7 KiB