AssignCustomerMiler and AssignCRMMiler retried five times, two minutes
apart, using time.Sleep inside a bare goroutine — about ten minutes of
state held only in one pod's memory. Any restart dropped every retry in
flight, and nothing recorded it: the booking just stayed unassigned
forever with no failure event, because publishAssignmentFailed only runs
at the end of a loop that no longer existed. Deploying during a quiet
patch was enough to lose bookings this way, and it happened during
today's rollout.
Retries now run on the ASSIGNMENTS stream. Each entry point publishes
one booking.assignment_requested message; a durable consumer performs a
single attempt per delivery and NAKs with retryDelay when no miler is
available, so JetStream owns both the waiting and the delivery count.
A pod dying mid-wait costs nothing — the message is still on the server
and another replica takes it.
Behaviour is deliberately unchanged from the caller's side: same five
attempts, same two-minute spacing, same publishAssignmentFailed handoff
to the DispatchAgent. The failure event is fired explicitly on the last
delivery, since JetStream stops redelivering at MaxDeliver and would
otherwise let the booking fail silently again.
runInline keeps the old loop as a fallback for when JetStream is down.
Assignment is how a booking reaches a rider, so it must not become
dependent on the event bus: an outage should cost durability, which is
what we had before, not stop bookings being assigned at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Half of this binary's js.Publish calls were bound to no stream at all.
The streams were declared only by an external Python script on another
machine (Birock/doormile-bookings/setup_jetstream.py) and had drifted
from the code: booking.cancelled, booking.outcome and
booking.assignment_failed had no stream, and CHAT declared the literal
"chat.room.closed" while chat.go publishes "chat.room.closed.<id>",
which it does not match. Every publish site is best-effort
(`if db.Js != nil` + warn-log), so those events were failing and being
dropped silently — every cancellation, delivery outcome and assignment
failure since the streams were created.
db.EnsureStreams now declares the streams at startup from a map that
sits next to the code that publishes, so the contract cannot drift
again. It only ever adds: existing streams keep their storage type,
retention, limits and every subject they already have. Nothing is
deleted. Losing the create race against a sibling replica is expected
and reconciles rather than erroring.
Alongside that, /internal/miler/* ingests rider telemetry still arriving
over the jupiter NATS chain. The forwarding worker holds no rider JWT —
the rider app is still jupiter-shaped — so LegacyMilerIdentity resolves
an identity from a header into c.Locals("userid") behind the existing
X-Internal-Key guard. That lets the routes reuse the miler handlers
unchanged instead of growing a parallel set that would drift.
Identity comes from a header, never the body: the telemetry handlers
overwrite a body-supplied userid precisely so one rider cannot write
another's GPS trail, and reading it from the body here would reopen that
from behind the internal key. MilerProfile.Legacyuserid (nullable,
indexed) maps a jupiter userid to a Doormile one.
Only fire-and-forget telemetry is exposed. Transactional actions stay
synchronous — a rider needs a real answer from pickup-complete, which a
queue in front of it cannot give.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>