Assignment decided who carried a booking but never what order to run
several of them in -- the one capability jupiter had that Doormile did
not. It turns out we already own the solver: routes.workolik.com is a
live in-house Route Optimization API backed by Valhalla road-network
routing, not the paid third party I had assumed. This is a client for
it, not a solver.
internal/routing posts a rider's active stops and writes back step,
previouskms, cumulativekms and ETA onto BookingAssignment. HubBatchAssign
calls it after committing a batch, which is exactly the case it exists
for: a rider used to walk away with several bookings and no order to run
them in. GetMilerAssignments now returns sequenced stops in step order,
falling back to newest-first for anything unsequenced.
Contract discovered by probing the live service -- the OpenAPI schema
types the body as a bare object array, so the field names are not
documented anywhere. They are pickuplat/pickuplong/deliverylat/
deliverylong, NOT pickuplatitude/deliverylatitude. Sending the wrong
names does not fail: it returns HTTP 200 with every coordinate defaulted
to "0.0", no reordering and all distances zero. That trap is recorded in
a comment so the next person does not lose an afternoon to it.
Numeric fields come back inconsistently typed -- previouskms as a number,
actualkms and eta as strings, some decimal -- so they are decoded loosely
and coerced, with tests pinning the coercion. Steps for deliveryids we
did not send are discarded rather than written, so an echoed or stale id
cannot reorder another rider's work.
Sequencing is best-effort throughout and runs after assignments commit.
The optimizer is a separate service over the network; it being down must
leave bookings assigned but unordered, never undo the batch. Step 0 means
"not sequenced", not "first".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
AssignCustomerMiler and AssignCRMMiler retried five times, two minutes
apart, using time.Sleep inside a bare goroutine — about ten minutes of
state held only in one pod's memory. Any restart dropped every retry in
flight, and nothing recorded it: the booking just stayed unassigned
forever with no failure event, because publishAssignmentFailed only runs
at the end of a loop that no longer existed. Deploying during a quiet
patch was enough to lose bookings this way, and it happened during
today's rollout.
Retries now run on the ASSIGNMENTS stream. Each entry point publishes
one booking.assignment_requested message; a durable consumer performs a
single attempt per delivery and NAKs with retryDelay when no miler is
available, so JetStream owns both the waiting and the delivery count.
A pod dying mid-wait costs nothing — the message is still on the server
and another replica takes it.
Behaviour is deliberately unchanged from the caller's side: same five
attempts, same two-minute spacing, same publishAssignmentFailed handoff
to the DispatchAgent. The failure event is fired explicitly on the last
delivery, since JetStream stops redelivering at MaxDeliver and would
otherwise let the booking fail silently again.
runInline keeps the old loop as a fallback for when JetStream is down.
Assignment is how a booking reaches a rider, so it must not become
dependent on the event bus: an outage should cost durability, which is
what we had before, not stop bookings being assigned at all.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The hub console has its own subdomain; without it browsers block every
cross-origin call from the hub UI.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Hardening pass over the API surface. No route's auth requirements change.
Resilience:
- Add recover middleware. There was none, so an unhandled panic in any
handler propagated out of the process instead of becoming a 500.
- Add a centralized ErrorHandler so errors and recovered panics return the
same {success,message} envelope as the utils helpers, not Fiber's default
plain-text body. 5xx responses are logged with method and path.
Rate limiting:
- Global 300/min per IP as an abuse backstop, exempting health/readiness
probes and websocket upgrades.
- 10/min shared across every credential endpoint (customer/miler/admin/hub
login, verify-pin, reset-pin, email OTP). PINs are 4 digits, so the whole
keyspace was previously walkable in seconds. One shared limiter instance
means rotating between endpoints doesn't reset the budget.
- Add TRUSTED_PROXIES config. Limits key on c.IP(), which behind a TLS
terminator is the proxy, collapsing every client into one bucket. When set,
X-Forwarded-For is honoured only from those proxies so the header can't be
spoofed to dodge the limit. Logs a warning when unset.
Transactions:
- Check the error on all 51 previously-unchecked tx.Save/Create/Delete/
Model(...).Update/Commit calls across 6 controllers. A failed write inside
a transaction was silently ignored and the request still reported success;
an unchecked Commit could fail with the caller told everything worked.
Each site now rolls back and returns a specific message.
Pagination:
- Add utils.ParsePage/Paginated, reusing the pageno/pagesize convention
GetAdminBookings already established. Default 500, hard cap 1000.
- Apply to the previously unbounded consignments, tripsheets, exceptions,
app-users and clients endpoints. Defaults are high so existing consoles
that don't paginate keep working; the cap only stops a growing table from
being loaded wholesale. total is now a real COUNT, not len(data).
- GetClients also loaded the entire auth table to join in memory; it now
fetches only the current page's rows.
Tests (first in the repo):
- Extract the hyperlocal pincode rule out of BookingPickupComplete into
isHyperlocal so it is testable, covering the short/empty pincode fallback.
- Cover calculateVolumetricWeight and the ParsePage clamping rules.
Repo hygiene:
- Tag scratch/*.go with //go:build ignore. Each declared its own main(), so
`go build ./...` failed on redeclaration; it now passes repo-wide.
- Untrack scratch/node_modules (216 files) and ignore node_modules, test
artifacts, and the `doormile` binary `go build .` emits.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Bugs found during live Coimbatore testing against production:
- GetMilerAssignments returned every assignment ever made to a miler with
no status filter, so weeks-old Rejected assignments still showed up as
actionable in the app. Now filters to Assigned/Accepted only.
- AcceptMilerAssignment overwrote assignmentstatus unconditionally, letting
a stale Rejected assignment be silently reactivated (and pushing its
booking back to Pickup_Scheduled). Now 400s unless currently Assigned,
reporting the actual status.
- RejectMilerAssignment read `reason` from the query string instead of the
JSON body, contradicting the API contract and every sibling endpoint.
- BookingPickupComplete resolved the origin hub via db.First(&hub) with no
Where clause — i.e. the lowest hub ID in the table, unrelated to the
booking or miler. Now uses the miler's own MilerProfile.Hubid, falling
back to the old behaviour with a warning only when unassigned.
- No code path ever set a consignment to Out_for_Delivery, making
MilerDeliverConsignment unreachable. BookingPickupComplete now goes
straight to Out_for_Delivery when pickup and delivery pincodes share a
3-digit postal-area prefix (same-miler hyperlocal), reusing the
hubPincodePrefix convention. Cross-hub still lands at Inwarded_at_Hub.
Also allow https://app.doormile.com in CORS, and drop a stray Windows-path
log file that was committed by accident.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>