Files
doormile_backend/CLAUDE.md

43 KiB
Raw Blame History

Doormile — Full Project Memory

Portable context for a fresh Claude session working on this project on any machine or account. Covers the whole Doormile system, not just one slice of it. Read fully before touching the codebase.

A note on provenance: sections marked [verified this session] were confirmed by directly reading this repo's code. Sections marked [carried forward] come from Suriya's own prior-session notes and have NOT been independently re-verified against source — treat them as reported state, not confirmed state, until checked.


1. What Doormile is

Suriya is building Doormile, a real commercial logistics platform for the Indian market (Coimbatore, Hyderabad, Bengaluru, Chennai) — hub-based parcel courier/delivery, not a prototype. Treat everything as production-grade from the start. Stakeholders: Doormile's own ops team, partner tenants (client logistics/travel companies), hub staff, milers (delivery riders), and end customers.

Suriya's standing preferences (apply without being asked):

  • Direct, honest technical assessments over diplomatic framing. If something isn't done, say so plainly. Don't declare victory early.
  • Production-grade code from the start, not "refactor later."
  • Warn before any consequential server/schema change before making it.
  • Concise answers — don't over-explain.
  • Once a decision is made, don't keep re-asking for confirmation on the same scope — proceed and report back. But do ask when a decision is genuinely ambiguous or touches money/data-correctness in a way that can't be safely guessed.
  • Minimal-effort, highest-leverage fixes over big rewrites, until he explicitly asks for the big rewrite.
  • Works across multiple Claude sessions/machines simultaneously, treating Claude as a co-architect across the full stack — this file exists so any of those sessions can pick up full context.

2. System architecture — the whole stack

[carried forward, substantial infra reported as built and verified in prior sessions]

  • Backend: Go + Fiber, deployed on Kubernetes, at api.doormile.com. 220 registered routes [verified 2026-09-02, exact count] — see §7. This is the primary booking/assignment API and the primary trigger for miler assignment, calling the AI decision layer with a 5-second timeout fallback so a slow AI response never blocks a booking.
  • AI dispatch layer: a Python autonomous agent swarm (8 agents, JARVIS as master orchestrator) on a dedicated server, connected to NATS JetStream. The DispatchAgent operates as a NATS watcher, not the assignment trigger — the Go backend triggers assignment directly (see §4); the agent swarm reasons about it, doesn't gate it.
  • AI decision engine: routemate.workolik.com — runs Claude Sonnet for miler-assignment reasoning, scoring candidates using hub load, on-time rate, and hub capacity. Uses pgvector-backed RAG memory with local embeddings (all-MiniLM-L6-v2, 384-dim) specifically to avoid ongoing per-call API cost. [verified this session]: the Go side of this is models.AgentDecision (agentdecisions table) with a context_embedding vector(1536) column and an ivfflat index (raw SQL in migrations/migrate.go, not GORM AutoMigrate) — note the dimension mismatch worth flagging (1536 in the Go model/migration vs 384 for MiniLM); worth confirming which embedding model is actually populating that column before trusting vector search quality.
  • Event bus: NATS JetStream, 6 persistent streams, dedicated server. [verified this session]: db.Js is the package-level JetStream handle; convention across the codebase is if db.Js != nil { ... } and warn-log on publish failure rather than failing the request — publishing is best-effort, never blocking. Confirmed publish call sites include booking.assigned (from internal/assignment) and NATS publishes inside AssignMilerToBooking (controllers/booking_assignment_service.go).
  • Data stores: PostgreSQL + pgvector 0.7.4, dedicated server, 37 tables under GORM AutoMigrate [verified this session, exact count from migrations/migrate.go]. Redis (go-redis/v9) for high-frequency ephemeral state — GPS pings, live miler status/location (GEO-indexed at milers:locations, used by queryNearbyMilers for assignment radius search), booking cache. Durable business state always lands in Postgres; Redis is never the system of record for anything that needs to survive a flush.
  • Ingress/infra: Kubernetes, Traefik, nginx.

3. The client-facing surfaces

Five distinct front doors into this system. Only the backend API surface for each has been directly inspected this session (via routes.go); the actual client codebases (Flutter, React) have not been opened in this session.

  1. Customer app (B2C) — Flutter, being replaced: doormile_customer_app (PIN auth, single-destination bookings) is retired in favour of doormile_cx. Backend surface rebuilt 2026-09-05 to the Customer App v1 contract: 28 customer routes (customer/customerAuth groups) — OTP auth with refresh, serviceability/slots/limits, place proxy, fare estimate, multi-destination pickups, per-order tracking, push devices. See §8.5. Auth is a 4-digit OTP to phone or email, NOT Firebase and NOT the miler PIN flow; the SMS gateway is still unplugged, which is the same wall §9 describes.
  2. Miler app — Flutter, for delivery riders. Backend surface: 38 routes [verified this session] (miler/milerAuth groups) — duty start/stop, GPS pings, assignment accept/reject/cancel, delivery complete/skip, break logs, support tickets. This session added ResetMilerPin, MilerCancelAssignment, MilerSkipDelivery to close gaps found against the old Nearle rider app (§8).
  3. Admin console — React, converted from the old NearlExpress /doormile_express_console codebase. Backend surface: 89 routes [verified this session] (admin/adminAuth groups) — bookings, partner management, reports, pricing, express (formerly "CRM") bookings. This session added partner CRUD, bulk booking create/cancel, reports, password change, miler notify (§8).
  4. Hub console — React, separate from the admin console. Backend surface: 31 routes [verified this session] (hub/hubAuth groups) — per-hub booking queues, tripsheet building, manual/batch assignment, hub messaging. Auth is separate from admin (middlewares.HubStaffAuth, sets c.Locals("hubid")). This session added tenant-scoping (closed a real cross-tenant leak) and HubBatchAssign (§8).
  5. CRM (field sales) — a genuinely separate feature from "CRM bookings" (see §8 naming note). POST/GET /crm/clients, GET/PUT/DELETE /crm/clients/:id — 5 routes, all open/no-auth by design: field reps register leads from the Flutter side, the web console reads without a separate CRM login. Models: DoormileClient, DoormileAuth. Not deeply audited — only the route registration was read this session.

Total: 200 routes = miler 38 + admin 89 + hub 31 + customer 19 + crm 5 + internal 5 + redisUsers 4 + bookingCache 3 + 2 top-level pricing routes + websocket routes. [verified this session]


4. AI dispatch & assignment — how a booking actually gets a miler

[verified this session] — read directly from internal/assignment/crm_assignment.go and booking_assignment_service.go.

  • Two entry points, same core: AssignCustomerMiler (B2C bookings) and AssignCRMMiler (express/console-originated bookings — name predates this session's "CRM"→"express" rename, deliberately left unrenamed since it's a working assignment engine, not just a label). Both are fire-and-forget goroutines called after transaction commit, retrying up to 5 times, 2 minutes apart (~10 minutes total), before giving up and publishing a failure event (publishAssignmentFailed, reasonNoMilerAvailable) to NATS so the failure isn't silently invisible.
  • TryAssignOnce is the synchronous single-attempt variant — used when a hub staff member manually triggers "auto-assign" from the console and needs an immediate answer rather than the ~10-minute retry loop. Returns a rich AutoAssignResult (assigned/escalated, miler name, distance, AI reasoning text, candidates found) for the UI to show.
  • Core flow (tryAssign): Redis GEOSEARCH on milers:locations (10km radius, top 10, sorted nearest-first) → selectMilerWithAI (the actual Claude Sonnet call to routemate.workolik.com, logs an AgentDecision row with reasoning + optional embedding) → commitAssignment (single DB transaction: create BookingAssignment, update PickupBooking.status to BookingMilerAssigned, flip MilerProfile.availabilitystatus) → publish booking.assigned to NATS (best-effort) → push notifications to miler and customer.
  • HubBatchAssign (built this session, controllers/hubController.go) is a separate, simpler path: a greedy nearest-available-rider heuristic (haversine distance, no AI call, no retry loop) for clearing a hub's pending-pickup queue in one batch call, capped per rider (default 5), reusing AssignMilerToBooking for the actual transactional assignment. Explicitly not a replacement for selectMilerWithAI's reasoning, and not a multi-stop route/VRP solver — see the routing gap below.

Logistics/routing — what exists vs what doesn't:

  • What exists: hub network (Hub model, originhubid/currenthubid/ destinationhubid on bookings/consignments for tracking a parcel's hub path), Tripsheet/TripsheetItem for hub-to-hub batch transport, the nearest-rider assignment logic above, HubBatchAssign's greedy queue clearer.
  • What does not exist in Doormile yet: any multi-stop route sequencing / vehicle-routing-problem solver. Nearle used external paid services (routes.workolik.com, and routemate.workolik.com was itself originally a Nearle-side concept) for that; nothing in Doormile replaces true stop-sequencing optimization today. HubBatchAssign only decides who gets which single booking, not what order a rider should run multiple stops in. Flag this clearly if asked whether routing is "done" — it isn't, by design, pending a real conversation about whether it's needed yet.

5. Data model reference (models/*.go)

[verified this session]

  • PickupBooking (pickupbookings) — customer's pickup request, pre-hub. Has Tenantid *int (added this session, see §8). Bookingsource: "Customer_App" (B2C) or "CRM_Console" (console-created — deliberately left as "CRM_Console" even though the outward API/route naming was changed to "express", see §8).
  • BookingParcel, BookingServiceOption, BookingPayment, BookingAssignment (status: Assigned/Accepted/Rejected/Reassigned/ Completed/Cancelled, carries AgentDecisionID *uint64 linking back to the AI reasoning that produced it), BookingVehicleRequirement.
  • Consignment (consignments) — the shipment once picked up. Has its own Tenantid int (not nullable, pre-existing). Attemptcount int (used by MilerSkipDelivery, added this session). ConsignmentHistory (event log), ConsignmentException (Lost/Damaged/Misrouted/Receiver_Refused/ Missing_Contents/Undeliverable).
  • Hub, Vehicle, Tripsheet, TripsheetItem, DeliveryProof. Hub is what the rider app calls a Base — same row, different word (see §12).
  • AppUser (appusers) — shared login table for staff/miler/admin roles (Roleid: 1 admin, 3 manager, 4 rep/exec, 5 miler, 6 hub staff via a separate HubStaffAccount table). MilerProfile — actual rider profile (availability status, current lat/lng, rating). MilerDutyLog, MilerBreakLog, MilerSupportTicket.
  • AppCustomer/AppCustomerLocation — B2C app customers (separate from legacy Customer/CustomerLocation kept for backward compat — flagged as a future duplicate-data risk, not resolved).
  • HubStaffAccount — Tenantid *int: nil = Doormile staff (sees everything), set = partner-tenant staff (own tenant's data only). Load-bearing distinction — see the tenant-leak fix in §8.
  • PartnerInfo (partnerinfo) — fleet/vehicle-supplying partner companies. Distinct concept from Tenant (tenants — client companies Doormile delivers for). Confusingly close names, flagged, not acted on.
  • DoormileClient/DoormileAuth — the real, separate CRM feature (§3.5), not to be confused with "CRM bookings" (renamed to "express bookings" this session specifically to end that naming collision).
  • Pricing, DoormilePricing, CompetitorBranch, CarrierPricing — pricing engine + competitive intel.
  • AgentDecision — AI dispatch-reasoning log (§4), pgvector context_embedding column.
  • Redis-only ephemeral structs (models/redis.go): MilerLog, MilerStatus, ConsignmentLog, CachedUser — never durable state.

6. Current environment / seeded data state

[carried forward, not re-verified this session]

  • Database seeded with 16 hubs (4 per city across the 4 target cities), 10+ milers, 10 tenants (4 Doormile-ops-owned + 6 named partner tenants), 5 hub staff accounts, 55+ pricing rules.
  • Credentials on file from prior sessions (values not repeated here — confirm current values rather than assuming): admin login at suriya@doormile.com; hub staff accounts pattern hub.[city]@doormile.in plus partner variants; miler test phone numbers + PINs.
  • Auth: not Firebase. Customer login is a 4-digit OTP to phone or email, issued by issueCxOtp and stored in Redis (§8.5). CX_STAGING_OTP makes it a fixed code outside production, which is what lets it be scripted — the earlier "cannot be bypassed from the command line" note (§9) predates that and predates the /customer/* rebuild.

7. Codebase conventions (read before writing any Go here)

[verified this session]

  • Response helpers (utils package): utils.OK(c, data), utils.Created(c, data), utils.Message(c, "text"), utils.List(c, slice, total), utils.Paginated(c, slice, total, page), utils.BadRequest/NotFound/Internal/Unauthorized/Forbidden/Conflict(c, msg). Always use these, never hand-roll c.JSON.
  • DB: db.DB is the package-level *gorm.DB. db.Rdb is Redis (*redis.Client, go-redis/v9). db.Js is NATS JetStream (nil-check before publishing, warn-log on failure, never fail the request over it). db.Ctx for Redis calls needing a context.
  • Auth: c.Locals("userid") (int), c.Locals("tenantid") (int, only set via middlewares.AuthMiddleware — the requesting user's own tenant, not necessarily the resource's tenant; conflating these two was exactly the bug fixed this session, §8). Hub staff auth is separate (middlewares.HubStaffAuth), sets c.Locals("hubid") and c.Locals("userid") = hubstaffaccountid.
  • DTOs: dto/*.go (admin.go, auth.go, booking.go, client.go), one struct per request shape.
  • Constants (constants/constants.go): every status enum lives here as typed string constants. Always use these, never inline string literals.
  • Soft delete: only some models have Deletedat *time.Time (Hub, Vehicle, Consignment, Tripsheet, TripsheetItem, ConsignmentException, DoormilePricing, CarrierPricing). Others (PartnerInfo, Tenant) have no soft-delete column — hard Delete. Check the struct before assuming either way.
  • Assignment logic reuse: AssignMilerToBooking(bookingID, milerUserID int, assignedByUserID *int) in controllers/booking_assignment_service.go is the one transactional path for "assign this miler to this booking." Reuse it, don't reimplement — both HubAssignMiler/AdminAssignMiler and HubBatchAssign call it. The AI-driven commitAssignment in internal/assignment is a separate transactional writer used by the auto-assignment retry path — don't conflate the two; know which one a given call site needs.
  • Distance calc: haversineKM(lat1, lon1, lat2, lon2) defined once in hubController.go, used package-wide. Don't redefine it.
  • Hub tenant scoping: scopeBookingsToOwnTenant(c, query) in hubController.go (added this session) — restricts a bookings query to the requesting hub staff's own tenant if partner-scoped, no-ops for Doormile staff. Use on any new hub-console bookings query.
  • Go toolchain was unavailable in the sandbox that wrote this session's changes. Every change was verified by hand (field/column names cross-checked against real structs, brace-balance via grep -o "{" | wc -l vs }) at write time. [verified separately, on Suriya's own machine]: go build ./... and go vet ./... were both run afterward and passed clean (only pulled two missing indirect modules, tinylib/msgp and philhofer/fwd). This was the first actual compiler verification of this code — commit c272a33, pushed to origin/main.

8. The Jupiter → Doormile migration (this session's scoped work)

This section is what one specific session did: closing API/schema gaps between the legacy Nearle ("jupiter") system and the new Doormile backend, for the rider app and console specifically (hub console gap-closure was also done incidentally while fixing the tenant leak, but wasn't the primary target).

8.1 Why jupiter/Nearle was being replaced

Legacy Go+Fiber backend (backend_jupiter, module nearle), Postgres, Redis, base URL jupiter.nearle.app (+ a separate write path queue.workolik.com with TLS verification disabled and a hardcoded IP pin). Consumed by a Flutter rider app and a React admin console (doormile_express_console — literal ancestor of the current Doormile admin console). Found to be structurally broken on live-DB analysis:

  • riderlogs: 1.17M rows, 643MB, zero indexes, 68M lifetime UPDATEs vs 762K inserts, caused by an unindexed UPDATE ... WHERE userid=? rewriting ~97K rows per call — the real root cause of read timeouts previously blamed on Redis.
  • Every table had only its primary key indexed; orders/deliveries never autovacuumed.
  • getdeliveries returned every row 21× (unconstrained LEFT JOIN tenantpricing, DISTINCT over 87 columns that didn't dedupe anything).
  • createdeliveries had a quadratic insert bug (slice declared outside a loop kept accumulating) — confirmed live: 66,446 deliveries → 132,826 deliveryqueues rows (~2×) with duplicate deliveryids.
  • v2 endpoints wrote only to Redis, invisible to v1/v3 Postgres reads — genuine split-brain, with a Redis INCR ID space independent of the Postgres sequence (collision risk).
  • orders had 75 columns (~20 never populated once across 137K rows). deliveries had 92 columns, six lat/lng pairs for 3 real points, status spread across 6 separate text+timestamp columns instead of an events table.
  • PUT /deliveries/updatedelivery was overloaded for 11 different real actions (8 rider status transitions + 3 unrelated console actions), distinguished only by which JSON fields happened to be non-empty.
  • The "exhaustive" API docs undercounted real usage by ~8 endpoints (including the login endpoint itself and a whole /v1/substitutions CRUD feature), found only by cross-checking against actual console source.

Doormile's new schema was already solving most of this structurally before this session (Redis-only telemetry, real event-log tables, 37 normalized tables with real FKs and typed columns, parcel/courier domain model instead of retail/food-delivery shaped).

8.2 What was built this session (14 new endpoints)

Method Path Handler
POST /miler/reset-pin ResetMilerPin
POST /miler/bookings/:bookingid/cancel MilerCancelAssignment
POST /miler/consignments/:id/skip MilerSkipDelivery
GET /admin/reports GetAdminReports
PUT /admin/profile/password AdminChangePassword
GET /admin/partners GetPartners
POST /admin/partners CreatePartner
GET /admin/partners/:id GetPartnerDetails
PUT /admin/partners/:id UpdatePartner
DELETE /admin/partners/:id DeletePartner
POST /admin/milers/:id/notify AdminNotifyMiler
POST /admin/expressbooking/bulk AdminBulkCreateBookings
POST /admin/bookings/bulk-cancel AdminBulkCancelBookings
POST /hub/bookings/batch-assign HubBatchAssign

Plus two real bugs found and fixed incidentally while doing tenant- separation work (not requested, found along the way):

  1. Data mis-attribution: BookingPickupComplete set the resulting Consignment.Tenantid from the completing miler's own tenant rather than the booking's actual tenant — silently mis-attributed shipments for any miler carrying parcels across tenants. Fixed to use booking.Tenantid when set.
  2. Cross-tenant data leak: GetHubUnassignedBookings/ GetHubBookingsRange had no tenant scoping — a partner tenant's hub staff could see every other tenant's bookings at the same hub. Fixed via scopeBookingsToOwnTenant.

Naming cleanup: CreateCRMBooking→CreateExpressBooking, /admin/crmbooking→/admin/expressbooking (+ /bulk), because "CRM bookings" collided with the genuinely separate CRM feature (§3.5). Deliberately not renamed: the stored Bookingsource: "CRM_Console" value and internal/assignment/crm_assignment.go's AssignCRMMiler — a working assignment engine and an already-written DB value, not just a label; renaming those is a deeper change than was asked for.

Deliberately skipped: rider substitutions (/v1/substitutions in Nearle) — Suriya's own call, looked low-traffic in the old system.

8.3 Verified vs not, for this migration slice

Verified: every new handler's field/column names checked by hand against real structs; brace-balance confirmed after every edit; no duplicate symbol definitions. go build ./... and go vet ./... both pass clean (verified on Suriya's machine, not the sandbox that wrote the code) — committed as c272a33, pushed to origin/main. The hand-verified code compiled correctly the first time it hit a real toolchain. Not verified:

  1. No integration test has hit any of the 14 new endpoints — a clean compile says the code is well-formed, not that it behaves correctly against a real DB/Redis/NATS.
  2. The Tenantid migration hasn't executed against a real DB yet — will run automatically via AutoMigrate next deploy (additive, nullable, safe).
  3. Nothing on the client side has changed. The rider Flutter app and doormile_express_console still call jupiter.nearle.app. API coverage existing on Doormile does not mean traffic is using it. Not a flip-a-URL cutover either — response shapes are completely different (flat 87-column Nearle rows vs nested Doormile JSON) — every screen that parses a response needs rewriting, not just repointing.

8.4 Logistics pickup-source & base-handover flow (2026-09-02)

[verified this session] — closes requests 25–31 on the Miler logistics line. Full contract, state-transition tables and wire values: docs/logistics-base-handover.md.

Vocabulary. The wire says hub; the rider app renders it as Base. Never change a wire value to match the app's wording: inward_at_hub, Inwarded_at_Hub, next_hub, pickup_source_type: "hub" stay exactly as spelt.

Feature flag MILER_HUB_HANDOVER_ENABLED (default off, read per request, same pattern as MILER_COLLECTED_STATE_ENABLED). On, a hub-routed parcel stops at Created at pickup-complete and only reaches Inwarded_at_Hub when the handover is recorded. Off (today), pickup-complete marks it Inwarded_at_Hub immediately — which is what the deployed rider app expects. Do not turn it on until a rider build that calls inward-at-hub is live, or every intercity parcel strands on Created with no way to advance it. Everything else in this work is ungated.

New endpoints (4):

Method Path Handler
POST /miler/consignments/:id/inward-at-hub MilerInwardConsignmentAtHub
GET /miler/bases MilerGetBases
GET /hub/inbound/expected GetHubInboundExpected
POST /hub/inbound/:id/reconcile ReconcileHubInbound

New columns (additive, nullable, AutoMigrate; no CHECK constraint needed widening — Created was already permitted on consignments): pickupbookings.pickupsourcetype, pickupbookings.pickuphubid, consignments.inwardedat.

Conventions added — reuse these, don't reimplement:

  • renderBase(hub) (controllers/logisticsHandoverController.go) is the ONE shape a base is returned in — all six fields, everywhere. A test enforces the count, because five of six leaves a rider unable to navigate.
  • nextActionForConsignment(status) is the ONE definition of what a rider does next. pickup-complete, the queue read and the consignment read all call it, so a poll can never disagree with the pivot.
  • resolveHandoverHub(booking, riderHubID) decides which base a parcel goes to. Backend decides; the app never picks a base.
  • pickupSource(booking, customerName) resolves type/id/name/address for any booking row, in the miler queue, the hub dispatch board and the admin detail.
  • scopeConsignmentsToOwnTenant(c, query) (hubInboundController.go) is the consignment counterpart of scopeBookingsToOwnTenant — use it on any new hub-console consignment query.

Two pre-existing bugs fixed in passing: a hub-routed pickup left its BookingAssignment open forever, so the rider could never go off duty (MilerEndDuty refuses while any assignment is Assigned/Accepted); and the no-rider-hub fallback took whichever hub row an unordered query returned first, now nearest-active-base by haversine.

Not verified: no integration test has hit the 4 new endpoints; the migration has not run against a real DB. go build, go vet and go test ./... all pass.


8.5 Customer app v1 — the /customer/* rebuild (2026-09-05)

[verified this session] — implements Doormile — Backend Requirements (Customer App v1) for the new doormile_cx Flutter client. Full contract, decisions and the written answers to the requirement doc's open questions: docs/customer-app-api.md. Spec: docs/openapi-customer.yaml.

The structural change: a customer books a PICKUP, not a shipment. One booking → 1..N destinations → one consignment and one tracking number per destination, minted when the miler completes the pickup. pickupbookings carried exactly one delivery address in its own columns, so there was nowhere to put a second; bookingdestinations is what closes that.

The compatibility rule that makes it safe — do not break it: destination 0 is mirrored onto the booking's flat delivery* columns. The miler app, the hub console, the routing code and the hyperlocal check all read those columns and none of them changed. A booking with no destination rows (every console/express booking, every pre-existing row) produces exactly one consignment through the same loop, byte-for-byte as before. Single-destination is one leg, never a special case.

Two prior surfaces were replaced, on Suriya's call (2026-09-05). The PIN auth (/customer/register|login|verify-pin|reset-pin, plus the email-OTP pair) and the single-destination booking create/list/detail/cancel/price and /customer/track/:trackingno are gone — doormile_customer_app is being retired in favour of doormile_cx. controllers/otpController.go was deleted with them. Customer routes: 19 → 28.

Identifier formats changed platform-wide. generateBookingNo() now mints DM-###### and generateTrackingNo() mints DMX########, both off Postgres sequences (cx_booking_reference_seq, cx_tracking_seq, created in migrations/migrate.go). The old generators used four random bytes; both columns are UNIQUE and a random short id collides long before the space runs out. Existing rows keep their DM-BK-/DM-TRK- strings — nothing parses either format, so the two coexist and the console just shows the new one for new work.

Conventions added — reuse these, don't reimplement

  • utils.CxOK / CxCreated / CxList / CxFail (utils/response_cx.go) are the ONLY response helpers for /customer/*. Deliberately separate from utils.OK/Fail: the customer contract always sends message (empty on success) and nests the code under error.code, while miler/console put code at the top level. Never mix them on one surface.
  • utils.EpochMillis(t) (utils/epoch.go) is the ONLY way a timestamp leaves /customer/*. This DB stores IST wall-clock digits (see DBNow), so t.UnixMilli() is off by 5h30m — the same defect that produced "yesterday's work shown as today" on the miler app. utils/epoch_test.go asserts it for both taggings the driver can produce.
  • internal/cxstage is the ONE place a customer stage is written. Record takes the caller's *gorm.DB — a stage event must commit or roll back with the operational write it describes. It dedupes per (booking, destination, stage), and Notify fires only after commit.
  • renderCxBooking + loadCxBundle (controllers/cxBookingView.go) build the canonical booking object. Every read that returns a booking goes through them; loadCxBundle is a fixed number of queries regardless of page size.
  • cxDestinationForConsignment(id) resolves a consignment to its booking. Use it instead of WHERE consignmentid = ? on pickupbookings — that column names only the FIRST order of a multi-destination pickup (see the bugs below).
  • cxPickupLegs(tx, booking) splits a booking into the journeys to create at pickup-complete. It is what decides single-vs-fan-out; nothing downstream needs to know which it got.

Stage derivation (the actual work)

Nine stages, lowercase snake_case, in constants.CxStage*. The client parses them verbatim and silently falls back to booked on an unknown key — never add or rename one without a client release. A booking rolls up from its slowest order once parcels split, or a customer sees "Delivered" while a parcel is still at a hub. Nothing is backfilled: a pre-existing booking gets a short honest history rather than an invented one.

cxstage.Release is the one place a stage moves backwards — a miler cancelling returns the pickup to the pool rather than cancelling it, and without walking the stage back the customer keeps seeing a rider who is not coming.

Four pre-existing bugs fixed in passing

All the same root cause, all found because the fan-out forced every consignment lookup to be re-read. Each would have broken multi-destination pickups outright:

  1. MilerDeliverConsignment could not close orders 2..N — its ownership check was WHERE consignmentid = ? AND assignedmileruserid = ? on pickupbookings, so a rider delivering the second parcel of a three-stop visit got "assigned consignment not found" and could not complete at all.
  2. MilerStartDelivery notified nobody for orders 2..N — same join, so no push and no receiver OTP.
  3. MilerInwardConsignmentAtHub left assignments open for orders 2..N — the rider could not go off duty (MilerEndDuty refuses on an open assignment) and the leg's distance/earnings recorded as zero.
  4. GET /miler/bookings showed only the first order — one row per booking keyed on that same column, so the fan-out would have minted orders no rider could see or deliver. milerStopsForBooking now emits one stop per order after collection, one visit before it, and exactly one row (unchanged) for a booking with no destination rows.

Also: CityGateMiddleware was a no-op for customer bookings. It sniffs the body for pickuppincode, which the new request shape does not carry, so every customer booking sailed past the operating-city gate. Now checked in the handler via the exported middlewares.PincodeInOperatingCity.

Blockers and gaps — state these plainly if asked

  • No SMS provider exists. internal/sms is the seam (a Sender interface, a logging sink, sms.Register()); until a gateway is plugged in, OTP codes go to the application log and nowhere else. This is the single blocker on real customer sign-in — and it is the same wall §9's E2E test hit. Staging has CX_STAGING_OTP (refused when ENV=production), which unblocks automated tests.
  • No integration test has hit any of these endpoints. go build, go vet and go test ./... pass; new unit tests cover the pure logic (stage rollup, epoch conversion, phone normalisation, weight fallback). None of that proves behaviour against a real DB/Redis/NATS.
  • The migration has not run against a real database. Additive, so it should be safe — but that is not the same as having run.
  • Failed delivery is invisible to the customer. MilerSkipDelivery works operationally, but there is no tenth stage key for it and an unknown key renders as booked, so a failed attempt leaves the parcel showing "Out for delivery". Needs product + a client release.
  • Latency (p95 ≤ 400ms) is unmeasured. Reads are batched and pricing is Redis-warmed, but that is an argument, not a measurement.
  • No retention policy for parcel photos or PII — nothing prunes either. The 30-minute signed-URL TTL limits link lifetime, not object lifetime.

New env vars

GEOCODER_URL, GEOCODER_EMAIL (place proxy — the app is never handed a map key, after the legacy rider app's key had to be revoked), MILER_CALL_PROXY (masked calling; empty exposes the rider's real number — set before launch), CX_STAGING_OTP, CX_ALLOW_STAGE_OVERRIDE.


8.6 Customer sign-in, config hardening & the ordering fix (2026-09-11)

[verified this session] — six defects were reported against customer bookings. Five were real; one was not. Everything below was verified against the code, and the schema question against the production database.

The one that locked every customer out (DM-01)

CxVerifyOtp read json:"code". The customer app sent otp — because docs/customer-app-api-crisp.md documented otp, while docs/openapi-customer.yaml correctly said code. The two docs disagreed and the app was built from the wrong one. req.Code was therefore always empty, the empty-code branch always fired, and every sign-in failed with 400 "Enter the code we sent you" — a correct code failed exactly like a wrong one. No amount of SMS-gateway credit would have fixed it.

The handler now accepts otp as a deprecated alias; code wins when both are present. This deliberately fixes builds already in customers' hands, which an app-side fix alone cannot. controllers/cxOtpFieldAlias_test.go pins both names, the precedence, and that a blank/missing code is still rejected. Remove the alias once the install base has moved on.

A failed OTP send no longer punishes the customer (DM-05)

issueCxOtp writes the code, the resend cooldown and the rate-limit slot before attempting delivery — it has to, the code must exist to be sent. But on failure it kept all three: the customer was told something went wrong, could not resend until the cooldown expired, had spent one of their five hourly codes, and a valid code they never received sat live in Redis for its full TTL. All three are now rolled back on a delivery error (DEL the code and cooldown, DECR the rate counter), so a retry is immediate and nothing usable is left.

Production credentials are no longer defaults (DM-06)

config.Load() defaulted JWT_SECRET_KEY to a literal, and NATS_URL / NATS_USER / NATS_PASSWORD / AI_LAYER_BASE_URL / ROUTE_OPTIMIZER_URL / DB_PASSWORD to the real production values. Two consequences, both live: a clone of this repository could mint a valid token for any user id and any role against any deployment that had not overridden the secret; and go run . on a laptop silently joined the production NATS cluster and competed with the real workers for the same durable consumer. This was hit accidentally on 2026-09-11 — a local instance pulled api.v1.bookings.update for real booking ids and caused redelivery churn on production for ~40 seconds.

All now default to empty, and the empty case is handled rather than assumed: InitNATS skips connecting, routing.BaseURL == "" already disabled sequencing, and the AI layer returns an error so the caller's existing AI_LAYER_FALLBACK path takes over with legacy scoring. A second hardcoded production URL in internal/assignment/ai_layer.go (not in the original report) was removed too.

JWT_SECRET_KEY is special-cased because an empty signing key is worse than a shared one: cfg.Validate(), called from main, refuses to start when it is unset in production. Outside production an ephemeral per-process key is generated with a warning, so local development needs no configuration while tokens stop surviving a restart. GEOCODER_URL deliberately keeps its default — Nominatim is a public service, not a Doormile host.

GET /admin/bookings is ordered (DM-04)

Added Order("bookingid DESC"). Without it the row order was unspecified — Postgres heap order, oldest first — which put the newest booking on the LAST page, outside the console's bounded drain window, and made OFFSET paging unstable enough to duplicate and skip rows. The primary key is unique, so the sort needs no tiebreaker.

The console half lives in the admin console repo (krow_talent_app — the directory name is stale; it is the Doormile Express Console): it requests pagesize=1000, receives 100, and stops after 12 pages, so it sees 1200 rows regardless. With the list now newest-first those 1200 are the most recent ones rather than the oldest, which turns a silent disappearance into a bounded view.

DM-03 (timestamp drift) is NOT a production bug — do not "fix" it

Reported as DBNow() relabelling IST wall-clock as UTC against timestamp with time zone columns, causing a +5:30 drift that hid evening bookings from the console. Verified against production and it is false there:

pickupbookings.createdat            timestamp without time zone
pickupbookings.preferredpickupfrom  timestamp without time zone
appcustomers.createdat              timestamp without time zone

Which is exactly what DBNow()'s own comment assumes. A round-trip confirmed it: a customer created at a known 12:10:34 IST stored as 12:10:34.094323. Zero drift. Changing DBNow() would introduce the bug, not fix it.

The real finding is the reporter's own fallback: GORM's Postgres driver maps time.Time to timestamptz, so a schema built fresh from AutoMigrate does NOT match production, and every new dev environment WILL show the +5:30 drift that production does not. That is why they saw it. Pin the column types explicitly in the models before this bites someone again.

Docs corrected

docs/customer-app-api-crisp.md had three request shapes that did not match their parsers, all failing silently through BodyParser — no error, just a zero value:

Endpoint Documented Actually parsed
auth/otp/verify otp code (now both)
fare/estimate pickup.latitude/longitude, packages[].weightKg pickup.lat/lng, packageCount
bookings flat destination fields, pickup.latitude nested details{}, pickup.lat/lng

The booking response block was wrong the same way (latitude/longitude where renderCxBooking emits lat/lng). docs/express-console-api.md also claimed pagination "default 500, cap 1000" when the code is default 20, cap 100 — which is what made the console size its page budget for twelve times the rows it actually receives.

When a doc and a parser disagree here, the parser has won every time. Three separate client teams have now built against wrong Doormile docs in one week.

Still open from this report

  • DM-02: email OTP returns 500 on production. SMTP_HOST, SMTP_USER and SMTP_PASSWORD all default to "" and are not set. Either configure SMTP or hide the app's Email tab — offering a path that always fails is worse than not offering it.
  • The committed secrets (.env and a live GCP service-account private key) are still tracked in git and pushed. Removing the defaults above does not help until those keys are rotated.
  • GET /api/v1/ready returns 503 while its body says "status":"ready" (routes/routes.go) — a monitor reading the body sees the opposite of the status code.

9. Current blockers & open work (whole-project level)

[carried forward]

  • Last live end-to-end system test stalled on customer JWT acquisition — customer login requires Firebase OTP on a real phone, can't be scripted from the command line. Needs either a real-phone run or a load-test workaround.
  • A load test targeting high concurrent bookings (capacity check before go-live) was in progress at the end of the last relevant session, not finished.
  • Go-live preparation across the 4 target cities is still ahead.

Migration-slice-specific (§8.3): go build verification is now done (commit c272a33). Still pending: integration testing of the 14 new endpoints, running the Tenantid migration, and the client-app rewrites.


10. Deliberately skipped / open decisions

  • Rider substitutions — skipped, low old-system traffic. Revisit if it turns out to matter.
  • B2C tenant attribution — PickupBooking.Tenantid stays nil for B2C bookings; whether direct-to-consumer traffic should be attributed to one of the 4 Doormile-ops tenants (per city) is a business decision, not something to guess at.
  • Batch route optimization — HubBatchAssign is a single-booking nearest-rider heuristic, not a multi-stop VRP solver (§4). Untested against real volume vs. whatever the old paid external services provided.
  • PartnerInfo vs Tenant naming confusion — flagged, not acted on.
  • Customer/CustomerLocation (legacy) vs AppCustomer (new B2C) — two customer-shaped tables coexisting, flagged as a future duplicate-data risk, not resolved.
  • AgentDecision.context_embedding dimension (1536) vs the reported MiniLM embedding size (384) — flagged this session (§2), not investigated further; worth resolving before relying on vector search quality from that column.

11. Suggested next steps, in order

  1. Resolve the customer-JWT E2E test blocker (real phone or load-test workaround) and finish the load test.
  2. go build ./... in DoormileBackend — done, clean pass, commit c272a33 pushed to origin/main.
  3. Stand up a test/staging DB, let AutoMigrate run, smoke-test the 14 new migration-session endpoints with real requests.
  4. Pick one low-risk console slice (e.g. Reports or Partner management — net-new UI, not replacing something live) and wire it to Doormile instead of jupiter — first real proof the cutover works end to end.
  5. Only after that: rewrite the higher-traffic screens (orders/deliveries list, rider status updates) and the rider app's request/response handling.
  6. Resolve the B2C tenant-attribution question before it's load-bearing for real revenue reporting.
  7. Confirm the AgentDecision embedding-dimension question.
  8. Decide on substitutions and batch-optimization sophistication once real usage data says whether they're actually needed.
  9. Go-live preparation across the 4 target cities.