Doormile AI: document the rules that are load-bearing

CLAUDE.md for the assistant folder, plus the capability roadmap.

The rules worth reading before editing either file:
- `truncated` is not optional to handle; an intent that reads scan.rows and
  ignores it reintroduces the silent under-reporting this replaced.
- Order status comes from utils/orderStatusGroups.js, not api.js's
  Deliveries taxonomy, and the two must not be merged.
- State questions ("how many are assigned") must not default to today;
  flow questions ("how many orders today") still do.
- orderQuery must not widen to claim single-dimension questions.
- Never put a React element in message state — localStorage round-trips it.

ROADMAP.md carries the analysis and the L0-L8 capability ladder, including
what is deliberately still out of scope: write actions pending a confirm
gate, and any LLM step pending a decision on where its key would live.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
2026-08-18 15:54:12 +05:30
parent 51667d1c9f
commit 1ea936c7d8
2 changed files with 521 additions and 0 deletions

View File

@@ -0,0 +1,131 @@
# CLAUDE.md — `src/pages/nearle/assistant/`
Rules for editing **Doormile AI** — the Operations Copilot (`intents.js`, `DoormileAI/`). Read this before touching either.
---
## 1. What this is
An in-console Q&A assistant that answers operator questions about live data — "how many orders today", "morning batch orders", "how many riders are active" — by calling the same API functions every other page in this console already uses. Product name is **Doormile AI**, subtitle **Operations Copilot**. Not a standalone page: it lives as a **right-side slide-over** opened from a header icon.
- **Mounted in**: `src/layout/MainLayout/AppTopNav.js` — a single `<DoormileAITrigger />`. There is **no route and no sidebar entry** for this feature — don't add one back. If you're tempted to give it a full page, re-read §2 first; that was tried and deliberately reverted.
- **`intents.js`** — all data logic: the intent catalog, keyword/phrase matching, and the API calls that produce answers. The UI never fetches.
- **`DoormileAI/`** — all UI:
- `index.js` — the trigger button; owns open/closed state and returns focus to itself on close.
- `AIPanel.js` — the portal, scrim, slide-over, focus/Escape handling, message state, history persistence, and the ask() flow.
- `AIWelcome.js` — greeting + suggestion cards (empty-thread state only).
- `AIMessage.js` — one turn. User turns are bubbles; assistant turns deliberately are NOT.
- `AIComposer.js` — auto-growing textarea, Enter to send, Shift+Enter for newline.
- `AIParts.js` — Spark / LiveIndicator / TypingIndicator / Metric / StatGrid / StateBlock.
- `pageContext.js` — route → context label + suggested questions.
- **`DoormileAI.css`** — the panel's stylesheet (same convention as `OrdersRedesign.css`).
### UI rules that are load-bearing, not cosmetic
- **Assistant turns must not become bubbles.** The no-bubble treatment is what keeps this reading as part of the dashboard rather than a bolted-on chatbot.
- **Never put a React element in message state.** Messages are JSON round-tripped through `localStorage`; elements don't survive it (`$$typeof` is a Symbol and is dropped) and the rehydrated value crashes the next render. Icons are referenced by *key* (`iconKey`) and resolved in `AIParts.js`. Same rule for anything new you add to a message.
- **Selectors that style an Astryx Stack need two classes.** `padding={0}` emits a StyleX atomic at the same (0,1,0) specificity as a bare class, so `.dai-header` can lose on stylesheet order. Those rules are written `.dai-root .dai-header`. Don't "simplify" them back to one class. This never applies to `.dai-panel`/`.dai-scrim`, which carry `.dai-root` on the *same* element.
- **`--dai-accent` is the single accent knob.** It resolves to the app accent (black, root CLAUDE.md §6.2). Switching the assistant to Doormile red is one line in `DoormileAI.css`, not a hunt through components.
- **Every suggestion in `pageContext.js` must actually resolve** against `INTENTS`. A chip that returns "I can't answer that yet" is worse than no chip — check it before adding.
---
## 2. Why this is deterministic, not LLM-based
This was a deliberate, explicit product decision (not a technical limitation worked around silently): **this app has zero backend of its own** — confirmed exhaustively (no `server/`, no Firebase Cloud Functions, no `firebase-admin`/`firebase-functions` dependency, `Dockerfile` just serves a static CRA build via nginx). An LLM call needs an API key held server-side; there is nowhere in this repo's infrastructure to put one without shipping it to the browser.
Two paths existed: add a new endpoint to `api.doormile.com` to hold the key (rejected — "don't need to create the new endpoints, use the existing ones"), or stay fully client-side with a much richer deterministic matcher (chosen). **Do not silently reach for an LLM/RAG library here** without first getting a decision on where its key would live — that conversation already happened once and the answer was no.
RAG (retrieval-augmented generation over a document/vector store) was also explicitly considered and rejected as the wrong tool: this bot's data isn't unstructured documents, it's structured operational data already reachable through typed API functions. The correct "grounded answer" pattern for that is what's already here — a fixed catalog of `{match, run}` pairs, not a vector search.
---
## 3. The intent pattern (`intents.js`)
Each entry in `INTENTS` is:
```js
{
id: 'someIntent',
label: 'Human-readable description — e.g. "example phrasing"',
match: (text) => params | null, // does this intent apply? extract params or refuse
run: async (params) => ({ headline, detail, sourceCalls }) | null // real answer, or "couldn't resolve"
}
```
- `answerQuestion(text)` walks `INTENTS` **in order** and returns the first intent whose `match` recognises the text **and** whose `run` resolves to a non-null result. `run` returning `null` means "the pattern matched but couldn't be resolved" (e.g. no tenant name in the question actually matched a real tenant) — the loop falls through to the next intent rather than answering with a guess.
- `sourceCalls` feeds `<ChatToolCalls>` behind the per-message "Sources" disclosure in `AIMessage.js` — every answer can still show which API was queried and what came back, so an operator can verify it wasn't invented. It is collapsed by default for visual quiet; **do not remove it**, that disclosure is the verifiability contract.
- `metric` and `stats` (optional) drive the panel's headline number and breakdown grid. `stats` comes from `statusStats()`, which tallies the *same* `mapBookingStatusToDeliveryStatus` classification as the prose `headline`, so the sentence and the cards can never disagree.
- Every `run` calls a real function from `pages/api/api.js` / `pages/api/doormileApi.js`. **Never fabricate a number** — if no existing function covers a question, either add a new intent that calls a real endpoint, or leave the question unanswered (falls through to the "I can't answer that one yet" state in `AIPanel.js`). A wrong number from this bot is worse than no answer.
### Coverage grows with the console, not ahead of it
The catalog covers orders/bookings, riders, tenants, and the fleet/ops resources (hubs, vehicles, tripsheets, exceptions, app users, customers, pricing, consignments, partners, competitor branches, carrier pricing). When a new admin resource gets its own page in this console, add a matching intent here too — and **always call the exact same `getX()` function that page's own table already calls** (e.g. `hubStatus` calls `getHubs()`, the same function `hubs.js` uses). Never write a bespoke fetch for the bot. This is what keeps the bot's numbers live and in agreement with what the corresponding page shows — the whole point of not hand-rolling a separate data path.
### Composite questions — `orderQuery`
`orderQuery` (ordered 3rd, ahead of `riderCounts`) is the one intent that composes filters: status x batch x tenant x rider, plus rankings ("top 5 tenants by orders"). Everything else in the catalog answers exactly one dimension and discards the rest of the sentence.
**It claims a question only when two or more of status/batch/tenant/rider are present, or a ranking is asked for.** A date is deliberately NOT counted as a dimension — every intent already handles dates, and counting it re-routed four working questions ("how many cancelled orders today") away from the intents that answer them better. If you widen this matcher, re-run the routing probe first; over-claiming here silently changes answers across the whole catalog.
`run` returns `null` when a named tenant or rider doesn't resolve, so an unrecognised name falls through rather than having its filter silently dropped — which is the exact bug this intent exists to fix.
Entity names resolve through `bestNameMatch`, which is bidirectional (the question may name a shorter or longer form than the record) and prefers the longest match, so "Acme" can't beat "Acme Foods" when both exist.
### Beyond single-question matching
A few layers sit on top of the plain `{match, run}` loop, all in `intents.js`, all still deterministic (no LLM):
- **Typo tolerance** — `correctTypos()` runs once before matching, correcting misspelled domain keywords (length ≥5, Levenshtein distance ≤1/≤2) against a fixed `KEYWORD_VOCAB`. It never touches order IDs, tenant names, or short words — only known keywords get "corrected," so it can't invent a wrong one.
- **Richer dates** — `explicitDateFromWords` (DD/MM/YYYY, ISO), `weekdayFromWords` (most recent past occurrence of a named day), and `rangeFromWords` (this/last week, this/last month, explicit "from X to Y") feed `dayFromWords`/`rangeFromWords`. Still a fixed vocabulary, not a date-parsing library — an unrecognised phrase falls back to today, never a guessed date.
- **Comparisons** — `comparisonIntent` (trigger: "vs"/"versus"/"compare[d] to") runs two `fetchBookingsInRange` calls and reports both counts/totals side by side. Ordered early (right after `tenantList`) since it must win before `totalOrders`/`revenueTotal` would otherwise swallow the question on the bare word "orders"/"revenue".
- **Multi-part answers** — `answerMultiPart()` splits on and/,/&, matches each segment independently through the same `INTENTS`, and only combines them if ≥2 segments resolve. A single-segment match falls through to the normal path untouched.
- **Follow-up context** — `answerQuestion(text, context)` takes `{ lastIntentId, lastParams }` from the previous turn (tracked in `AIPanel.js`'s state). If the new text is a bare date/range phrase ("what about yesterday?") with no other domain keyword, it re-runs the *same* intent with the date swapped rather than requiring the whole question again. This is pattern-matching on the phrase shape, not real conversational memory — a question that also names a different domain is treated as new, not a follow-up.
- **Entity lookups** — `riderLookup`/`hubLookup`/`vehicleLookup` require an explicit `LOOKUP_TRIGGER` phrase ("find"/"where is"/"status of"/"search for") before a name, and are ordered ahead of their aggregate counterparts (`riderCounts`/`hubStatus`/`vehicleStatus`) so a named-entity question doesn't get swallowed by the count intent.
### Ordering and cross-domain guards — read before adding an intent
A real bug shipped here once: `statusBreakdown` matched the word "active" (a valid order status), so "how many riders are active today" was swallowed by the order-status intent and called `getBookings` instead of `getallridersummary` — because `statusBreakdown` sat earlier in `INTENTS` than `riderCounts` and its `run` never returns `null` (it always finds *some* count, even 0), so it never yielded.
The fix, and the rule going forward:
1. **Domain-specific intents (rider, tenant) are ordered near the top**, ahead of the generic order/status/date intents, so an unambiguous keyword like "rider" always wins first-match.
2. **Generic intents explicitly refuse to match on another domain's keyword**, via helper guards like `mentionsRiders(text)` at the top of their `match`. This is deliberately redundant with (1) — if someone reorders `INTENTS` later without noticing the significance, the guards still hold.
If you add a new intent whose trigger words could plausibly appear in an unrelated intent's question (status words, "for", generic nouns), do both: place it appropriately in the order, and add a guard to anything downstream it could shadow — don't rely on ordering alone.
### Date/batch/status vocabulary — reuse, don't reinvent
- **Batch bucketing** (`morning`/`afternoon`/`evening`) comes from `src/utils/batchBucket.js`, extracted from `Dispatch.js`/`deliveries.js`'s canonical model (see `dispatch/CLAUDE.md` §1). Bucketing on anything other than `orderdate` (a booking's `createdat`) will disagree with what those two pages show — don't reintroduce `expecteddeliverytime`/`assigntime` bucketing here, they were both tried and rejected for the same reasons documented there.
- **Order status classification comes from `utils/orderStatusGroups.js`**, which is the SAME match set the Orders page's tabs count with (`orders.js` imports `statusesInGroup` for its `ORDERS_STATUS_TABS`). Use `groupForBookingStatus` / `isInGroup` / `statusesInGroup`; don't grow a third definition.
- It is deliberately **not** `mapBookingStatusToDeliveryStatus` (api.js), which is the *Deliveries* page's rider-centric taxonomy and keeps `miler_assigned` on `pending`. The two exist on purpose — Orders tracks the operator's action, Deliveries tracks the rider's. Don't merge them; that was tried and reverted per explicit product direction.
- The assistant answers order-status questions with the ORDERS taxonomy because that is the screen an operator compares its answers against. A live bug came from the mismatch: the Orders page showed 19 Assigned while the bot said 0.
- **Date words** are a fixed, small vocabulary — not a real date-parsing library. Don't guess at "the 5th" style phrasing; an unrecognised date phrase falls back to today rather than to a wrong date.
- **State questions vs flow questions — do not default a state question to today.** `mentionsAnyDate(text)` distinguishes "the question named a date" from "we defaulted to one":
- *State* ("how many orders are assigned / cancelled") describes the queue **right now** and must be unscoped, because the Orders page's tabs apply no date filter either. Scoping it to orders *created today* is what made the bot answer 0 against a page showing 19.
- *Flow* ("how many orders today", revenue, batches) genuinely needs a period and keeps the today default.
When a state question is answered unscoped, say so in the detail — the answer must never leave the operator guessing which window it covered.
- `GET /admin/bookings` has no server-side date/status/tenant filter, so every intent fetches and filters client-side. It does **not** fetch a single page: `pagesize` is capped at 1000 server-side, so a lone `getBookings(1, 1000)` silently under-reports the moment an account passes 1000 lifetime bookings. Use `fetchBookingsInRange(start, end)` / `fetchBookingsForDay(day)` / `scanBookings()`, which drain pages via `getBookingsPage` up to `MAX_PAGES` and return `{ rows, truncated, scanned, pagesFetched, total }`.
- **`truncated` is not optional to handle.** If you write a new intent, run its count through `countPhrase(scan, n)` ("At least 42"), append `truncationNote(scan)` to the detail, and build its audit entry with `scanCall(scan, ...)` — which reports `status: 'error'` when capped so the tool-call strip can't show a green "complete" beside a partial number. An intent that reads `scan.rows` and ignores `scan.truncated` reintroduces exactly the bug this replaced.
- **Revenue excludes cancelled orders and is labelled "estimated"** — `revenueOf(rows)` sums every `serviceoptions[].estimatedprice` on non-cancelled rows. It is a quote, not a settled amount; don't relabel it "revenue" flat.
- **"Assigned" does not go through the coarse bucket.** `mapBookingStatusToDeliveryStatus` collapses `miler_assigned` into `pending` alongside `pending_pickup` (orders with no rider at all), so `rawStatusFromWords` matches the backend enum directly. Any other question naming a raw enum should do the same rather than being forced into a delivery-status bucket.
---
## 4. What's deliberately out of scope right now
- **Write actions.** The original ask included "if I say create an order, it should create it" — deliberately **not built**. Giving a keyword-matched bot the ability to mutate data (order creation has real validation elsewhere: CityGate pincode checks, delivery-slot windows, the dispatch reconcile-before-commit rule) is a materially bigger risk than read-only Q&A. If this gets built, it needs its own guardrail — the bot proposes what it would submit, the operator explicitly confirms, only then does a real create-order call fire. Don't wire a write action straight from intent match to a mutation call.
- **Open-ended LLM understanding.** See §2. Revisit only with an explicit decision on where the LLM key lives.
- **Tenant/role-aware scoping.** Every intent currently queries the same data an unscoped admin session would see — there's no per-login "you only see your own tenant" filter applied inside `intents.js` itself. Needs a decision on how tenant-locked logins should be detected (`localStorage.tenantid`/`roleid`) and whether that's a hard filter or just a default, before it's built.
- **Proactive alerts.** Surfacing anomalies unprompted (e.g. "3 hubs inactive") via the notification bell is a different feature from Q&A — it needs a polling/watch mechanism, and the notification panel it would feed is currently static UI scaffolding, not wired to a real alert stream. Not started.
- **Automated tests for the intent matcher.** The matcher is pure functions (`match`/`run` per intent) and would be straightforward to unit-test, but the project's stated convention is "no tests of consequence, lint is the only gate" (root `CLAUDE.md`). Adding a test framework here is a scope decision for the user, not something to introduce silently.
---
## 5. Don'ts
- Don't re-add a route/sidebar entry for this feature — it's a header slide-over, not a page.
- Don't let a new intent's `match` fire without considering what other intents' trigger words it might contain (see §3's ordering rule).
- Don't bucket batches or classify statuses with page-local logic — reuse `utils/batchBucket.js` and `mapBookingStatusToDeliveryStatus`.
- Don't answer with a number that didn't come from `sourceCalls`-tracked real data — including anything rendered into a `metric` or `stats` card.
- Don't surface a raw API error string to the operator. Errors log to `console.error` and render as the polished error state; the toast that used to leak `err.response.data.message` is gone.

View File

@@ -0,0 +1,390 @@
# Doormile Bot — v3 development plan
Analysis of the shipped v2 (`BotPanel.js` + `intents.js`, 9 intents) and the staged plan to take it to an advanced operator assistant.
Authority: `src/pages/nearle/assistant/CLAUDE.md` constrains this work — no route/sidebar (§5), no LLM without a key-location decision (§2), no write action without a propose→confirm gate (§4). This plan does not override any of those; where it touches one, it says so explicitly.
---
## Part 1 — Analysis of v2
### Architecture today
```
answerQuestion(text)
└─ for each of 9 INTENTS, in array order
├─ intent.match(text) → params | null
└─ intent.run(params) → { headline, detail, sourceCalls } | null
first match that resolves wins; otherwise "I can't answer that yet"
```
One question maps to exactly one intent. Every data intent bulk-fetches `getBookings(1, 1000)` and filters client-side.
### Defects, ranked by damage
#### Tier A — the bot states wrong numbers confidently
**A1. The 1000-row ceiling is real and undetectable.**
`BULK_PAGESIZE = 1000` is not a chosen page size, it is the API's hard cap (`express-console-api.md` → Conventions: "Default 500, cap 1000"). Worse, `getBookings` returns `response.data.data` and discards the envelope's `total`, so no caller can even detect truncation. Past 1000 lifetime bookings every count intent under-reports with no warning, while `sourceCalls` displays `status: 'complete'` beside it. This directly violates the stated promise in CLAUDE.md §3 ("a wrong number from this bot is worse than no answer").
**A2. Range words are silently downgraded to "today".**
Only `revenueTotal` and `weekOrders` call `rangeFromWords`. The other seven use `dayFromWords`, which ignores "this week" and returns today. "How many delivered orders this week" is answered by `statusBreakdown` as *today's* count. The headline says "today", so it is disclosed — but the operator asked something else and got an answer to a different question.
**A3. `revenueTotal` is mislabelled and includes cancelled orders.**
It sums `serviceoptions[0].estimatedprice` over every row in range with no status filter. Cancelled bookings inflate it, only the first service option is counted, and the figure is an *estimate* presented as "Total revenue".
**A4. "assigned" resolves to the wrong bucket.**
`statusFromWords` maps `assigned → 'accepted'`, but `mapBookingStatusToDeliveryStatus` maps the backend's `miler_assigned → 'pending'`. "How many assigned orders today" therefore counts `pickup_scheduled` + `converted_to_consignment` and excludes the orders the operator means.
#### Tier B — the bot answers a different question
**B1. `riderCounts` swallows every sentence containing "rider".**
`match: (text) => (mentionsRiders(text) ? {} : null)` and its `run` never returns null. "Which rider has order #4821", "how many orders did rider Suresh deliver" and "rider performance this week" all return the same fleet availability summary. This is the mirror image of the bug CLAUDE.md §3 documents as fixed — the guard stopped `statusBreakdown` stealing rider questions, but nothing stops `riderCounts` stealing everything else.
**B2. `orderLookup`'s fall-through answers a different question entirely.**
It searches only page 1. An order not in the most recent 1000 returns `null`, the loop continues, and `totalOrders` answers "142 orders created today" to the question "status of order #9931". Fall-through is right when an intent *mismatched*; it is wrong when the intent matched and the lookup failed.
**B3. No composite filters.** "Delivered orders for Acme this week" is answered by `statusBreakdown` alone — tenant and range discarded.
**B4. `resolveTenant` is one-directional substring matching.** The question must contain the tenant's full name. "Orders for acme foods" against tenant "Acme Foods Pvt Ltd" fails, falls through, and `totalOrders` answers with the all-tenant count.
**B5. `orderIdFromWords` treats any bare 4+ digit run as an order id** — years, pincodes, quantities.
#### Tier C — architecture
**C1. No TanStack Query.** CLAUDE.md §7 makes it the rule for all reads. `answerQuestion` calls the API functions raw. Clicking the five suggestion chips issues five separate 1000-row fetches; `tenantCount` issues two sequentially.
**C2. `sourceCalls` are hand-written after the fact** with `status: 'complete'` hardcoded. They cannot represent a failed or partial call and will drift from the code the moment a `run` is edited. `ChatToolCalls` already supports `'pending' | 'running' | 'complete' | 'error'` plus `duration`, `errorMessage` and `resultDetail` — none used.
**C3. The two aggregation endpoints are unused.** `GET /admin/reports` (`from`/`to`/`tenantid`/`locationid`/`hubid`, with `by_location`/`by_hub`/`by_tenant`/`by_rider` blocks per `getReports`'s comment) and `GET /admin/dashboard` ("counts + today's numbers") are server-computed, uncapped, and already wrapped in `doormileApi.js`. They are the correct source for every counting question and the answer to A1.
**C4. No abort or timeout.** A 1000-row fetch cannot be cancelled; the composer is simply disabled.
#### Tier D — UX
- **D1.** Suggestion chips are gated on `messages.length === 0`, so they vanish permanently after the first question. No reset control.
- **D2.** History dies when the popover closes (accepted in CLAUDE.md §1, but a liability once answers get expensive).
- **D3.** Answers are two strings. `detail` truncates at 10 ids with "…and N more" and there is no way to see the rest, and no way to jump to the matching rows on the Orders page.
- **D4.** Unanswered questions are dropped — no signal on what to build next.
- **D5.** Fixed `PANEL_WIDTH = 400` / `MESSAGES_HEIGHT = 420` raw px.
---
## Part 2 — Target architecture
Replace one-question-one-intent with a three-stage pipeline:
```
parse(text) → Query pure, no I/O, fully unit-testable
resolve(Query) → Dataset cached, paginated, audited, abortable
render(Query, Dataset) → Answer headline + blocks + real sourceCalls
```
`Query` is a slot bag, not an intent id:
```js
{
subject : 'orders' | 'riders' | 'tenants' | 'revenue' | 'order',
metric : 'count' | 'sum' | 'breakdown' | 'lookup' | 'top',
filters : { range: {start, end, label}, batch, status, tenantId, hubId, riderId, orderId },
groupBy : 'status' | 'batch' | 'tenant' | 'rider' | 'hour' | null,
limit : number
}
```
Filters compose. "Delivered orders for Acme this week" fills three slots and runs one query instead of picking one of three intents.
Proposed file layout inside `src/pages/nearle/assistant/`:
```
parse/
vocab.js date/range, batch, status, metric, groupBy vocabularies
entities.js tenant/rider/hub resolution + fuzzy scoring
parseQuery.js text → Query (+ confidence, + unresolved slots)
resolve/
source.js cached, paginated booking source; reports/dashboard source
aggregate.js count / sum / breakdown / top over a normalised row set
render/
answer.js Query + Dataset → { headline, blocks[], sourceCalls[] }
intents.js thin compatibility shim → parseQuery + resolve + render
BotPanel.js richer rendering, chips, reset, stop, persistence
```
`intents.js` keeps its exported surface (`answerQuestion`, `EXAMPLE_QUESTIONS`) so `BotPanel.js` and `AppTopNav.js` are unaffected during the swap.
---
## Part 3 — Phased plan
### Phase 0 — Truth foundation (blocking; ship before any new capability)
Nothing else matters while the numbers can be wrong.
| # | Task | Files | Acceptance |
|---|---|---|---|
| 0.1 | Expose the envelope `total`/`page` from `getBookings` — return `{ rows, total, page }` or add `getBookingsPage`. Keep the existing signature working for `fetchDeliveries`' four callers. | `pages/api/doormileApi.js` | A caller can detect that more rows exist than were returned. |
| 0.2 | Build a paginated booking source that drains pages until the oldest row predates the requested range, with a hard page budget. Emits a `truncated` flag when the budget is hit. | `resolve/source.js` | "How many orders today" is correct on an account with >1000 lifetime bookings. |
| 0.3 | Route counting/aggregate questions through `getReports(from, to, tenantid, locationid, hubid)` first; fall back to the paginated booking scan only when reports can't answer the shape. | `resolve/source.js` | Counts for a date range come from one server call, not a 1000-row scan. |
| 0.4 | Never present a truncated result as complete — if `truncated`, headline reads "at least N" and the tool call carries `status: 'error'` or an `errorMessage`. | `render/answer.js` | Truncation is visible in the answer, not just the console. |
| 0.5 | Fix A3: exclude `cancelled` from revenue, sum all `serviceoptions`, relabel as "estimated". | `render/answer.js` | "Total revenue today" excludes cancelled and says "estimated". |
| 0.6 | Fix A4: align status synonyms with `BOOKING_STATUS_TO_DELIVERY_STATUS`. "assigned" → the bucket `miler_assigned` actually lands in. | `parse/vocab.js` | "Assigned orders today" matches what the Orders page shows for the same filter. |
| 0.7 | Fix B2: when an intent matched but its lookup failed, answer "I couldn't find order X" — do not fall through to a broader intent. | `resolve/` + `render/` | Asking for a nonexistent order never returns a global count. |
**Risk:** 0.3 depends on the `/admin/reports` response shape, which is documented only in `jupiter2doormile.md` and is not in `express-console-api.md`. Confirm live before building on it; if the shape doesn't carry what's needed, 0.2 alone still fixes A1 at higher cost.
### Phase 1 — Slot parser (the capability multiplier)
| # | Task | Acceptance |
|---|---|---|
| 1.1 | `parseQuery(text) → Query` with independent slot extraction; unrecognised slots stay empty rather than defaulting. | Unit tests over a fixed corpus (Part 5). |
| 1.2 | Real relative-date vocabulary: today, yesterday, this/last week, last N days, this month, explicit `DD MMM` and `YYYY-MM-DD`. Anything unparsed → *ask*, don't assume today. | "Delivered orders this week" returns the week, not today. |
| 1.3 | Entity resolution with bidirectional + fuzzy matching and an ambiguity path: 0 matches → say so; 1 → use it; 2+ → ask which. | "Orders for acme foods" resolves to "Acme Foods Pvt Ltd". |
| 1.4 | Confidence scoring replaces first-match-wins. Below threshold → clarifying question listing what was understood. | "Rider Suresh's orders today" no longer returns a fleet summary (fixes B1). |
| 1.5 | Conversation context: carry the last `Query` forward; a follow-up mutates only the slots it names. Reset on explicit "start over" and on an entity switch. | "and yesterday?" / "just for Acme" work as follow-ups. |
| 1.6 | Guard B5: a bare number is an order id only with an order-ish trigger word nearby and no date/quantity reading. | "Orders in 2026" is not treated as an id lookup. |
Coverage after this phase is the product of the slots, not a list of nine — subject × range × status × batch × tenant × rider all compose.
### Phase 2 — Answers that are objects, not sentences
| # | Task | Notes |
|---|---|---|
| 2.1 | `Answer.blocks[]` — typed blocks (`stat`, `table`, `breakdown`, `link`) rendered by `BotPanel`. | Replaces the two-string shape. |
| 2.2 | Result table for row-returning answers: `Table` + `StatusBadge` cells instead of "…and N more". | Reuse `components/nearle_components/StatusBadge`. |
| 2.3 | Deep link — every answer carries the filter state that produced it, with a button that navigates to the Orders/Deliveries page pre-filtered. | The single biggest usability jump: answer → action. |
| 2.4 | Breakdown answers via `groupBy` (by status / batch / tenant / rider / hour). | Feeds off `/admin/reports` `by_*` blocks where available. |
| 2.5 | New subjects using already-exported functions: `getBookingTrack` (where is order X), `getMilerActivity` / `getMilerSummary` (what has rider X done), `getConsignments`, `getTripsheets`, `getHubs`, `getVehicles`. | No new endpoints needed. |
| 2.6 | Real `sourceCalls`: emitted by the fetch layer, streaming `pending → running → complete/error`, with `duration` and `errorMessage`. | `ChatToolCalls` already supports all four states. |
### Phase 3 — Panel UX
| # | Task |
|---|---|
| 3.1 | Keep suggestion chips available after the first message (collapse into a `ChatComposerDrawer` or a header affordance), plus a "New chat" reset. |
| 3.2 | Persist thread + last `Query` to `sessionStorage` so closing the popover doesn't lose it. |
| 3.3 | `onStop` / `isStopShown` on `ChatComposer` wired to an `AbortController` through the fetch layer. |
| 3.4 | `ChatSystemMessage` for context resets, day dividers and truncation notices. |
| 3.5 | `ChatLayoutScrollButton` + `useChatNewMessages` for long threads. |
| 3.6 | `useTriggerMenu` + `ChatComposerTokenElement`: `@tenant` / `@rider` / `/` commands so an operator *picks* a real entity instead of relying on fuzzy matching. Directly de-risks 1.3. |
| 3.7 | Replace raw `PANEL_WIDTH`/`MESSAGES_HEIGHT` px with tokens; keyboard/focus check inside the `Popover` (Escape currently closes the panel mid-typing). |
| 3.8 | Log unanswered questions locally (capped ring buffer) and surface them — this is the backlog for the next intent round. |
| 3.9 | i18n the bot's strings into `utils/locales/en.json` like the rest of the app. |
### Phase 4 — Actions, confirm-gated (needs sign-off)
CLAUDE.md §4 rules this out today and specifies the shape it must take if built: propose → operator confirms → execute. Plan accordingly:
1. Parse produces an `Action` (never executed at parse time).
2. Render shows exactly what will be submitted — target rows, field values, the endpoint — as a `ChatSystemMessage` with explicit confirm/cancel.
3. Execute only on confirm, through the same api.js functions the pages use, honouring the dispatch reconcile rule (root CLAUDE.md §4) and the notify-rider-after-mutation rule (§9).
4. Post-action, refetch the related queries and show the new state.
Safe first candidates: `notifyMiler` (broadcast to a rider), `cancelBooking` (single, with confirm). Deliberately last: order creation — CityGate pincode validation and delivery-slot windows live elsewhere and must not be bypassed.
### Phase 5 — LLM as parser only (BLOCKED on a decision)
CLAUDE.md §2 records that this was raised and rejected because there is nowhere to hold a key — no backend, static nginx build. That reasoning still stands, so this phase is blocked, not dismissed. If the key question is ever answered, the correct shape is narrow:
- The model does **slot extraction only** — text in, a validated `Query` JSON out via tool-use / structured output. It never produces a number, a row, or a sentence the operator reads as fact.
- Deterministic code still executes every fetch and every calculation.
- Slots that don't resolve against real tenants/hubs/riders are rejected and the regex parser runs as fallback.
This preserves the "no fabricated numbers" guarantee exactly, while removing the vocabulary ceiling. Note it also makes Phase 1 the fallback path rather than dead work.
### Phase 6 — Proactive
Once the query layer is trustworthy: watch for conditions rather than waiting to be asked — "14 morning-batch orders unassigned, 30 minutes to cutoff", "rider X offline mid-route". Surfaces as a badge on the bot icon and a `ChatSystemMessage`. Ties into the existing FCM path.
---
## Part 4 — Decisions needed
1. **`/admin/reports` response shape** — confirm live. Blocks task 0.3, which is the cheap fix for the 1000-row problem.
2. **Write actions** — in scope for this round, or stays read-only? Blocks Phase 4 entirely.
3. **LLM key location** — unchanged from CLAUDE.md §2? Blocks Phase 5.
4. **History persistence** — `sessionStorage` (dies with the tab) or Redux + `localStorage` (survives)? Affects 3.2.
5. **Tenant-scoped logins** — should the bot say "across your tenant" when the token carries a tenantid, rather than implying global figures?
---
## Part 5 — Regression corpus
Build this as a fixture the parser is tested against; every row is a question the bot must either answer correctly or explicitly decline.
| Question | Must produce |
|---|---|
| how many orders today | count, orders, today |
| how many orders this week | count, orders, 7-day range (currently → today) |
| how many delivered orders this week | count + status + range (currently drops range) |
| delivered orders for Acme this week | count + status + tenant + range (currently drops two) |
| morning batch orders yesterday | count + batch + day |
| how many riders are active | rider availability |
| which rider has order #4821 | order lookup → rider (currently fleet summary) |
| how many orders did rider Suresh deliver today | rider activity (currently fleet summary) |
| status of order #9931 (nonexistent) | "couldn't find it" (currently a global count) |
| orders in 2026 | not an id lookup |
| total revenue today | estimated, cancelled excluded |
| how many assigned orders today | must agree with the Orders page |
| and yesterday? (follow-up) | previous query, day shifted |
| how many tenants | tenant count |
| where is order DM-BK-123 | tracking |
| top 5 tenants by orders this week | breakdown + limit |
| unassigned orders right now | pending count |
| how many orders (account with >1000 bookings) | correct, or explicitly "at least N" |
---
# Part 6 — Capability levels
A different cut from the phases above: not *how* to build it, but *what the bot
could do*, ordered by how much has to exist underneath.
Baseline: `doormileApi.js` exports **96 functions, 37 of them reads**. The bot
calls **13** — all list endpoints. Every level below L6 is built from functions
that already exist and are already used by some page in this console.
---
## L0 — Foundation (not a feature; blocks everything)
Every count the bot gives is capped at `getBookings(1, 1000)` — page 1, at the
API's hard cap — and `getBookings` discards the envelope's `total`, so
truncation is undetectable. Fix pagination, surface `total`, route counts
through `GET /admin/reports`, and say "at least N" when truncated.
Until this lands, every level below inherits a silent wrong-number risk.
---
## L1 — Counts and lists — **SHIPPED**
25 intents over 13 API functions. Single-dimension questions, plus typo
tolerance, date vocabulary, comparisons, multi-part, and date follow-ups.
**Ceiling:** one filter at a time. "Delivered orders for Acme this week"
answers only the status.
---
## L2 — Composable queries
Slot filling replaces first-match-wins: `subject x range x status x batch x
tenant x rider x hub` all compose into one query.
- "Delivered orders for Acme this week"
- "Pending morning-batch orders at Coimbatore hub yesterday"
- "Cancelled orders for Acme vs Beta last month"
**New endpoints needed:** none. Coverage becomes the product of the slots
rather than a list of 25.
---
## L3 — Entity intelligence
Deep answers about *one* thing, using the detail endpoints the bot has never
touched.
| Subject | Functions available now | Unlocks |
|---|---|---|
| Order | `getBooking`, `getBookingTrack` | "Where is DM-BK-123", full status timeline, assigned rider, ETA |
| Parcel | `getConsignment`, `getConsignmentLogs`, `trackConsignment` | Scan history, exception trail |
| Rider | `getMiler`, `getMilerActivity`, `getMilerLogs`, `getMilerSummary` | "What has Suresh done today" — assigned/completed/rejected/cancelled, riderkms, last ping, live position |
| Hub / vehicle | `getHub`, `getVehicle` | Per-site detail, assigned fleet |
| Tenant | `getAdminTenant`, `getTenantLocations`, `getTenantCustomers` | Sites, customers, contact |
| Exception | `getException` | Why a delivery failed |
| Pricing | `quotePricing`, `simulatePricing` | "What would a 5kg parcel from 641001 to 600001 cost?" — a real calculation, not a lookup |
**Warning:** `riderLookup` today reads `found.status`, `found.phonenumber`,
`found.vehicletype`. The confirmed-live miler shape (documented in `api.js`)
has none of those — it carries `availabilitystatus`, `phone`,
`defaultvehicletype`. Fix against the real shape before extending this level.
---
## L4 — Analytics, ranking, anomaly
Built on `getReports` (`by_tenant` / `by_hub` / `by_rider` blocks),
`getLocationsSummary`, `getMilerSummary`.
- "Top 5 tenants by orders this week"
- "Which hub is busiest / which needs attention"
- "Which riders have the most cancellations"
- "Cancellation rate this week vs last"
- "Orders per hour today"
**"Which orders are delayed" is computable today** — `serviceoptions[0].
estimateddeliveryat` exists on a booking, so "past estimated delivery and not
yet delivered" is a real filter, not a guess. This is probably the single
highest-value question on the list and nothing currently answers it.
---
## L5 — Navigation and UI control
The assistant stops being a read-only oracle and starts driving the console.
- Every answer carries the filter state that produced it → "Open in Orders"
- "Show me cancelled orders" navigates and applies the filter, instead of
returning a count
- "Open order DM-BK-123" routes to the record
**Dependency:** the target pages must accept filter state from the URL. Check
what `orders.js` supports before committing to this.
---
## L6 — Write actions, confirm-gated
Out of scope per `CLAUDE.md` §4 until signed off, and that doc already fixes
the required shape: propose -> show the exact payload and affected rows ->
operator confirms -> execute -> refetch. Never straight from match to mutation.
Tiered by blast radius:
| Tier | Functions | Risk |
|---|---|---|
| T1 | `notifyMiler` | Sends a message. Reversible by sending another. |
| T2 | `assignMilerToBooking`, `assignVehicleToBooking`, `updateBookingStatus`, `cancelBooking`, `updateExceptionStatus`, `blockMiler` | Single record, real consequence |
| T3 | `batchAssignBookings`, `bulkCancelBookings` | Many records at once |
| T4 | `createTripsheet`, `addTripsheetItem`, `dispatchTripsheet`, `arriveTripsheet` | Multi-step workflow with ordering rules |
| T5 | `createExpressBooking`, `createExpressBookingBulk` | Last. CityGate pincode validation and delivery-slot windows live elsewhere and must not be bypassed. |
Must honour the dispatch reconcile rule (root `CLAUDE.md` §4) and the
notify-rider-after-mutation rule (§9) inside the executor, so the assistant
can't become a backdoor around either.
---
## L7 — Proactive / watch
Stops waiting to be asked. A watch loop evaluates threshold conditions and
pushes into the thread plus a badge on the trigger.
- "14 morning-batch orders unassigned, 30 minutes to cutoff"
- "Rider offline mid-route"
- "Hub with zero active riders"
**Dependency:** the notification panel in `AppTopNav` is currently static
scaffolding, not wired to a real alert stream.
---
## L8 — LLM as parser only — BLOCKED
`CLAUDE.md` §2 records the decision: no backend, nowhere to hold a key. If
that ever changes, the model does **slot extraction only** — text in, a
validated Query out. It never produces a number, a row, or a sentence read as
fact. Deterministic code still executes every fetch and every calculation, and
the L2 parser becomes the fallback.
---
## Suggested order
1. **L0** — one day, removes the wrong-number risk
2. **L4's delay detection** — highest value per unit of work, no new endpoints
3. **L2** — multiplies coverage, deletes code
4. **L3** — the 24 unused read functions
5. **L5** — makes answers actionable
6. **L6/L7** — only after sign-off