The previous commit treated a missing miler_status:<id> key as "not available".
In production only 2 of 34 milers have that key at all, so the agent would have
reported "no available rider" for nearly every failure — a confident wrong
answer in the opposite direction from the bug it fixed.
- _miler_presence returns available / unavailable / unknown. No key, an
unparseable value, or an unrecognised status reads as unknown.
- _find_zone prefers a confirmed-available miler, otherwise reports the nearest
unknown-presence one (it may well be assignable), and returns None only when
every nearby candidate is confirmed off duty.
- Facts carry nearest_miler_within_km + nearest_miler_presence; the prompt
states plainly that unknown presence is not evidence of a coverage gap and
should lean to monitor/escalate rather than ops_alert.
- New eval case for riders-nearby-but-no-presence-data; tests for all three
presence states.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
Prod findings (2026-09-22): a backend retry sweep failed 1,001 bookings in
60s; the agent called each a "coverage gap" because GEORADIUS found a miler
in the geo index (last seen in June), then sent 1,001 ops_alert tasks to
CUSTOMER_AGENT, which has no such handler.
- _find_zone filters GEORADIUS candidates by the backend's miler_status:<id>
key; only status=Available counts. Facts now carry
nearest_available_miler_within_km plus milers_in_geo_index_within_30km so
the decision can separate "no riders here" from "riders exist, none on duty".
- Rate limit per zone per day: after the first alert, further failures only
bump the counter (no LLM call); a summary re-alert goes out every
DISPATCH_REALERT_EVERY (default 100).
- _ops_alert / _escalate_dispatch send EXCEPTION_DETECTED to JARVIS (the path
that is actually handled); customer delay notice uses CUSTOMER_AGENT's real
send_notification contract.
- JARVIS: escalation inbox (_escalations, pending_escalations()) and
human_review/ops_alert task types are recorded instead of dropped.
- ExceptionAgent pull loops: also catch asyncio.TimeoutError (distinct from
nats.errors.TimeoutError on 3.11) and log the exception type — the blank
"pull loop error:" lines.
- Prompt + eval cases updated for the renamed facts; new case for the
observed index-full/nobody-on-duty pattern. Tests for liveness filtering,
burst suppression, fallback heuristic, sinks, and the JARVIS inbox.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin