fix(dispatch): liveness-aware coverage check, escalation rate limit, real alert sink
Prod findings (2026-09-22): a backend retry sweep failed 1,001 bookings in 60s; the agent called each a "coverage gap" because GEORADIUS found a miler in the geo index (last seen in June), then sent 1,001 ops_alert tasks to CUSTOMER_AGENT, which has no such handler. - _find_zone filters GEORADIUS candidates by the backend's miler_status:<id> key; only status=Available counts. Facts now carry nearest_available_miler_within_km plus milers_in_geo_index_within_30km so the decision can separate "no riders here" from "riders exist, none on duty". - Rate limit per zone per day: after the first alert, further failures only bump the counter (no LLM call); a summary re-alert goes out every DISPATCH_REALERT_EVERY (default 100). - _ops_alert / _escalate_dispatch send EXCEPTION_DETECTED to JARVIS (the path that is actually handled); customer delay notice uses CUSTOMER_AGENT's real send_notification contract. - JARVIS: escalation inbox (_escalations, pending_escalations()) and human_review/ops_alert task types are recorded instead of dropped. - ExceptionAgent pull loops: also catch asyncio.TimeoutError (distinct from nats.errors.TimeoutError on 3.11) and log the exception type — the blank "pull loop error:" lines. - Prompt + eval cases updated for the renamed facts; new case for the observed index-full/nobody-on-duty pattern. Tests for liveness filtering, burst suppression, fallback heuristic, sinks, and the JARVIS inbox. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
This commit is contained in:
12
core/llm.py
12
core/llm.py
@@ -181,10 +181,14 @@ fallback failed, so no rider was assigned. Decide how to react. Choose exactly o
|
||||
- "ops_alert": a genuine coverage gap in this zone (repeated failures, no nearby riders). Alert operations to onboard or redirect riders here.
|
||||
- "escalate": ambiguous or contradictory — e.g. riders ARE nearby yet assignment keeps failing, which suggests a systemic issue rather than a coverage gap. A human dispatcher should look.
|
||||
|
||||
Weigh how many times assignment has failed in this zone today, how far the nearest available rider is,
|
||||
the time of day, and whether coordinates were even available. A single failure with a rider nearby is
|
||||
usually transient; repeated failures with no nearby rider is a coverage gap; repeated failures *despite*
|
||||
nearby riders is not a coverage gap and warrants a human. Report an honest confidence in [0, 1]."""
|
||||
Weigh how many times assignment has failed in this zone today, how far the nearest *available* rider is,
|
||||
the time of day, and whether coordinates were even available. "nearest_available_miler_within_km" counts
|
||||
only riders whose live status is Available; "milers_in_geo_index_within_30km" is the raw count of riders
|
||||
ever seen nearby (it includes off-duty and stale entries), so a large index count with no available rider
|
||||
means riders exist here but nobody is on duty — a staffing gap, not a systemic fault. A single failure with
|
||||
an available rider nearby is usually transient; repeated failures with no available rider is a coverage
|
||||
gap; repeated failures *despite* available riders nearby is not a coverage gap and warrants a human.
|
||||
Report an honest confidence in [0, 1]."""
|
||||
|
||||
_ASSIGNMENT_SCHEMA = {
|
||||
"type": "object",
|
||||
|
||||
Reference in New Issue
Block a user