fix(dispatch): liveness-aware coverage check, escalation rate limit, real alert sink

Prod findings (2026-09-22): a backend retry sweep failed 1,001 bookings in
60s; the agent called each a "coverage gap" because GEORADIUS found a miler
in the geo index (last seen in June), then sent 1,001 ops_alert tasks to
CUSTOMER_AGENT, which has no such handler.

- _find_zone filters GEORADIUS candidates by the backend's miler_status:<id>
  key; only status=Available counts. Facts now carry
  nearest_available_miler_within_km plus milers_in_geo_index_within_30km so
  the decision can separate "no riders here" from "riders exist, none on duty".
- Rate limit per zone per day: after the first alert, further failures only
  bump the counter (no LLM call); a summary re-alert goes out every
  DISPATCH_REALERT_EVERY (default 100).
- _ops_alert / _escalate_dispatch send EXCEPTION_DETECTED to JARVIS (the path
  that is actually handled); customer delay notice uses CUSTOMER_AGENT's real
  send_notification contract.
- JARVIS: escalation inbox (_escalations, pending_escalations()) and
  human_review/ops_alert task types are recorded instead of dropped.
- ExceptionAgent pull loops: also catch asyncio.TimeoutError (distinct from
  nats.errors.TimeoutError on 3.11) and log the exception type — the blank
  "pull loop error:" lines.
- Prompt + eval cases updated for the renamed facts; new case for the
  observed index-full/nobody-on-duty pattern. Tests for liveness filtering,
  burst suppression, fallback heuristic, sinks, and the JARVIS inbox.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
This commit is contained in:
2026-09-22 16:26:27 +05:30
parent 58bfa07385
commit 4283c602f6
8 changed files with 376 additions and 54 deletions

View File

@@ -208,6 +208,10 @@ class MasterAgent(Agent):
self._sub_agents: Dict[str, Agent] = {}
self._active_orders: Dict[str, Dict] = {}
self._decision_log: List[Dict] = []
# Human-review inbox: every escalation/ops alert a sub-agent raised.
# The only sink for proposals the agents aren't authorised to act on;
# surfaced by pending_escalations() for the Command Center.
self._escalations: List[Dict] = []
def register_sub_agent(self, agent: Agent):
self._sub_agents[agent.agent_id] = agent
@@ -218,7 +222,7 @@ class MasterAgent(Agent):
must decide) are logged at WARNING and recorded so they are visible
rather than silently dropped — the endpoint of the human-review path."""
if message.message_type == MessageType.EXCEPTION_DETECTED:
logger.warning(f"JARVIS: exception escalation from {message.sender}: {message.payload}")
self._record_escalation(message.sender, message.payload)
else:
logger.info(f"JARVIS: {message.message_type.value} from {message.sender}")
self._decision_log.append({
@@ -240,9 +244,30 @@ class MasterAgent(Agent):
return await self._handle_exception(task)
elif task.task_type == "generate_report":
return await self._generate_report(task)
elif task.task_type in ("human_review", "ops_alert"):
# Escalations sent as tasks land here instead of dead-lettering.
self._record_escalation(task.data.get("source", "unknown"), task.data)
return {"status": "recorded", "escalations_pending": len(self._escalations)}
else:
logger.warning(f"JARVIS: unknown task type {task.task_type!r} — dropped")
return {"status": "unknown_task", "task_type": task.task_type}
def _record_escalation(self, sender: str, payload: Dict[str, Any]):
logger.warning(f"JARVIS: escalation from {sender}: {payload}")
self._escalations.append({
"timestamp": datetime.now(),
"from": sender,
"exception_type": payload.get("exception_type"),
"severity": payload.get("severity"),
"payload": payload,
})
if len(self._escalations) > 500:
self._escalations = self._escalations[-500:]
def pending_escalations(self, limit: int = 50) -> List[Dict[str, Any]]:
"""Most recent escalations first — the human-review inbox."""
return list(reversed(self._escalations[-limit:]))
async def _orchestrate_order(self, task: AgentTask) -> Dict[str, Any]:
order_data = task.data.get("order", {})
order_id = order_data.get("order_id", "unknown")

View File

@@ -181,10 +181,14 @@ fallback failed, so no rider was assigned. Decide how to react. Choose exactly o
- "ops_alert": a genuine coverage gap in this zone (repeated failures, no nearby riders). Alert operations to onboard or redirect riders here.
- "escalate": ambiguous or contradictory — e.g. riders ARE nearby yet assignment keeps failing, which suggests a systemic issue rather than a coverage gap. A human dispatcher should look.
Weigh how many times assignment has failed in this zone today, how far the nearest available rider is,
the time of day, and whether coordinates were even available. A single failure with a rider nearby is
usually transient; repeated failures with no nearby rider is a coverage gap; repeated failures *despite*
nearby riders is not a coverage gap and warrants a human. Report an honest confidence in [0, 1]."""
Weigh how many times assignment has failed in this zone today, how far the nearest *available* rider is,
the time of day, and whether coordinates were even available. "nearest_available_miler_within_km" counts
only riders whose live status is Available; "milers_in_geo_index_within_30km" is the raw count of riders
ever seen nearby (it includes off-duty and stale entries), so a large index count with no available rider
means riders exist here but nobody is on duty — a staffing gap, not a systemic fault. A single failure with
an available rider nearby is usually transient; repeated failures with no available rider is a coverage
gap; repeated failures *despite* available riders nearby is not a coverage gap and warrants a human.
Report an honest confidence in [0, 1]."""
_ASSIGNMENT_SCHEMA = {
"type": "object",