Prod findings (2026-09-22): a backend retry sweep failed 1,001 bookings in 60s; the agent called each a "coverage gap" because GEORADIUS found a miler in the geo index (last seen in June), then sent 1,001 ops_alert tasks to CUSTOMER_AGENT, which has no such handler. - _find_zone filters GEORADIUS candidates by the backend's miler_status:<id> key; only status=Available counts. Facts now carry nearest_available_miler_within_km plus milers_in_geo_index_within_30km so the decision can separate "no riders here" from "riders exist, none on duty". - Rate limit per zone per day: after the first alert, further failures only bump the counter (no LLM call); a summary re-alert goes out every DISPATCH_REALERT_EVERY (default 100). - _ops_alert / _escalate_dispatch send EXCEPTION_DETECTED to JARVIS (the path that is actually handled); customer delay notice uses CUSTOMER_AGENT's real send_notification contract. - JARVIS: escalation inbox (_escalations, pending_escalations()) and human_review/ops_alert task types are recorded instead of dropped. - ExceptionAgent pull loops: also catch asyncio.TimeoutError (distinct from nats.errors.TimeoutError on 3.11) and log the exception type — the blank "pull loop error:" lines. - Prompt + eval cases updated for the renamed facts; new case for the observed index-full/nobody-on-duty pattern. Tests for liveness filtering, burst suppression, fallback heuristic, sinks, and the JARVIS inbox. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
50 lines
1.7 KiB
Python
50 lines
1.7 KiB
Python
"""JARVIS is the only sink for escalations; nothing sent there may dead-letter."""
|
|
import pytest
|
|
|
|
from datetime import datetime
|
|
|
|
from core.agent import MasterAgent
|
|
from core.types import AgentMessage, AgentTask, MessageType, Priority
|
|
|
|
|
|
def make_jarvis():
|
|
j = MasterAgent.__new__(MasterAgent)
|
|
j.agent_id = "JARVIS"
|
|
j._escalations = []
|
|
j._decision_log = []
|
|
return j
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_exception_detected_message_lands_in_inbox():
|
|
j = make_jarvis()
|
|
await j.handle_message(AgentMessage(
|
|
message_id="m1", timestamp=datetime.now(),
|
|
sender="DISPATCH_AGENT", recipient="JARVIS",
|
|
message_type=MessageType.EXCEPTION_DETECTED,
|
|
payload={"exception_type": "coverage_gap", "severity": "medium", "zone_id": "z"},
|
|
))
|
|
inbox = j.pending_escalations()
|
|
assert len(inbox) == 1
|
|
assert inbox[0]["from"] == "DISPATCH_AGENT"
|
|
assert inbox[0]["exception_type"] == "coverage_gap"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_human_review_task_is_recorded_not_dropped():
|
|
j = make_jarvis()
|
|
task = AgentTask(task_id="t1", agent_type="master", task_type="human_review", priority=Priority.HIGH,
|
|
data={"source": "EXCEPTION_AGENT", "exception_type": "miler_stalled", "severity": "high"})
|
|
result = await j.handle_task(task)
|
|
assert result["status"] == "recorded"
|
|
assert j.pending_escalations()[0]["from"] == "EXCEPTION_AGENT"
|
|
|
|
|
|
@pytest.mark.asyncio
|
|
async def test_inbox_is_newest_first_and_capped():
|
|
j = make_jarvis()
|
|
for i in range(600):
|
|
j._record_escalation("X", {"exception_type": "t", "severity": "low", "n": i})
|
|
assert len(j._escalations) == 500
|
|
assert j.pending_escalations(limit=3)[0]["payload"]["n"] == 599
|