Prod findings (2026-09-22): a backend retry sweep failed 1,001 bookings in 60s; the agent called each a "coverage gap" because GEORADIUS found a miler in the geo index (last seen in June), then sent 1,001 ops_alert tasks to CUSTOMER_AGENT, which has no such handler. - _find_zone filters GEORADIUS candidates by the backend's miler_status:<id> key; only status=Available counts. Facts now carry nearest_available_miler_within_km plus milers_in_geo_index_within_30km so the decision can separate "no riders here" from "riders exist, none on duty". - Rate limit per zone per day: after the first alert, further failures only bump the counter (no LLM call); a summary re-alert goes out every DISPATCH_REALERT_EVERY (default 100). - _ops_alert / _escalate_dispatch send EXCEPTION_DETECTED to JARVIS (the path that is actually handled); customer delay notice uses CUSTOMER_AGENT's real send_notification contract. - JARVIS: escalation inbox (_escalations, pending_escalations()) and human_review/ops_alert task types are recorded instead of dropped. - ExceptionAgent pull loops: also catch asyncio.TimeoutError (distinct from nats.errors.TimeoutError on 3.11) and log the exception type — the blank "pull loop error:" lines. - Prompt + eval cases updated for the renamed facts; new case for the observed index-full/nobody-on-duty pattern. Tests for liveness filtering, burst suppression, fallback heuristic, sinks, and the JARVIS inbox. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
32 lines
1.1 KiB
Python
32 lines
1.1 KiB
Python
"""
|
|
Offline eval for the DispatchAgent assignment-failure decision
|
|
(core.llm.decide_assignment_failure).
|
|
|
|
Same scoring as the stall eval (see evals/_harness.py): acceptable-rate is the
|
|
headline, exact-rate secondary. The cases use the real gatherer keys
|
|
(zone_id, failures_today, nearest_available_miler_within_km,
|
|
milers_in_geo_index_within_30km, has_coordinates) so the eval
|
|
reflects what DispatchAgent actually sends.
|
|
|
|
Usage:
|
|
export ANTHROPIC_API_KEY=...
|
|
python -m evals.assignment_eval --runs 3
|
|
python -m evals.assignment_eval --dry-run
|
|
python -m evals.assignment_eval --model claude-haiku-4-5 --min-pass-rate 0.9
|
|
"""
|
|
from pathlib import Path
|
|
|
|
import core.llm as llm
|
|
from core.llm import build_assignment_failure_context, decide_assignment_failure
|
|
from evals._harness import case_now, run_cli
|
|
|
|
CASES_PATH = Path(__file__).with_name("assignment_cases.jsonl")
|
|
|
|
|
|
def to_context(case):
|
|
return build_assignment_failure_context(case.get("facts", {}), now=case_now(case))
|
|
|
|
|
|
if __name__ == "__main__":
|
|
run_cli(CASES_PATH, to_context, decide_assignment_failure, llm)
|