The previous commit treated a missing miler_status:<id> key as "not available". In production only 2 of 34 milers have that key at all, so the agent would have reported "no available rider" for nearly every failure — a confident wrong answer in the opposite direction from the bug it fixed. - _miler_presence returns available / unavailable / unknown. No key, an unparseable value, or an unrecognised status reads as unknown. - _find_zone prefers a confirmed-available miler, otherwise reports the nearest unknown-presence one (it may well be assignable), and returns None only when every nearby candidate is confirmed off duty. - Facts carry nearest_miler_within_km + nearest_miler_presence; the prompt states plainly that unknown presence is not evidence of a coverage gap and should lean to monitor/escalate rather than ops_alert. - New eval case for riders-nearby-but-no-presence-data; tests for all three presence states. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
32 lines
1.1 KiB
Python
32 lines
1.1 KiB
Python
"""
|
|
Offline eval for the DispatchAgent assignment-failure decision
|
|
(core.llm.decide_assignment_failure).
|
|
|
|
Same scoring as the stall eval (see evals/_harness.py): acceptable-rate is the
|
|
headline, exact-rate secondary. The cases use the real gatherer keys
|
|
(zone_id, failures_today, nearest_miler_within_km, nearest_miler_presence,
|
|
milers_in_geo_index_within_30km, has_coordinates) so the eval
|
|
reflects what DispatchAgent actually sends.
|
|
|
|
Usage:
|
|
export ANTHROPIC_API_KEY=...
|
|
python -m evals.assignment_eval --runs 3
|
|
python -m evals.assignment_eval --dry-run
|
|
python -m evals.assignment_eval --model claude-haiku-4-5 --min-pass-rate 0.9
|
|
"""
|
|
from pathlib import Path
|
|
|
|
import core.llm as llm
|
|
from core.llm import build_assignment_failure_context, decide_assignment_failure
|
|
from evals._harness import case_now, run_cli
|
|
|
|
CASES_PATH = Path(__file__).with_name("assignment_cases.jsonl")
|
|
|
|
|
|
def to_context(case):
|
|
return build_assignment_failure_context(case.get("facts", {}), now=case_now(case))
|
|
|
|
|
|
if __name__ == "__main__":
|
|
run_cli(CASES_PATH, to_context, decide_assignment_failure, llm)
|