Commit Graph

12 Commits

Author SHA1 Message Date
aa3bb35733 fix(deploy): pass NATS_HOST/NATS_PORT and autonomy gates through compose
The agents connect with NATS_HOST/NATS_PORT (NATS_URL is informational), but
docker-compose only forwarded NATS_URL. That was masked while system_config
carried the real host as a hardcoded default; after the secrets scrub the
default is localhost, so a deploy would have sent every NATS connection —
message bus, DispatchAgent, ExceptionAgent — to localhost.

Also forwards the autonomy gates (all default false) plus DISPATCH_REALERT_EVERY
and LLM_MODEL, so production behaviour is set in .env rather than by code
defaults. .env.example updated to match.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
2026-09-22 16:43:51 +05:30
008df49605 fix(dispatch): bucket failures by geo cell while the backend omits zone_id
Every booking.assignment_failed observed in production carries no zone_id, so
the per-zone failure counter and the once-a-day alert limit collapsed into a
single global "unknown" bucket — one alert per day for the whole country.

zone_key() prefers the backend's zone_id and falls back to a coarse lat/lon
grid cell (1 dp, ~11 km, DISPATCH_GEO_BUCKET_DP). The failure reason is now
logged too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
2026-09-22 16:42:15 +05:30
5159c8d1a7 fix(dispatch): presence is three-state — absent status key is unknown, not off-duty
The previous commit treated a missing miler_status:<id> key as "not available".
In production only 2 of 34 milers have that key at all, so the agent would have
reported "no available rider" for nearly every failure — a confident wrong
answer in the opposite direction from the bug it fixed.

- _miler_presence returns available / unavailable / unknown. No key, an
  unparseable value, or an unrecognised status reads as unknown.
- _find_zone prefers a confirmed-available miler, otherwise reports the nearest
  unknown-presence one (it may well be assignable), and returns None only when
  every nearby candidate is confirmed off duty.
- Facts carry nearest_miler_within_km + nearest_miler_presence; the prompt
  states plainly that unknown presence is not evidence of a coverage gap and
  should lean to monitor/escalate rather than ops_alert.
- New eval case for riders-nearby-but-no-presence-data; tests for all three
  presence states.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
2026-09-22 16:39:33 +05:30
4283c602f6 fix(dispatch): liveness-aware coverage check, escalation rate limit, real alert sink
Prod findings (2026-09-22): a backend retry sweep failed 1,001 bookings in
60s; the agent called each a "coverage gap" because GEORADIUS found a miler
in the geo index (last seen in June), then sent 1,001 ops_alert tasks to
CUSTOMER_AGENT, which has no such handler.

- _find_zone filters GEORADIUS candidates by the backend's miler_status:<id>
  key; only status=Available counts. Facts now carry
  nearest_available_miler_within_km plus milers_in_geo_index_within_30km so
  the decision can separate "no riders here" from "riders exist, none on duty".
- Rate limit per zone per day: after the first alert, further failures only
  bump the counter (no LLM call); a summary re-alert goes out every
  DISPATCH_REALERT_EVERY (default 100).
- _ops_alert / _escalate_dispatch send EXCEPTION_DETECTED to JARVIS (the path
  that is actually handled); customer delay notice uses CUSTOMER_AGENT's real
  send_notification contract.
- JARVIS: escalation inbox (_escalations, pending_escalations()) and
  human_review/ops_alert task types are recorded instead of dropped.
- ExceptionAgent pull loops: also catch asyncio.TimeoutError (distinct from
  nats.errors.TimeoutError on 3.11) and log the exception type — the blank
  "pull loop error:" lines.
- Prompt + eval cases updated for the renamed facts; new case for the
  observed index-full/nobody-on-duty pattern. Tests for liveness filtering,
  burst suppression, fallback heuristic, sinks, and the JARVIS inbox.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
2026-09-22 16:26:27 +05:30
58bfa07385 fix: tool registry emits Anthropic schema, question lifecycle cleanup
- Tool.to_schema() emits input_schema (Anthropic Messages API) instead
  of OpenAI-style parameters
- QuestionManager: remove questions from _pending on answer and timeout,
  reject double answers, add list_pending()
- core/skills/__init__.py so skills are importable as a package
- revert optional-import fallbacks in logger/message_bus: nats-py and
  loguru are hard requirements, a broken install should fail loudly
- tests for wait_for_answer happy path and timeout cleanup

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
2026-09-22 15:52:06 +05:30
be8103c1d2 security: remove hardcoded credentials, untrack .env
- config/system_config.py: defaults are localhost with empty credentials;
  real hosts/secrets must come from env (.env or docker-compose)
- main.py: help text lists env var names instead of real NATS host/user
- doormile_test.py: reads infra config from env instead of literals
- untrack .env, ignore .env/.env.*, add .env.example with keys only
- pytest.ini: testpaths=tests so doormile_test.py isn't collected

Credentials remain in git history and must be rotated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AJLYcbTHCe45fyFnMfEin
2026-09-22 15:52:06 +05:30
7227d2d2bd updates on the agent 2026-09-22 15:29:39 +05:30
90cdb57a38 route optimizer agent with dailygrubs ai assign 2026-09-01 13:48:58 +05:30
Suriyakumarvijayanayagam
aa4d4c6549 feat: ExpressDispatchAgent — tenant-scoped batch assignment + sequencing
New agent that owns the DoormileExpress dispatch flow. The normal B2C booking
flow is untouched; this handles express orders that a console operator has let
accumulate batch by batch, then dispatches in one go.

Triggered by express.dispatch_requested (the console's manual dispatch action).
For the tenant it: reads the pending bookings and the tenant's available riders
(Go internal API), distributes them across riders by proximity + load with a
radius guard, asks routes.workolik.com (Valhalla) for each rider's stop order,
and writes the assignments back through the Go internal API — which stays the
single writer of assignment state; the agent only decides who and what order.

- agents/express_dispatch_agent.py: the agent (mirrors DispatchAgent's NATS
  binding; pure consumer of the backend-owned EXPRESS stream).
- config/system_config.py: ROUTE_OPTIMIZER_URL.
- registered under JARVIS in main.py / agents package.
- tests: distribution (radius, load, cap, no-GPS fallback) and sequencing
  (single-stop, remap, optimizer-down, unknown-id) — deterministic, no network.

Gated off by default on the backend (EXPRESS_AGENT_ENABLED); EXPRESS_AGENT_
AUTONOMOUS=false runs it observe-only (logs the plan, writes nothing).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-08-17 19:33:36 +05:30
ad753be7a5 deploy: add Dockerfile and docker-compose for production 2026-07-01 13:08:42 +05:30
377ee710a7 refactor: DispatchAgent watcher, ExceptionAgent consumer fix, JARVIS cleanup 2026-07-01 13:07:54 +05:30
f49193ee73 Initial commit 2026-06-26 16:08:31 +05:30