# LogiFlow AI — What Was Actually Built Here **One-line summary:** a Python multi-agent "operations brain" that sits beside a production last-mile delivery platform (a Go backend called Doormile), watches real infrastructure events (rider GPS, booking assignments), uses an LLM to make judgment calls about operational exceptions, and acts on them through the backend's API — with careful autonomy gating, deterministic fallbacks, an offline eval suite, and a live mission-control dashboard. > ⚠️ **Commercially sensitive — fix before publishing this repo or portfolio:** `config/system_config.py` contains **hardcoded fallback credentials** (NATS, Redis, and Postgres passwords, an internal API key) and **real public server IPs**; `main.py`'s help text also prints the NATS host/user, and a `.env` file is present in the working tree. None of that is reproduced in this document, but it must be scrubbed and the credentials rotated before the repo is shared. The `docs/` handoff note and code comments also reveal internals of the Doormile backend (package names, NATS topology) — fine for an internal doc, worth reviewing before public posting. The codebase has two distinct layers, and an honest portfolio should distinguish them: - A **production integration layer** (the real engineering): the message bus, the ExceptionAgent stall pipeline, the DispatchAgent coverage-gap watcher, the LLM decision layer, the eval harness, the HTTP client, and the Command Center. These talk to real NATS/Redis/Postgres and are tested. - A **demo/simulation layer**: FleetAgent, HubAgent, RouteOptimizerAgent, and the two Streamlit UIs operate on hardcoded in-memory data (a fictional fleet across Indian cities). They demonstrate the agent framework but don't touch real infrastructure. --- ## 1. The Agent Runtime and Message Bus (`core/agent.py`, `core/message_bus.py`) **What it is.** The framework every agent runs on. Each agent is an async worker with its own task queue; agents talk to each other through a central message bus rather than calling each other directly. The problem this solves: in a logistics operation, many concerns (orders, dispatch, exceptions, customer comms) need to react to the same events without being tangled together. Decoupling them behind a bus means any agent can be restarted, replaced, or run in a different process without the others knowing. **How it's built.** `Agent` (abstract base) owns an `asyncio.Queue` of tasks and a processing loop: pull a task, dispatch to a handler, record success/failure, emit telemetry. `MasterAgent` ("JARVIS") is the orchestrator — it registers sub-agents, fans out order work, and is the sink for human-review escalations. `MessageBus` is a singleton with **two transport modes**: when connected, it publishes over **NATS JetStream** (durable, persisted streams); when not, it silently degrades to in-process dispatch so the whole system still runs on a laptop with no infrastructure. ``` publish(msg) ├─ NATS connected? ── JetStream publish │ direct msg → subject logistics.direct. (durable consumer per agent) │ broadcast → subject logistics. (durable consumer per type) └─ not connected ── local dispatch straight into the recipient's queue ``` **The approach.** Directed messages carrying a `task_type` are converted into `AgentTask`s and enqueued — so a message from another agent is processed by exactly the same code path as a locally submitted task. Messages without a `task_type` go to `handle_message()` as notifications. A comment in `deliver()` records the bug this design fixed: previously, directed messages "sat unread in the bus queue forever." If a recipient isn't present in the process, the message is retained in a pull queue rather than dropped. Task IDs are derived from the message's correlation ID, which threads one booking's journey across every agent that touches it. **Trade-offs.** Message history is an in-memory list capped at 1,000 entries (observability, not durability — durability is JetStream's job). The local-fallback mode has weaker delivery guarantees than JetStream, which is the price of a zero-infrastructure dev mode. --- ## 2. Miler Stall Detection and LLM-Gated Exception Handling (`agents/exception_agent.py`) — the flagship **What it is.** "Milers" are delivery riders. Sometimes a rider stops moving mid-delivery — traffic, a handoff, a breakdown, a dead phone. The customer is waiting either way. This agent detects the stall from GPS data, figures out *why it probably happened*, and chooses among four responses: wait, reassure the customer, reassign the booking to another rider, or escalate to a human dispatcher. The hard part isn't detection — it's that the right response is a **judgment call** (a rider parked at the delivery address is fine; a rider frozen on a highway for 50 minutes is not), and one of the possible actions (reassignment) is customer-visible and irreversible. **How it's built.** Detection runs on two independent paths that feed one decision pipeline: ``` Path A (event-driven): Path B (sweep, every 60s): NATS TRACKING stream Postgres: all active bookings miler.location.updated (pull-sub) ─┐ └─ for each, read Redis movement compare to last position in │ hash; if last GPS ping is Redis hash miler::movement │ stale ≥10 min ──────────────┐ if unmoved ≥10 min ─────────────┤ │ ▼ ▼ Redis SET NX "stall_notified:" ← dedup claim (6h TTL) │ winner only ▼ publish miler.stalled (JetStream) ▼ pull-consumer of miler.stalled Redis SET NX "stall_handled:" ← redelivery guard ▼ gather facts (Postgres booking phase/age, Redis GPS freshness, per-miler stall count today) ▼ Claude decides: wait │ notify_only │ reassign │ escalate ▼ wait → log notify → POST /internal/notify (Go API) reassign → gated: only if EXCEPTION_AGENT_AUTONOMOUS=true AND confidence ≥ 0.7, else escalate to JARVIS escalate → message to JARVIS (human review endpoint) ``` **The approach, step by step:** 1. **Why two detection paths?** The event path catches a rider whose phone keeps reporting the *same* position. The sweep path catches the opposite failure: a phone that stops reporting entirely (no events → the event path is blind). Together they cover both "frozen position" and "gone dark." 2. **Why the Redis `SET NX` claims?** Both paths can detect the same stall, and every ping after minute 10 re-detects it. The old design used an in-memory set (unbounded, lost on restart, racy). The replacement is an atomic Redis `SET NX` with a 6-hour TTL — bounded memory, restart-safe, and race-free between the two paths. There are *two* claims: one dedupes the *alert* (producer side), one guards *handling* (consumer side), because JetStream is at-least-once and a redelivered event must never re-run an irreversible reassign. Both claims deliberately **fail open** on Redis errors — a genuine stall must never be silently muted; a duplicate alert is the cheaper mistake. 3. **Context gathering before deciding.** The agent pulls real signals — booking phase (assigned vs. pickup scheduled), booking age, minutes since last GPS ping, position-unchanged duration, how many times this rider has stalled today — all read-only and best-effort. A thin-context decision still runs, and the prompt is written so low information skews toward escalation. 4. **The LLM decides, but with a blast-radius hierarchy.** The four actions are ordered by reversibility. The prompt explicitly instructs "prefer the least disruptive action that fits the evidence" and to report honest confidence. The only irreversible action is double-gated in *code*, not in the prompt: an env-var autonomy switch (default **off**) AND a confidence threshold. Below the bar, the model's proposal is packaged with its reasoning and confidence and escalated to a human via JARVIS — the AI proposes, the human disposes. 5. **Deterministic fallback.** If the LLM call fails entirely (network, refusal, truncation, bad JSON), the agent falls back to the pre-LLM behavior — reassign + notify — so detection never silently stops acting just because the model is down. 6. **Timezone discipline.** All stall math goes through `_utcnow()` (tz-aware UTC) and a parser that assumes naive timestamps are UTC — explicitly so pre-existing Redis entries stay comparable during rollout and container timezone can't corrupt the arithmetic. There are dedicated unit tests for exactly this. **Trade-offs.** The dedup TTL (6h) means a rider who stalls, recovers, and stalls again within the window won't re-alert — accepted to avoid alert storms. Exception records live in an in-memory dict (observability, not system of record — Postgres owns booking truth). --- ## 3. The Assignment-Failure Watcher (`agents/dispatch_agent.py` + `docs/handoff-assignment-failed-event.md`) **What it is.** When a new booking comes in, the Go backend tries to assign a rider (AI assignment plus a fallback). Sometimes both fail — nobody gets assigned and, previously, nothing happened except a synchronous flag. This agent consumes assignment events and answers a subtle operational question: **is this failure noise, a real delay, a coverage gap (not enough riders in this zone), or something systemically broken?** Each answer warrants a different response. **How it's built.** The agent binds durable JetStream consumers to the backend-owned `ASSIGNMENTS` stream. It is deliberately a **pure consumer** — a code comment documents the bug this fixed: creating its own stream over the backend's subjects raised a silent overlap error that left the consumer unbound. Instead it *discovers* which existing stream carries the subject (`find_stream_name_by_subject`) and binds to it, with a clear actionable log if no stream carries it. On `booking.assigned`, it forwards hub-preparation work. On `booking.assignment_failed`, the interesting path runs: ``` booking.assignment_failed (from Go backend) ▼ gather coverage facts: • Redis GEORADIUS on milers:locations at 10 / 20 / 30 km rings → "nearest rider within N km" or "none within 30km" • Redis INCR failed_assignment:: → failures-today counter (48h TTL) ▼ Claude decides: monitor │ notify_customer │ ops_alert │ escalate ▼ notify_customer is gated behind DISPATCH_AGENT_AUTONOMOUS (default off: it becomes an internal ops proposal instead of a customer message) escalate → JARVIS with the model's proposed action + confidence LLM unavailable → deterministic fallback: ops alert only if failures ≥ 3 today ``` **The approach.** The clever part is the *fact design*. The expanding GEORADIUS rings turn a geo query into a coarse feature the model can reason about ("nearest rider within 20km"), and the per-zone daily counter distinguishes one-off from chronic. The prompt encodes the key diagnostic asymmetry: repeated failures with **no** nearby rider = coverage gap → alert ops; repeated failures **despite** nearby riders = something is broken in the assignment system itself → escalate to a human. Missing coordinates degrade gracefully to "coverage unknown," which the prompt treats as a reason for caution. **A notable engineering artifact:** the consumer, decision logic, evals, and tests are all built, but the feature is currently **inert** — the Go backend doesn't publish the failure event yet. `docs/handoff-assignment-failed-event.md` is a precise cross-team handoff spec (exact subject, payload schema, stream change, verification steps, discovered by inspecting the running Go binary's symbols) telling the backend team the one thing they need to publish to light it up. That doc is good evidence of working across a language/team boundary. --- ## 4. The LLM Decision Layer (`core/llm.py`) **What it is.** A single, thin module that owns every model interaction, so agent code never touches the Anthropic SDK directly. It solves the problem of making free-form model output safe to act on in an automated pipeline. **The approach.** - **Constrained decisions, not open generation.** Every call uses structured output against a strict JSON Schema — the model must return `{action ∈ fixed enum, reasoning, confidence}`. The action is re-validated in Python even after schema enforcement. - **Fail-to-None contract.** Any failure — network error (one retry), refusal, `max_tokens` truncation, unparseable JSON, out-of-enum action — collapses to `None`, and every caller has a documented deterministic fallback. The model can therefore *only ever* improve behavior over the pre-LLM baseline, never break it. - **Async client, lazily constructed** so a missing API key can't break agent startup and a slow model call can't freeze the event loop. - **Prompt-building functions are shared with the evals** (`build_stall_context`, `build_assignment_failure_context`) so the offline eval scores *exactly* the prompt that runs in production — no drift between what's tested and what ships. - Model, effort, token budget, and timeout are all env-configurable; a comment records the operational lesson that adaptive thinking eats into `max_tokens`, so the budget must be generous or the JSON gets truncated. --- ## 5. The Offline Eval Harness (`evals/`) **What it is.** Before letting the model act autonomously, you need a number that says how good its judgment is. This is a small purpose-built eval framework: labeled scenario datasets (JSONL), a shared scoring harness, and one runner per decision type. **The approach.** - **Cases encode that judgment problems have multiple right answers.** Each case has an `ideal` action and an `acceptable` set (e.g., a 25-minute stall with no GPS ping for 25 minutes: ideal `escalate`, but `reassign` also acceptable). The headline metric is *acceptable-rate* — "was the decision defensible?" — with exact-match as secondary. This is a more honest metric design than forcing a single gold label onto ambiguous scenarios. - **Majority voting across N samples** per case (`--runs 3`) measures decision *stability*, not just a lucky single sample. Model errors are counted separately from wrong answers. - **Reproducibility:** cases pin a fixed date so only hour-of-day varies between runs; `--dry-run` prints every prompt without API calls; `--model` swaps models for comparison; `--min-pass-rate` turns the eval into a CI gate. - The 25 cases (15 stall, 10 assignment) cover the genuinely tricky territory: rush-hour ambiguity, a rider at the delivery door, thin overnight coverage, a repeat-staller pattern, a dead phone vs. a frozen position, and the riders-nearby-yet-failing systemic case. Several are labeled as using the *real* fact keys the production gatherer emits. - **The intended workflow is explicit in the docstrings:** collect real production cases into the JSONL, get the acceptable-rate trustworthy, and only then flip the autonomy env vars. The eval is the graduation exam for autonomy. --- ## 6. The JARVIS Command Center (`dashboard/command_center.py` + one static HTML page) **What it is.** A real-time mission-control web dashboard (port 8600) showing every agent's live status, every message crossing the bus, task throughput/failure counts, per-agent task-duration stats, and message rate — the answer to "what is this distributed swarm doing *right now*?" **How it's built.** ``` Agents ──(JetStream)── logistics.> ┐ Agents ──(plain NATS)── telemetry.> ┤→ Command Center process (FastAPI) │ in-memory aggregation (deques, counters) │ REST: /api/state, /api/agent/{id} └──→ WebSocket fan-out → browser UI ``` **The approach.** The key design decision is that the dashboard is a **read-only tap that cannot perturb the system it observes**: it uses plain (non-JetStream) subscriptions with no durable consumers and no acks, so it never steals or holds messages the real agents need. Symmetrically, on the producer side, telemetry is published on plain NATS — deliberately *not* JetStream — and is a no-op when disconnected, so observability can never accumulate in the persistent stream or affect agent behavior (the base `Agent` class emits state every 5 seconds and a lifecycle event per task). The server reconnects to NATS forever, marks agents offline after 15 seconds of telemetry silence, computes a rolling messages-per-minute rate from a timestamp deque, and every buffer is a bounded `deque` so it can run indefinitely. New browsers get a full state snapshot, then deltas. --- ## 7. Supporting Cast (briefer, and honestly labeled) - **HTTP client (`core/http_client.py`)** — one shared `aiohttp` session with connection pooling, and retry logic that encodes real HTTP semantics: 4xx returns `None` immediately (retrying a client error is pointless), 5xx retries with exponential backoff, and only 2xx/3xx yields a body. The docstring nails why: callers gate side effects on a non-None result, so "server rejected it" must never look like "it worked." - **OrderAgent** — order intake/validation/categorization (required fields, phone/pincode regex, weight limits), zone typing via pincode prefix (same prefix → last-mile; same region → hub-to-spoke; else hub-to-hub), persisting through the Go backend's CRM booking API and adopting the backend's booking ID as the canonical order ID. - **FleetAgent / HubAgent / RouteOptimizerAgent** — the simulation layer: an in-memory fleet with hub capacity accounting (with visible bug-fix comments about counter drift on double-release), and route estimation using **haversine distance with time-of-day traffic multipliers and a TTL route cache**. These demonstrate the framework, not production logic. - **Tests (`tests/`, ~40 cases)** — genuinely targeted at the failure modes, not happy paths: dedup claim semantics including fail-open on Redis errors, timezone parsing edge cases, HTTP retry matrices (5xx→2xx recovery, 4xx no-retry), LLM refusal/truncation/invalid-action handling, stream binding without creation, and message-bus delivery routing. All use hand-rolled async fakes (FakeRedis with real `SET NX` semantics, fake PG pools, fake Anthropic clients) rather than mocking frameworks — the fakes model the *behavior* being relied on. - **Deployment** — Dockerfile + docker-compose with resource limits, all secrets via env (the compose file does this correctly; it's the config file's *fallback defaults* that leak). --- ## The Stack - **Language:** Python 3.12, `asyncio` throughout - **Messaging:** NATS JetStream (durable streams, push + pull consumers, at-least-once delivery) - **State/coordination:** Redis (GEO radius queries, hashes for GPS state, atomic `SET NX` claims, daily counters with TTL) - **Database:** PostgreSQL via `asyncpg` (read-only against the backend's booking tables) - **AI:** Anthropic Claude via the async SDK — structured output (JSON Schema), adaptive thinking, configurable effort - **Web:** FastAPI + uvicorn + WebSockets (command center); Streamlit (admin dashboard & customer portal demos) - **HTTP:** aiohttp (pooled shared session) - **Integration target:** a Go backend ("Doormile") via internal REST API - **Testing:** `unittest` + `IsolatedAsyncioTestCase`; custom JSONL eval harness - **Deployment:** Docker, docker-compose --- ## What's Genuinely Hard Here 1. **Exactly-once *effects* on an at-least-once bus.** JetStream redelivers; two detection paths race; every GPS ping re-detects a stall. Getting from that to "the customer is reassigned at most once" required the two-tier atomic Redis claim design (alert claim + handled claim), with the deliberate fail-open asymmetry — duplicate alerts are cheaper than muted stalls. This is the classic distributed-systems dedup problem, solved correctly and unit-tested. 2. **Making an LLM safe to put in an actuation loop.** The whole shape of the safety story — actions as a closed enum with schema-enforced output, a reversibility hierarchy where only the irreversible action is code-gated behind autonomy-flag + confidence-threshold, human escalation carrying the model's proposal and reasoning, and a deterministic fallback so model downtime degrades to the old behavior instead of inaction. The gates live in code, not in the prompt — the prompt is advice, the gate is enforcement. 3. **Evaluating judgment, not correctness.** Recognizing that stall handling has no single right answer and designing the metric around a defensible-action set with majority-vote stability — then wiring the eval to consume the *identical* prompt-building code as production, and making autonomy contingent on the eval passing on real collected cases. That's a disciplined promote-to-autonomy pipeline in miniature. 4. **Being a good citizen in someone else's infrastructure.** The agents consume streams owned by a Go backend they don't control. The stream-discovery-instead-of-creation fix (with the documented overlap-error war story), the read-only no-ack dashboard tap, telemetry deliberately kept off JetStream, and the reverse-engineered handoff spec for the missing event all show real cross-system integration work — the unglamorous kind that actually breaks projects when done wrong. 5. **Dual detection with complementary blind spots.** The frozen-position event path and the gone-dark database sweep each catch what the other structurally cannot. Recognizing that "no data" is itself a signal requiring a separate mechanism is a nice piece of failure-mode thinking. --- *Notes for portfolio framing:* the git history and comments show honest iteration (bug-fix comments explaining *why* the old design was wrong — dead code after a `return`, unbound consumers, the in-memory dedup set), which reads well in an engineering narrative. And once more so it isn't missed: **rotate and remove the credentials in `config/system_config.py` and the committed `.env` before this repo goes anywhere public.**