I got this wrong first time. "Head office cannot show live video cheaply" conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs WebRTC and a TURN server. The shop PC is behind a router with no inbound route, so head office cannot pull the engine's MJPEG. It can answer the agent's outbound requests, which is the shape of everything else here: the server holds a poll open, the agent asks "is anyone watching?", and pushes JPEGs up for exactly as long as somebody is. Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at 640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds - 4 fps, as configured. The UI says "about 4 frames a second" rather than letting anyone conclude the camera stutters. Nothing is uploaded when nobody is looking, which is the whole cost argument: Publish returns false once the last viewer goes, interest lapses on a timer each viewer refreshes as it reads (so a closed tab stops the upload within seconds), one push is capped at five minutes, and the UI streams one camera at a time. LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell pushes nothing because nothing may be lost; here a dropped frame is the correct outcome, so each viewer has a one-slot buffer that is overwritten - the only frame worth having is the newest, and a queue would show an ever-growing delay behind the shop instead of dropping back to live. Ownership is proved once, before anything streams: the relay is keyed on a camera id, a hub does not know whose camera it holds, and a camera id is not a secret. Verified: another tenant gets 404, no session gets 401, and an agent cannot push into another site's camera. Also fixes a bug I introduced with it - the Live button was gated on `connected`, which is head office's last report and up to two minutes stale, so it hid itself during every reconnect. "Is that camera really down?" is exactly when somebody wants to look, and a hidden control says "you cannot" where the honest answer is "here is why". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Behavision agent (Go)
The half of the edge install that touches the network. The Python engine keeps the cameras and the models; this keeps the tray icon, the UI shell, the MQTT connection and the offline queue.
┌─ agent (Go) ───────────────────┐ ┌─ engine (Python) ────────┐
│ tray icon + WebView2 window │ │ RTSP capture │
│ supervises the engine process │───────▶│ YuNet / ArcFace / FAISS │
│ MQTT publish + offline spool │◀───────│ SQLite (biometric) │
│ S3 handoff, tenant config │ local │ localhost API + events │
└────────────────────────────────┘ HTTP └──────────────────────────┘
Why the split
Go cannot run ONNX, OpenCV or FAISS, so the engine stays Python and ships frozen. Go is here for what it is actually better at: a durable queue that survives a store's internet dropping, a supervised child process, and one language shared with the server so the MQTT contract has a single definition.
Why not a Windows service
A service runs in session 0 and cannot draw a tray icon — that is Windows session isolation, not a library limitation. Since the product is "user starts and stops it from the tray", the agent is a normal user-session process that spawns the engine as a child. That also means it never needs elevation at runtime: starting a child process does not, controlling a service does.
internal/engine is written so a service wrapper can be added later without
touching the supervision logic.
Layout
main.go entry point, mode dispatch
internal/engine start/stop/supervise the Python engine, health polling
internal/spool durable event queue (survives restart and outage)
internal/mqtt broker client, publishes from the spool
internal/config tenant identity, broker settings, credentials
frontend/ React UI served into the WebView
Build
go build ./... # agent alone
wails build # agent + frontend, once the UI is added