Two changes, and the second was found by verifying the first. ## Watching a camera from the app, in another building Snapshots answer "is that camera working". They do not answer "what is happening in my shop right now", which is what somebody who opens the app away from the counter is asking. Head office's browser already had that answer - LiveHub plus cameras.Live, where the shop PC asks outbound whether anybody is watching and pushes JPEG frames for as long as somebody is - and the app could not reach it. cloud.CameraLive opens that feed and the app's own loopback relay re-emits it as multipart MJPEG. That is the trick: frames arrive base64 over SSE, an <img> cannot render that, and an <img> renders MJPEG natively - so a tile is an ordinary <img> pointed at loopback whether the camera is in this room or another city. - Reconnecting happens in the relay, not the page. The server caps one push at five minutes, so doing it here means the <img> never sees the stream end. - The headers are flushed before the first frame. Go writes them on the first body write, so without that the whole response waits for the shop PC to start pushing. Measured against production: 30 seconds and not even a Content-Type, which surfaces as the request timing out. - One camera at a time. Watching makes a shop PC upload, so a grid that went live at once would put an estate's worth of cameras on the wire because somebody opened a page. - live.mjpeg is behind the same per-run token as the engine routes, and a wrong token is a 404 that never reaches head office at all. - CameraLive uses its own HTTP client: the shared one's 30s timeout covers the whole response and would sever a working view every thirty seconds - the trap that made the server set WriteTimeout to zero for its own SSE endpoint. ## A camera read "Connected" for 34 minutes after the shop PC went blind Which is why the verification above looked like a failure: head office registered the viewer and no frame ever came. reportWith returns early when the engine is unreachable - correctly, it has nothing to say - so the last state it sent stays in the database looking current. Measured live: cam2 and entrance both reading Connected, in green, with last_seen_at 34 minutes old, while the heartbeat from the same PC said cameras_up 0 of 0. Two surfaces reading two stored fields and disagreeing. false could not be the answer. It means "this camera is not connecting", which sends an installer to check cabling on a camera that was working perfectly the last time anybody could ask it. So there are four states and one function: connected reported recently, and working not_connecting reported recently, and the stream will not open waiting no shop PC has ever reported this camera stale reported once, and not lately - Connected is CLEARED when stale or waiting. A stale true left in place stays available to every client reading the field directly, and leaves two fields on one object disagreeing - how the shops screen once came out labelled Working, in green, above "2 of 3 cameras not connecting". - Computed in scanCamera, so every camera anybody reads passes through it. A state computed per handler is one a handler forgets, and this had already reached three screens. - CameraStaleAfter is 5 minutes: five missed reports, not one. Same reasoning as three missed heartbeats - an indicator that cries wolf gets ignored. - An unparseable last_seen_at is stale. It should be impossible, which is why it must not fall through to the state that says everything is fine. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Behavision
Production face recognition over RTSP. Watches camera streams, detects and tracks faces, recognizes known people, auto-enrolls new visitors, records visit history, and serves a live dashboard + JSON API.
Clean-room rewrite of the previous Camera/ and pattern_reg/ projects:
same core ideas, correct engineering.
Quick start (Windows)
cd D:\NEARLE\Behavision
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
python -m behavision setup-models # downloads YuNet, copies ArcFace etc. from the old project
python -m behavision run # dashboard at http://localhost:8010
Camera credentials live in .env (gitignored) — never in code or YAML.
So do the dashboard credentials: set BEHAVISION_API_USER and
BEHAVISION_API_PASSWORD, or let the server generate one into
data/api_credentials.txt on first boot. A routable api.host is never
served without HTTP Basic auth; 127.0.0.1 is left open.
To test without a camera, set webcam: 0 on a camera in
config/default.yaml.
Enroll a person by name from photos:
python -m behavision enroll --name "Alice" --images C:\photos\alice\
Architecture
behavision/
├── config.py typed config: YAML + ${ENV} expansion, validated (pydantic)
├── capture.py RTSP/webcam reader thread: latest-frame slot, TCP transport,
│ exponential-backoff reconnect, percent-encoded credentials
├── detection.py YuNet face detector (OpenCV) → boxes + 5 landmarks, clipped
├── recognition.py ArcFace ONNX encoder (correct (x-127.5)/127.5 RGB preprocessing,
│ unit-norm output) + clamped face-quality scoring
├── tracking.py IoU tracker: identity decided once per TRACK, not per frame
├── gallery/
│ ├── index.py FAISS IndexFlatIP (exact cosine) with identical numpy fallback
│ ├── store.py SQLite (WAL): identities, embeddings, sightings — source of truth
│ └── service.py three-zone matching: match / ambiguous(do nothing) / enroll
├── attributes.py optional age, gender, emotion on the aligned chip
├── events.py async event bus → log / webhook / rate-limited email sinks
├── engine.py one worker thread per camera, shared models + gallery
├── api.py FastAPI: dashboard, MJPEG stream, identities, events, stats
└── __main__.py CLI: run | enroll | setup-models
Pipeline
RTSP ──► capture ──► detect (YuNet) ──► track (IoU)
│ once per track, quality-gated
▼
align (Umeyama 5-pt) ──► ArcFace ──► cosine search
│
┌───────────────────────────┼──────────────────────────┐
sim ≥ 0.42 0.32 ≤ sim < 0.42 sim < 0.32
known person ambiguous → retry new visitor
sighting + event on a better frame auto-enroll + event
Design decisions (and the failure they prevent)
| Decision | Prevents |
|---|---|
Per-camera FaceDetector, shared thread-safe encoder |
cv2 input-size race between camera workers |
| HTTP Basic on every route, escaped dashboard output | open biometric API on the LAN; stored XSS via identity labels |
| Percent-encoded credentials, URL built from parts | @ in password silently breaking the stream (old bug) |
| Track-level identity, sighting cooldown | one user registered per frame (old bug) |
Exact IndexFlatIP on unit vectors, -1 guarded |
inverted L2 threshold + wrong-person metadata[-1] (old bugs) |
| SQLite as source of truth, index rebuilt at boot | index/metadata drift, untrained-IVF crash (old bugs) |
| Three-zone thresholds with ambiguous no-op | duplicate identities and wrong merges |
| One color conversion, ArcFace-native normalization | off-distribution embeddings making thresholds meaningless (old bug) |
| All quality terms clamped to [0,1] | unreachable registration threshold (old bug) |
| Readiness-guarded API, sinks off the hot path | startup crashes, notification stalls |
API
| Method | Path | Purpose |
|---|---|---|
| GET | / |
live dashboard |
| GET | /api/health, /api/stats |
liveness / metrics |
| GET | /api/cameras/{id}/stream.mjpeg |
annotated live stream |
| GET | /api/cameras/{id}/frame.jpg |
latest annotated frame |
| GET | /api/identities, /api/sightings, /api/events |
data |
| PATCH | /api/identities/{id} |
rename a visitor ({"label": "Alice"}) |
| DELETE | /api/identities/{id} |
forget a person (embeddings removed) |
Tests
pip install pytest
pytest tests -q
Configuration
Everything lives in config/default.yaml; ${VAR} placeholders resolve
from the environment (.env is loaded first). Thresholds:
recognition.match_threshold(default 0.42): raise for fewer false matches, lower for fewer duplicates.recognition.min_enroll_quality(0.65): how good a face must look (sharpness, size, lighting, frontality) before a new identity is minted.tracking.min_hits_for_id(4): frames a face must persist before we spend an embedding on it — filters passers-by and phantom detections.cameras[].max_width(1280): frames are downscaled at ingest — full 3MP streams waste memory and detector time.
Recognition models
The encoder picks the first usable model in models/:
arcface_int8.onnx → w600k_mbf.onnx (MobileFaceNet, 13 MB, downloaded
automatically) → arcface.onnx (r100, 260 MB, copied from the old project;
needs ~1.5 GB free RAM to load). Pin one with recognition.model_file.
Every stored embedding is tagged with the model that produced it, and only
embeddings from the active model are searched — different encoders'
vectors are numerically incompatible and never mix.