Three reported from a colleague's machine, plus one the fixing uncovered. Every one produced a message that was true and useless. ## behavision-setup chose the Python least likely to work findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a floor with NO ceiling, which is exactly backwards. The newest Python on a machine is the one least likely to have binary wheels. It picked 3.14, pip found no numpy wheel for cp314, fell back to building numpy from source and produced "Unknown compiler(s)"; once the operator had installed Xcode's command line tools to get past that, ten minutes of compiling ended in "<arm_neon.h> is intended only for ARM and AArch64 targets". maxMinor refuses in one line before anything is downloaded, and "too new" is a different message from "too old" - telling somebody holding Python 3.14 that no Python was found sends them to install a newer one, which is the direction that just failed. ## numpy<2.0 was the cap; OpenCV was the hazard Widening it needed proof, and the proof found something else. Nine runs of the detector guard per combination, one machine, one sitting: numpy 1.26 / cv2 4.11 9 passed, 0 crashed numpy 2.0 / cv2 4.11 8 passed, 1 crashed numpy 1.26 / cv2 4.14 3 passed, 6 crashed numpy 2.0 / cv2 4.14 2 passed, 7 crashed numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash was test_a_shared_detector_really_does_race, which races a shared cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11 usually turned it into an exception, 4.14 usually turns it into a segfault, and 4.11 crashing once says the hazard was always there. It never reached the product - Engine._build_worker builds a detector per camera. It reached the suite: two runs in three died with no failing assertion in them. The race runs in a subprocess now, and one clean attempt proves nothing, so the premise holds if any of several attempts misbehaves. 226 passed / 2 skipped on numpy 2.0.2, five runs of five. opencv stays capped below 5: everything above was measured on 4.x, and an uncapped >=4.8.1 gives every NEW install a major release this project has never run a real camera through. ## One MQTT client id for a whole shop, so two PCs fought over it behavision-<client>-<site> is the same string on every computer claimed to one site. MQTT requires unique client ids and a broker enforces it by disconnecting the older session, so the colleague's Mac and the shop's own till took turns kicking each other off: broker connected / broker connection lost: EOF / broker connected / EOF ... The damage is not confined to the new machine. The till is the other half of that loop, so signing in on a laptop to look at the product stops a live shop delivering visits - and from each end it reads as an unstable network. MQTTClientID() appends a per-installation id, minted on first load and written back so an existing install gets one without anybody doing anything. The site stays in the name because that is what a broker log is read by. An unwritable config falls back to a per-run id rather than a shared one. ## "no such file or directory" for an engine nobody had installed Pressing Start went straight to the supervisor, which reported what exec reported: a 200-character path ending in "no such file or directory". Every word true, none of it saying "run the setup tool" - the startup path had that sentence, in a log file nobody on a shop counter opens. engineMissing() is the one function the startup path, the Start button and the status panel all consult. It also names App Translocation, which was in that path and is unguessable: macOS runs a downloaded unsigned app from a random read-only copy, so relative paths resolve inside it and an install there would not survive a restart. The product is unsigned, so that is the normal first-run state on every Mac, not an edge case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Behavision
Production face recognition over RTSP. Watches camera streams, detects and tracks faces, recognizes known people, auto-enrolls new visitors, records visit history, and serves a live dashboard + JSON API.
Clean-room rewrite of the previous Camera/ and pattern_reg/ projects:
same core ideas, correct engineering.
Quick start (Windows)
cd D:\NEARLE\Behavision
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
python -m behavision setup-models # downloads YuNet, copies ArcFace etc. from the old project
python -m behavision run # dashboard at http://localhost:8010
Camera credentials live in .env (gitignored) — never in code or YAML.
So do the dashboard credentials: set BEHAVISION_API_USER and
BEHAVISION_API_PASSWORD, or let the server generate one into
data/api_credentials.txt on first boot. A routable api.host is never
served without HTTP Basic auth; 127.0.0.1 is left open.
To test without a camera, set webcam: 0 on a camera in
config/default.yaml.
Enroll a person by name from photos:
python -m behavision enroll --name "Alice" --images C:\photos\alice\
Architecture
behavision/
├── config.py typed config: YAML + ${ENV} expansion, validated (pydantic)
├── capture.py RTSP/webcam reader thread: latest-frame slot, TCP transport,
│ exponential-backoff reconnect, percent-encoded credentials
├── detection.py YuNet face detector (OpenCV) → boxes + 5 landmarks, clipped
├── recognition.py ArcFace ONNX encoder (correct (x-127.5)/127.5 RGB preprocessing,
│ unit-norm output) + clamped face-quality scoring
├── tracking.py IoU tracker: identity decided once per TRACK, not per frame
├── gallery/
│ ├── index.py FAISS IndexFlatIP (exact cosine) with identical numpy fallback
│ ├── store.py SQLite (WAL): identities, embeddings, sightings — source of truth
│ └── service.py three-zone matching: match / ambiguous(do nothing) / enroll
├── attributes.py optional age, gender, emotion on the aligned chip
├── events.py async event bus → log / webhook / rate-limited email sinks
├── engine.py one worker thread per camera, shared models + gallery
├── api.py FastAPI: dashboard, MJPEG stream, identities, events, stats
└── __main__.py CLI: run | enroll | setup-models
Pipeline
RTSP ──► capture ──► detect (YuNet) ──► track (IoU)
│ once per track, quality-gated
▼
align (Umeyama 5-pt) ──► ArcFace ──► cosine search
│
┌───────────────────────────┼──────────────────────────┐
sim ≥ 0.42 0.32 ≤ sim < 0.42 sim < 0.32
known person ambiguous → retry new visitor
sighting + event on a better frame auto-enroll + event
Design decisions (and the failure they prevent)
| Decision | Prevents |
|---|---|
Per-camera FaceDetector, shared thread-safe encoder |
cv2 input-size race between camera workers |
| HTTP Basic on every route, escaped dashboard output | open biometric API on the LAN; stored XSS via identity labels |
| Percent-encoded credentials, URL built from parts | @ in password silently breaking the stream (old bug) |
| Track-level identity, sighting cooldown | one user registered per frame (old bug) |
Exact IndexFlatIP on unit vectors, -1 guarded |
inverted L2 threshold + wrong-person metadata[-1] (old bugs) |
| SQLite as source of truth, index rebuilt at boot | index/metadata drift, untrained-IVF crash (old bugs) |
| Three-zone thresholds with ambiguous no-op | duplicate identities and wrong merges |
| One color conversion, ArcFace-native normalization | off-distribution embeddings making thresholds meaningless (old bug) |
| All quality terms clamped to [0,1] | unreachable registration threshold (old bug) |
| Readiness-guarded API, sinks off the hot path | startup crashes, notification stalls |
API
| Method | Path | Purpose |
|---|---|---|
| GET | / |
live dashboard |
| GET | /api/health, /api/stats |
liveness / metrics |
| GET | /api/cameras/{id}/stream.mjpeg |
annotated live stream |
| GET | /api/cameras/{id}/frame.jpg |
latest annotated frame |
| GET | /api/identities, /api/sightings, /api/events |
data |
| PATCH | /api/identities/{id} |
rename a visitor ({"label": "Alice"}) |
| DELETE | /api/identities/{id} |
forget a person (embeddings removed) |
Tests
pip install pytest
pytest tests -q
Configuration
Everything lives in config/default.yaml; ${VAR} placeholders resolve
from the environment (.env is loaded first). Thresholds:
recognition.match_threshold(default 0.42): raise for fewer false matches, lower for fewer duplicates.recognition.min_enroll_quality(0.65): how good a face must look (sharpness, size, lighting, frontality) before a new identity is minted.tracking.min_hits_for_id(4): frames a face must persist before we spend an embedding on it — filters passers-by and phantom detections.cameras[].max_width(1280): frames are downscaled at ingest — full 3MP streams waste memory and detector time.
Recognition models
The encoder picks the first usable model in models/:
arcface_int8.onnx → w600k_mbf.onnx (MobileFaceNet, 13 MB, downloaded
automatically) → arcface.onnx (r100, 260 MB, copied from the old project;
needs ~1.5 GB free RAM to load). Pin one with recognition.model_file.
Every stored embedding is tagged with the model that produced it, and only
embeddings from the active model are searched — different encoders'
vectors are numerically incompatible and never mix.