Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
128 lines
6.0 KiB
Markdown
128 lines
6.0 KiB
Markdown
# Behavision
|
|
|
|
Production face recognition over RTSP. Watches camera streams, detects and
|
|
tracks faces, recognizes known people, auto-enrolls new visitors, records
|
|
visit history, and serves a live dashboard + JSON API.
|
|
|
|
Clean-room rewrite of the previous `Camera/` and `pattern_reg/` projects:
|
|
same core ideas, correct engineering.
|
|
|
|
## Quick start (Windows)
|
|
|
|
```powershell
|
|
cd D:\NEARLE\Behavision
|
|
python -m venv .venv
|
|
.venv\Scripts\activate
|
|
pip install -r requirements.txt
|
|
python -m behavision setup-models # downloads YuNet, copies ArcFace etc. from the old project
|
|
python -m behavision run # dashboard at http://localhost:8010
|
|
```
|
|
|
|
Camera credentials live in `.env` (gitignored) — never in code or YAML.
|
|
So do the dashboard credentials: set `BEHAVISION_API_USER` and
|
|
`BEHAVISION_API_PASSWORD`, or let the server generate one into
|
|
`data/api_credentials.txt` on first boot. A routable `api.host` is never
|
|
served without HTTP Basic auth; `127.0.0.1` is left open.
|
|
To test without a camera, set `webcam: 0` on a camera in
|
|
`config/default.yaml`.
|
|
|
|
Enroll a person by name from photos:
|
|
|
|
```powershell
|
|
python -m behavision enroll --name "Alice" --images C:\photos\alice\
|
|
```
|
|
|
|
## Architecture
|
|
|
|
```
|
|
behavision/
|
|
├── config.py typed config: YAML + ${ENV} expansion, validated (pydantic)
|
|
├── capture.py RTSP/webcam reader thread: latest-frame slot, TCP transport,
|
|
│ exponential-backoff reconnect, percent-encoded credentials
|
|
├── detection.py YuNet face detector (OpenCV) → boxes + 5 landmarks, clipped
|
|
├── recognition.py ArcFace ONNX encoder (correct (x-127.5)/127.5 RGB preprocessing,
|
|
│ unit-norm output) + clamped face-quality scoring
|
|
├── tracking.py IoU tracker: identity decided once per TRACK, not per frame
|
|
├── gallery/
|
|
│ ├── index.py FAISS IndexFlatIP (exact cosine) with identical numpy fallback
|
|
│ ├── store.py SQLite (WAL): identities, embeddings, sightings — source of truth
|
|
│ └── service.py three-zone matching: match / ambiguous(do nothing) / enroll
|
|
├── attributes.py optional age, gender, emotion on the aligned chip
|
|
├── events.py async event bus → log / webhook / rate-limited email sinks
|
|
├── engine.py one worker thread per camera, shared models + gallery
|
|
├── api.py FastAPI: dashboard, MJPEG stream, identities, events, stats
|
|
└── __main__.py CLI: run | enroll | setup-models
|
|
```
|
|
|
|
### Pipeline
|
|
|
|
```
|
|
RTSP ──► capture ──► detect (YuNet) ──► track (IoU)
|
|
│ once per track, quality-gated
|
|
▼
|
|
align (Umeyama 5-pt) ──► ArcFace ──► cosine search
|
|
│
|
|
┌───────────────────────────┼──────────────────────────┐
|
|
sim ≥ 0.42 0.32 ≤ sim < 0.42 sim < 0.32
|
|
known person ambiguous → retry new visitor
|
|
sighting + event on a better frame auto-enroll + event
|
|
```
|
|
|
|
### Design decisions (and the failure they prevent)
|
|
|
|
| Decision | Prevents |
|
|
|---|---|
|
|
| Per-camera `FaceDetector`, shared thread-safe encoder | cv2 input-size race between camera workers |
|
|
| HTTP Basic on every route, escaped dashboard output | open biometric API on the LAN; stored XSS via identity labels |
|
|
| Percent-encoded credentials, URL built from parts | `@` in password silently breaking the stream (old bug) |
|
|
| Track-level identity, sighting cooldown | one user registered per frame (old bug) |
|
|
| Exact `IndexFlatIP` on unit vectors, `-1` guarded | inverted L2 threshold + wrong-person `metadata[-1]` (old bugs) |
|
|
| SQLite as source of truth, index rebuilt at boot | index/metadata drift, untrained-IVF crash (old bugs) |
|
|
| Three-zone thresholds with ambiguous no-op | duplicate identities *and* wrong merges |
|
|
| One color conversion, ArcFace-native normalization | off-distribution embeddings making thresholds meaningless (old bug) |
|
|
| All quality terms clamped to [0,1] | unreachable registration threshold (old bug) |
|
|
| Readiness-guarded API, sinks off the hot path | startup crashes, notification stalls |
|
|
|
|
## API
|
|
|
|
| Method | Path | Purpose |
|
|
|---|---|---|
|
|
| GET | `/` | live dashboard |
|
|
| GET | `/api/health`, `/api/stats` | liveness / metrics |
|
|
| GET | `/api/cameras/{id}/stream.mjpeg` | annotated live stream |
|
|
| GET | `/api/cameras/{id}/frame.jpg` | latest annotated frame |
|
|
| GET | `/api/identities`, `/api/sightings`, `/api/events` | data |
|
|
| PATCH | `/api/identities/{id}` | rename a visitor (`{"label": "Alice"}`) |
|
|
| DELETE | `/api/identities/{id}` | forget a person (embeddings removed) |
|
|
|
|
## Tests
|
|
|
|
```powershell
|
|
pip install pytest
|
|
pytest tests -q
|
|
```
|
|
|
|
## Configuration
|
|
|
|
Everything lives in `config/default.yaml`; `${VAR}` placeholders resolve
|
|
from the environment (`.env` is loaded first). Thresholds:
|
|
|
|
- `recognition.match_threshold` (default 0.42): raise for fewer false
|
|
matches, lower for fewer duplicates.
|
|
- `recognition.min_enroll_quality` (0.65): how good a face must look
|
|
(sharpness, size, lighting, frontality) before a new identity is minted.
|
|
- `tracking.min_hits_for_id` (4): frames a face must persist before we
|
|
spend an embedding on it — filters passers-by and phantom detections.
|
|
- `cameras[].max_width` (1280): frames are downscaled at ingest — full
|
|
3MP streams waste memory and detector time.
|
|
|
|
## Recognition models
|
|
|
|
The encoder picks the first usable model in `models/`:
|
|
`arcface_int8.onnx` → `w600k_mbf.onnx` (MobileFaceNet, 13 MB, downloaded
|
|
automatically) → `arcface.onnx` (r100, 260 MB, copied from the old project;
|
|
needs ~1.5 GB free RAM to load). Pin one with `recognition.model_file`.
|
|
Every stored embedding is tagged with the model that produced it, and only
|
|
embeddings from the active model are searched — different encoders'
|
|
vectors are numerically incompatible and never mix.
|