Files
Behavision/README.md
Suriyakumarvijayanayagam dad04e8cda Behavision: face recognition for retail, edge to head office
Five components that ship as one product:

- behavision/  the recognition engine. RTSP ingest, YuNet detection, IoU
               tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
               FastAPI dashboard. Identity is decided once per TRACK from an
               average of at least three embeddings, never per frame.
- agent/       the Go edge agent: supervises the engine, holds a durable
               spool, and drains it to MQTT. Nothing is acked before the
               broker confirms.
- desktop/     the shop PC application (Wails + React + tray).
- server/      the cloud API, MQTT consumer, reports and assistant.
- web/         platform.loyaly.ai, the head-office app, embedded in the
               server binary.

The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.

CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 11:14:18 +05:30

128 lines
6.0 KiB
Markdown

# Behavision
Production face recognition over RTSP. Watches camera streams, detects and
tracks faces, recognizes known people, auto-enrolls new visitors, records
visit history, and serves a live dashboard + JSON API.
Clean-room rewrite of the previous `Camera/` and `pattern_reg/` projects:
same core ideas, correct engineering.
## Quick start (Windows)
```powershell
cd D:\NEARLE\Behavision
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
python -m behavision setup-models # downloads YuNet, copies ArcFace etc. from the old project
python -m behavision run # dashboard at http://localhost:8010
```
Camera credentials live in `.env` (gitignored) — never in code or YAML.
So do the dashboard credentials: set `BEHAVISION_API_USER` and
`BEHAVISION_API_PASSWORD`, or let the server generate one into
`data/api_credentials.txt` on first boot. A routable `api.host` is never
served without HTTP Basic auth; `127.0.0.1` is left open.
To test without a camera, set `webcam: 0` on a camera in
`config/default.yaml`.
Enroll a person by name from photos:
```powershell
python -m behavision enroll --name "Alice" --images C:\photos\alice\
```
## Architecture
```
behavision/
├── config.py typed config: YAML + ${ENV} expansion, validated (pydantic)
├── capture.py RTSP/webcam reader thread: latest-frame slot, TCP transport,
│ exponential-backoff reconnect, percent-encoded credentials
├── detection.py YuNet face detector (OpenCV) → boxes + 5 landmarks, clipped
├── recognition.py ArcFace ONNX encoder (correct (x-127.5)/127.5 RGB preprocessing,
│ unit-norm output) + clamped face-quality scoring
├── tracking.py IoU tracker: identity decided once per TRACK, not per frame
├── gallery/
│ ├── index.py FAISS IndexFlatIP (exact cosine) with identical numpy fallback
│ ├── store.py SQLite (WAL): identities, embeddings, sightings — source of truth
│ └── service.py three-zone matching: match / ambiguous(do nothing) / enroll
├── attributes.py optional age, gender, emotion on the aligned chip
├── events.py async event bus → log / webhook / rate-limited email sinks
├── engine.py one worker thread per camera, shared models + gallery
├── api.py FastAPI: dashboard, MJPEG stream, identities, events, stats
└── __main__.py CLI: run | enroll | setup-models
```
### Pipeline
```
RTSP ──► capture ──► detect (YuNet) ──► track (IoU)
│ once per track, quality-gated
▼
align (Umeyama 5-pt) ──► ArcFace ──► cosine search
│
┌───────────────────────────┼──────────────────────────┐
sim ≥ 0.42 0.32 ≤ sim < 0.42 sim < 0.32
known person ambiguous → retry new visitor
sighting + event on a better frame auto-enroll + event
```
### Design decisions (and the failure they prevent)
| Decision | Prevents |
|---|---|
| Per-camera `FaceDetector`, shared thread-safe encoder | cv2 input-size race between camera workers |
| HTTP Basic on every route, escaped dashboard output | open biometric API on the LAN; stored XSS via identity labels |
| Percent-encoded credentials, URL built from parts | `@` in password silently breaking the stream (old bug) |
| Track-level identity, sighting cooldown | one user registered per frame (old bug) |
| Exact `IndexFlatIP` on unit vectors, `-1` guarded | inverted L2 threshold + wrong-person `metadata[-1]` (old bugs) |
| SQLite as source of truth, index rebuilt at boot | index/metadata drift, untrained-IVF crash (old bugs) |
| Three-zone thresholds with ambiguous no-op | duplicate identities *and* wrong merges |
| One color conversion, ArcFace-native normalization | off-distribution embeddings making thresholds meaningless (old bug) |
| All quality terms clamped to [0,1] | unreachable registration threshold (old bug) |
| Readiness-guarded API, sinks off the hot path | startup crashes, notification stalls |
## API
| Method | Path | Purpose |
|---|---|---|
| GET | `/` | live dashboard |
| GET | `/api/health`, `/api/stats` | liveness / metrics |
| GET | `/api/cameras/{id}/stream.mjpeg` | annotated live stream |
| GET | `/api/cameras/{id}/frame.jpg` | latest annotated frame |
| GET | `/api/identities`, `/api/sightings`, `/api/events` | data |
| PATCH | `/api/identities/{id}` | rename a visitor (`{"label": "Alice"}`) |
| DELETE | `/api/identities/{id}` | forget a person (embeddings removed) |
## Tests
```powershell
pip install pytest
pytest tests -q
```
## Configuration
Everything lives in `config/default.yaml`; `${VAR}` placeholders resolve
from the environment (`.env` is loaded first). Thresholds:
- `recognition.match_threshold` (default 0.42): raise for fewer false
matches, lower for fewer duplicates.
- `recognition.min_enroll_quality` (0.65): how good a face must look
(sharpness, size, lighting, frontality) before a new identity is minted.
- `tracking.min_hits_for_id` (4): frames a face must persist before we
spend an embedding on it — filters passers-by and phantom detections.
- `cameras[].max_width` (1280): frames are downscaled at ingest — full
3MP streams waste memory and detector time.
## Recognition models
The encoder picks the first usable model in `models/`:
`arcface_int8.onnx` → `w600k_mbf.onnx` (MobileFaceNet, 13 MB, downloaded
automatically) → `arcface.onnx` (r100, 260 MB, copied from the old project;
needs ~1.5 GB free RAM to load). Pin one with `recognition.model_file`.
Every stored embedding is tagged with the model that produced it, and only
embeddings from the active model are searched — different encoders'
vectors are numerically incompatible and never mix.