Behavision: face recognition for retail, edge to head office

Five components that ship as one product:

- behavision/  the recognition engine. RTSP ingest, YuNet detection, IoU
               tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
               FastAPI dashboard. Identity is decided once per TRACK from an
               average of at least three embeddings, never per frame.
- agent/       the Go edge agent: supervises the engine, holds a durable
               spool, and drains it to MQTT. Nothing is acked before the
               broker confirms.
- desktop/     the shop PC application (Wails + React + tray).
- server/      the cloud API, MQTT consumer, reports and assistant.
- web/         platform.loyaly.ai, the head-office app, embedded in the
               server binary.

The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.

CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
This commit is contained in:
2026-09-04 11:14:18 +05:30
commit dad04e8cda
216 changed files with 40473 additions and 0 deletions

102
config/default.yaml Normal file
View File

@@ -0,0 +1,102 @@
# Behavision configuration.
# ${VAR} placeholders are resolved from the environment (.env is loaded first).
app:
data_dir: data
models_dir: models
log_level: INFO
debug_faces: false # dump aligned chips to data/debug (diagnosis only)
# Write one face image per visit to data/outbox for the agent to upload.
# Off by default on purpose: with this off the machine holds no photographs,
# which is a data-protection position, not a missing feature.
store_faces: false
api:
host: 0.0.0.0
port: 8010
# HTTP Basic credentials for the dashboard and the whole JSON API.
# Leave blank and a credential is generated into
# data/api_credentials.txt on first boot (and logged) — a routable
# host is never served unauthenticated. Blank + host 127.0.0.1 is
# open, since it is unreachable from off-box.
username: ${BEHAVISION_API_USER}
password: ${BEHAVISION_API_PASSWORD}
cameras:
- id: cam1
# Either give a full `url` (must be percent-encoded yourself), or give
# parts below and the URL is built with proper encoding ('@' in the
# password is handled correctly).
url: ""
host: ${BEHAVISION_CAM1_HOST}
port: 554
path: /ch0_0.264
username: ${BEHAVISION_CAM1_USERNAME}
password: ${BEHAVISION_CAM1_PASSWORD}
# For quick testing without a camera, set `webcam: 0` to use a local
# webcam instead of RTSP.
webcam: null
# Per-camera overrides for the recognition gates. Anything left out uses
# the global `recognition:` block below. The gates describe a *view*, so
# an overhead corridor camera and an entrance camera at head height need
# different numbers — measure each with:
# python -m behavision calibrate --person NAME --seconds 25
# python -m behavision calibrate --report
# Quality is safe to loosen per camera (it only judges this view).
# match/enroll are not: every camera writes into one shared gallery, so a
# loose camera can merge two people into an identity a strict one trusts.
tuning:
min_enroll_quality: null
match_threshold: null
enroll_threshold: null
detection:
score_threshold: 0.82 # measured: frosted-glass false positives pass 0.75
nms_threshold: 0.3
min_face_px: 48 # ignore faces smaller than this (short side, px)
max_faces: 20
recognition:
# Cosine similarity on L2-normalised ArcFace embeddings.
match_threshold: 0.42 # >= this -> same person (higher = stricter)
enroll_threshold: 0.32 # < this -> safe to treat as a brand-new person
reinforce_threshold: 0.55
max_embeddings_per_identity: 5
auto_enroll: true
# Measured on this camera: real frontal faces score 0.70-0.82, glass
# blurs/silhouettes peak at 0.54 — 0.65 separates them cleanly.
min_enroll_quality: 0.65
sighting_cooldown_seconds: 30
tracking:
iou_threshold: 0.3
max_misses: 25 # frames a track survives without a detection
min_hits_for_id: 4 # frames before a track can be identified
min_embeddings_for_id: 3 # embeddings averaged before deciding identity
min_quality_to_encode: 0.35
max_id_attempts: 8
# Bounds how often an ambiguous track re-decides, not how often it
# encodes: embeddings still accumulate every frame, so the 8 attempts
# above span ~4s of genuinely different frames instead of ~0.3s.
id_retry_interval_seconds: 0.5
# Keep learning a person's other angles for the rest of their visit
# instead of freezing the identity on its first embedding.
reinforce_during_track: true
reinforce_interval_seconds: 1.0
attributes:
enabled: true # age / gender / emotion (needs optional models)
# Gate for collecting one age/gender/emotion sample. Separate from
# recognition.min_enroll_quality on purpose: that gate guards creating a
# permanent identity, this one only guards a measurement, and sharing it
# meant no track ever gathered the several samples the median needs.
min_quality: 0.35
events:
webhook_url: ${BEHAVISION_WEBHOOK_URL}
email:
smtp_host: ${BEHAVISION_SMTP_HOST}
smtp_port: ${BEHAVISION_SMTP_PORT}
username: ${BEHAVISION_SMTP_USER}
password: ${BEHAVISION_SMTP_PASSWORD}
to: ${BEHAVISION_SMTP_TO}
min_interval_seconds: 300