POST /api/customers and POST /api/visitors/{id}/merge. They ship together
because the first creates the need for the second: a customer typed in at
a counter has no face template, so when a camera sees that person later
the matcher has nothing to compare against and enrols them as somebody
new. That is the design working, not failing - and it means every
hand-created customer is a duplicate waiting to happen. Shipping the
create alone would manufacture duplicates into the state CLAUDE.md
already flags: "there is no merge endpoint server-side, so its
duplicates would be unrecoverable."
The number comes from clients.visitor_seq, taken exactly as RecordVisit
takes it. Two sources of visitor numbers that could disagree would be
worse than none: V-42 has to mean one person whichever way they arrived.
The label is the typed name, or "Visitor N" when they gave none - the
same string the engine writes, so a record created by hand is
indistinguishable from an enrolled one afterwards.
The merge is one transaction over FIVE tables, and the count is the
point. visits, purchases, visitor_embeddings, consents and
visitor_profiles all reference visitors ON DELETE CASCADE, so a table
this forgets to re-point is not an error - those rows are destroyed with
the source and nobody finds out until a customer's history is short.
visitor_profiles is UNIQUE on visitor_id, so the two cannot simply both
move and something has to win. Blanks on the survivor are filled from the
source and nothing it already holds is overwritten, which is exactly
right for the case this exists for: a hand-typed name and phone joining
the face that was recognised a week later.
Policies carried over from the edge gallery's merge, which had to settle
all of this once already: a human-assigned name outranks an auto
"Visitor N" whichever direction the operator merged; visit_count is
recomputed with COUNT(*) and never summed, because the stored counter may
be stale and the row count cannot be; first_seen_at takes the earlier of
the two, since it is one person and always was.
Two things that are this side's own:
- The source is deleted for real, not soft-deleted. A tombstone would
leave its number resolving to a record holding nothing, which reads as
"this customer exists and has never been here" - a worse answer than
"no such customer".
- The response names the RETIRED reference. Staff write V-42 on cards and
read it aloud; a merge that does not say which one stopped working
leaves somebody to discover it at a counter.
Manager and above, not staff. Apart from erasure this is the only
irreversible operation on a customer: two people welded together cannot
be separated, because nothing records which visit came from whom. It logs
at WARNING and writes an audit row for the same reason.
Also fixed while here: two s.Log.Printf calls - one of them mine, from
the password endpoint - that would panic on a nil logger. The package has
a nil-guarded s.logf and those were the only two not using it. The
password one sat in an error path no test reaches, which is exactly where
that bug waits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Behavision
Production face recognition over RTSP. Watches camera streams, detects and tracks faces, recognizes known people, auto-enrolls new visitors, records visit history, and serves a live dashboard + JSON API.
Clean-room rewrite of the previous Camera/ and pattern_reg/ projects:
same core ideas, correct engineering.
Quick start (Windows)
cd D:\NEARLE\Behavision
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
python -m behavision setup-models # downloads YuNet, copies ArcFace etc. from the old project
python -m behavision run # dashboard at http://localhost:8010
Camera credentials live in .env (gitignored) — never in code or YAML.
So do the dashboard credentials: set BEHAVISION_API_USER and
BEHAVISION_API_PASSWORD, or let the server generate one into
data/api_credentials.txt on first boot. A routable api.host is never
served without HTTP Basic auth; 127.0.0.1 is left open.
To test without a camera, set webcam: 0 on a camera in
config/default.yaml.
Enroll a person by name from photos:
python -m behavision enroll --name "Alice" --images C:\photos\alice\
Architecture
behavision/
├── config.py typed config: YAML + ${ENV} expansion, validated (pydantic)
├── capture.py RTSP/webcam reader thread: latest-frame slot, TCP transport,
│ exponential-backoff reconnect, percent-encoded credentials
├── detection.py YuNet face detector (OpenCV) → boxes + 5 landmarks, clipped
├── recognition.py ArcFace ONNX encoder (correct (x-127.5)/127.5 RGB preprocessing,
│ unit-norm output) + clamped face-quality scoring
├── tracking.py IoU tracker: identity decided once per TRACK, not per frame
├── gallery/
│ ├── index.py FAISS IndexFlatIP (exact cosine) with identical numpy fallback
│ ├── store.py SQLite (WAL): identities, embeddings, sightings — source of truth
│ └── service.py three-zone matching: match / ambiguous(do nothing) / enroll
├── attributes.py optional age, gender, emotion on the aligned chip
├── events.py async event bus → log / webhook / rate-limited email sinks
├── engine.py one worker thread per camera, shared models + gallery
├── api.py FastAPI: dashboard, MJPEG stream, identities, events, stats
└── __main__.py CLI: run | enroll | setup-models
Pipeline
RTSP ──► capture ──► detect (YuNet) ──► track (IoU)
│ once per track, quality-gated
▼
align (Umeyama 5-pt) ──► ArcFace ──► cosine search
│
┌───────────────────────────┼──────────────────────────┐
sim ≥ 0.42 0.32 ≤ sim < 0.42 sim < 0.32
known person ambiguous → retry new visitor
sighting + event on a better frame auto-enroll + event
Design decisions (and the failure they prevent)
| Decision | Prevents |
|---|---|
Per-camera FaceDetector, shared thread-safe encoder |
cv2 input-size race between camera workers |
| HTTP Basic on every route, escaped dashboard output | open biometric API on the LAN; stored XSS via identity labels |
| Percent-encoded credentials, URL built from parts | @ in password silently breaking the stream (old bug) |
| Track-level identity, sighting cooldown | one user registered per frame (old bug) |
Exact IndexFlatIP on unit vectors, -1 guarded |
inverted L2 threshold + wrong-person metadata[-1] (old bugs) |
| SQLite as source of truth, index rebuilt at boot | index/metadata drift, untrained-IVF crash (old bugs) |
| Three-zone thresholds with ambiguous no-op | duplicate identities and wrong merges |
| One color conversion, ArcFace-native normalization | off-distribution embeddings making thresholds meaningless (old bug) |
| All quality terms clamped to [0,1] | unreachable registration threshold (old bug) |
| Readiness-guarded API, sinks off the hot path | startup crashes, notification stalls |
API
| Method | Path | Purpose |
|---|---|---|
| GET | / |
live dashboard |
| GET | /api/health, /api/stats |
liveness / metrics |
| GET | /api/cameras/{id}/stream.mjpeg |
annotated live stream |
| GET | /api/cameras/{id}/frame.jpg |
latest annotated frame |
| GET | /api/identities, /api/sightings, /api/events |
data |
| PATCH | /api/identities/{id} |
rename a visitor ({"label": "Alice"}) |
| DELETE | /api/identities/{id} |
forget a person (embeddings removed) |
Tests
pip install pytest
pytest tests -q
Configuration
Everything lives in config/default.yaml; ${VAR} placeholders resolve
from the environment (.env is loaded first). Thresholds:
recognition.match_threshold(default 0.42): raise for fewer false matches, lower for fewer duplicates.recognition.min_enroll_quality(0.65): how good a face must look (sharpness, size, lighting, frontality) before a new identity is minted.tracking.min_hits_for_id(4): frames a face must persist before we spend an embedding on it — filters passers-by and phantom detections.cameras[].max_width(1280): frames are downscaled at ingest — full 3MP streams waste memory and detector time.
Recognition models
The encoder picks the first usable model in models/:
arcface_int8.onnx → w600k_mbf.onnx (MobileFaceNet, 13 MB, downloaded
automatically) → arcface.onnx (r100, 260 MB, copied from the old project;
needs ~1.5 GB free RAM to load). Pin one with recognition.model_file.
Every stored embedding is tagged with the model that produced it, and only
embeddings from the active model are searched — different encoders'
vectors are numerically incompatible and never mix.