Suriyakumarvijayanayagam ecc8bbba6f The app on a laptop reported an engine that was never meant to be there
Signing in on a second Mac showed "engine not reachable at
http://127.0.0.1:8010" and 0 of 0 cameras, on an account whose shops were
running and recognising people the whole time. Nothing was broken: Live() and
Cameras() read only the engine on loopback, so the app answered as though the
person had never signed in - and camera sync goes through the engine, which is
why the count was zero rather than stale.

Having no engine is a normal state. A shop PC watches cameras; an owner's
laptop, a manager's machine and a second till being set up do not, and all
three are signed in to the same estate. Both methods now fall back to head
office when loopback fails and somebody is signed in. Loopback is still tried
first: a real shop PC must never be shown a minute-old summary when the engine
two milliseconds away has the live one.

Decisions worth keeping:

- Viewing is on the snapshot, not inferred per screen. Three surfaces read it,
  and a screen that computed it separately is how the shops screen once came
  out labelled Working, in green, above "2 of 3 cameras not connecting".
- fraction_below_gate takes the WORST shop, never an average. 0.10 against
  0.73 averages to 0.42 and hides the only shop anyone needs to visit.
- A remote camera is flagged, and Edit, Remove and Check placement are
  withheld. They talk to a camera on a LAN this computer cannot reach, and a
  button that cannot work is worse than one that is absent.
- connected is three states. null is "no shop computer has reported yet" and
  reads as waiting; false is "Not connecting". A bare false sends somebody to
  check cabling on a camera nobody has tried to reach.
- Snapshots are fetched in Go as data: URIs and cached by snapshot_at. A
  webview <img> resolves a relative src against wails:// and cannot send the
  bearer - the problem VisitorImage already solved - and this screen polls
  every 8 seconds at ~90 KB a camera.
- With no engine AND no session, the engine error is still the answer. The
  person is most likely setting this PC up.

The picture is the last snapshot and the banner says so: there is no live
video from here, because the engine's MJPEG stream is on the shop PC's
loopback behind a router with no inbound route. The LiveHub relay head office
uses is the answer to that and is a further step for this client.

Verified against production: five arrivals and two cameras parsed from the
real API. viewing_test.go covers the fallback, the worst-shop rule, the
withheld credentials and that an unchanged snapshot is fetched once across
two polls.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 16:33:32 +05:30
2026-09-29 16:08:52 +05:30

Behavision

Production face recognition over RTSP. Watches camera streams, detects and tracks faces, recognizes known people, auto-enrolls new visitors, records visit history, and serves a live dashboard + JSON API.

Clean-room rewrite of the previous Camera/ and pattern_reg/ projects: same core ideas, correct engineering.

Quick start (Windows)

cd D:\NEARLE\Behavision
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
python -m behavision setup-models   # downloads YuNet, copies ArcFace etc. from the old project
python -m behavision run            # dashboard at http://localhost:8010

Camera credentials live in .env (gitignored) — never in code or YAML. So do the dashboard credentials: set BEHAVISION_API_USER and BEHAVISION_API_PASSWORD, or let the server generate one into data/api_credentials.txt on first boot. A routable api.host is never served without HTTP Basic auth; 127.0.0.1 is left open. To test without a camera, set webcam: 0 on a camera in config/default.yaml.

Enroll a person by name from photos:

python -m behavision enroll --name "Alice" --images C:\photos\alice\

Architecture

behavision/
├── config.py        typed config: YAML + ${ENV} expansion, validated (pydantic)
├── capture.py       RTSP/webcam reader thread: latest-frame slot, TCP transport,
│                    exponential-backoff reconnect, percent-encoded credentials
├── detection.py     YuNet face detector (OpenCV) → boxes + 5 landmarks, clipped
├── recognition.py   ArcFace ONNX encoder (correct (x-127.5)/127.5 RGB preprocessing,
│                    unit-norm output) + clamped face-quality scoring
├── tracking.py      IoU tracker: identity decided once per TRACK, not per frame
├── gallery/
│   ├── index.py     FAISS IndexFlatIP (exact cosine) with identical numpy fallback
│   ├── store.py     SQLite (WAL): identities, embeddings, sightings — source of truth
│   └── service.py   three-zone matching: match / ambiguous(do nothing) / enroll
├── attributes.py    optional age, gender, emotion on the aligned chip
├── events.py        async event bus → log / webhook / rate-limited email sinks
├── engine.py        one worker thread per camera, shared models + gallery
├── api.py           FastAPI: dashboard, MJPEG stream, identities, events, stats
└── __main__.py      CLI: run | enroll | setup-models

Pipeline

RTSP ──► capture ──► detect (YuNet) ──► track (IoU)
                                          │  once per track, quality-gated
                                          ▼
                       align (Umeyama 5-pt) ──► ArcFace ──► cosine search
                                          │
              ┌───────────────────────────┼──────────────────────────┐
        sim ≥ 0.42                0.32 ≤ sim < 0.42            sim < 0.32
        known person              ambiguous → retry            new visitor
        sighting + event          on a better frame            auto-enroll + event

Design decisions (and the failure they prevent)

Decision Prevents
Per-camera FaceDetector, shared thread-safe encoder cv2 input-size race between camera workers
HTTP Basic on every route, escaped dashboard output open biometric API on the LAN; stored XSS via identity labels
Percent-encoded credentials, URL built from parts @ in password silently breaking the stream (old bug)
Track-level identity, sighting cooldown one user registered per frame (old bug)
Exact IndexFlatIP on unit vectors, -1 guarded inverted L2 threshold + wrong-person metadata[-1] (old bugs)
SQLite as source of truth, index rebuilt at boot index/metadata drift, untrained-IVF crash (old bugs)
Three-zone thresholds with ambiguous no-op duplicate identities and wrong merges
One color conversion, ArcFace-native normalization off-distribution embeddings making thresholds meaningless (old bug)
All quality terms clamped to [0,1] unreachable registration threshold (old bug)
Readiness-guarded API, sinks off the hot path startup crashes, notification stalls

API

Method Path Purpose
GET / live dashboard
GET /api/health, /api/stats liveness / metrics
GET /api/cameras/{id}/stream.mjpeg annotated live stream
GET /api/cameras/{id}/frame.jpg latest annotated frame
GET /api/identities, /api/sightings, /api/events data
PATCH /api/identities/{id} rename a visitor ({"label": "Alice"})
DELETE /api/identities/{id} forget a person (embeddings removed)

Tests

pip install pytest
pytest tests -q

Configuration

Everything lives in config/default.yaml; ${VAR} placeholders resolve from the environment (.env is loaded first). Thresholds:

  • recognition.match_threshold (default 0.42): raise for fewer false matches, lower for fewer duplicates.
  • recognition.min_enroll_quality (0.65): how good a face must look (sharpness, size, lighting, frontality) before a new identity is minted.
  • tracking.min_hits_for_id (4): frames a face must persist before we spend an embedding on it — filters passers-by and phantom detections.
  • cameras[].max_width (1280): frames are downscaled at ingest — full 3MP streams waste memory and detector time.

Recognition models

The encoder picks the first usable model in models/: arcface_int8.onnx → w600k_mbf.onnx (MobileFaceNet, 13 MB, downloaded automatically) → arcface.onnx (r100, 260 MB, copied from the old project; needs ~1.5 GB free RAM to load). Pin one with recognition.model_file. Every stored embedding is tagged with the model that produced it, and only embeddings from the active model are searched — different encoders' vectors are numerically incompatible and never mix.

Description
No description provided
Readme 6 MiB
Languages
Go 60.6%
Python 19%
JavaScript 11.2%
CSS 3.5%
PLpgSQL 2.3%
Other 3.3%