Suriyakumarvijayanayagam 5e544eee3d The camera tiles put a password in the page, and loaded nothing
StreamURL built http://user:pass@127.0.0.1:8010/api/cameras/<id>/
stream.mjpeg and handed it to an <img>, with a comment saying the
credentials were inline "so an <img> tag can load it".

It cannot. Chromium strips credentials from subresource URLs and has
since M59, and WebView2 is Chromium - so on the one platform this
product ships to, every camera tile on a shop counter was a broken
image. Measured against a running engine: the app's Go-side calls
returned stats and people while an <img> on that very URL failed, and
curl proved the URL answered 200. The engine was never the problem.

The password now stays on this side of the process boundary. A loopback
relay attaches Basic auth and streams the engine's bytes back
unchanged - the same reasoning Shot.jsx already follows at head office,
where an <img> equally cannot carry a session.

What the relay is careful about, since it is a door onto the biometric
API with a credential attached:

  - loopback only, on a port the OS picks; a fixed one would collide
    with whatever else a shop PC runs and read as "the cameras broke"
  - a per-run random token in the path. The engine's own credential
    exists so the live face feed is never served open; an
    unauthenticated relay would hand that feed to any other process on
    the PC. Compared in constant time, and a wrong one is 404, not 403
  - an allow-list of stream.mjpeg and frame.jpg. Holding the token does
    not reach the identity list, the gallery, or erasure
  - camera ids validated, not interpolated
  - every chunk flushed; a buffered MJPEG stream is a tile that never
    paints, which looks identical to the bug being fixed

Two of those were written after a test failed, not before:

  - `..` MATCHES the id pattern, because real camera ids contain dots.
    `/api/cameras/../stream.mjpeg` is not the endpoint anyone intended.
    The id can never hold a slash, so `.` and `..` are the whole
    remaining traversal surface and are now refused by name.
  - the serve goroutine read p.srv off the struct while stop() was
    nilling it, so a quick start/stop dereferenced nil and took the
    process down. Captured before launching now.

FrameURL is deliberately not added. No screen asks for a still, and a
bound method nothing calls is the same defect as a capability the UI
cannot reach, only pointing the other way.

Verified: nine unit tests, plus a live test against the real engine and
the real office camera - two MJPEG frames, 90,793 bytes, no credential
in the URL. Windows and darwin both build; vet clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pcn9asw19WGBfCEaHvNug6
2026-09-10 19:53:28 +05:30

Behavision

Production face recognition over RTSP. Watches camera streams, detects and tracks faces, recognizes known people, auto-enrolls new visitors, records visit history, and serves a live dashboard + JSON API.

Clean-room rewrite of the previous Camera/ and pattern_reg/ projects: same core ideas, correct engineering.

Quick start (Windows)

cd D:\NEARLE\Behavision
python -m venv .venv
.venv\Scripts\activate
pip install -r requirements.txt
python -m behavision setup-models   # downloads YuNet, copies ArcFace etc. from the old project
python -m behavision run            # dashboard at http://localhost:8010

Camera credentials live in .env (gitignored) — never in code or YAML. So do the dashboard credentials: set BEHAVISION_API_USER and BEHAVISION_API_PASSWORD, or let the server generate one into data/api_credentials.txt on first boot. A routable api.host is never served without HTTP Basic auth; 127.0.0.1 is left open. To test without a camera, set webcam: 0 on a camera in config/default.yaml.

Enroll a person by name from photos:

python -m behavision enroll --name "Alice" --images C:\photos\alice\

Architecture

behavision/
├── config.py        typed config: YAML + ${ENV} expansion, validated (pydantic)
├── capture.py       RTSP/webcam reader thread: latest-frame slot, TCP transport,
│                    exponential-backoff reconnect, percent-encoded credentials
├── detection.py     YuNet face detector (OpenCV) → boxes + 5 landmarks, clipped
├── recognition.py   ArcFace ONNX encoder (correct (x-127.5)/127.5 RGB preprocessing,
│                    unit-norm output) + clamped face-quality scoring
├── tracking.py      IoU tracker: identity decided once per TRACK, not per frame
├── gallery/
│   ├── index.py     FAISS IndexFlatIP (exact cosine) with identical numpy fallback
│   ├── store.py     SQLite (WAL): identities, embeddings, sightings — source of truth
│   └── service.py   three-zone matching: match / ambiguous(do nothing) / enroll
├── attributes.py    optional age, gender, emotion on the aligned chip
├── events.py        async event bus → log / webhook / rate-limited email sinks
├── engine.py        one worker thread per camera, shared models + gallery
├── api.py           FastAPI: dashboard, MJPEG stream, identities, events, stats
└── __main__.py      CLI: run | enroll | setup-models

Pipeline

RTSP ──► capture ──► detect (YuNet) ──► track (IoU)
                                          │  once per track, quality-gated
                                          ▼
                       align (Umeyama 5-pt) ──► ArcFace ──► cosine search
                                          │
              ┌───────────────────────────┼──────────────────────────┐
        sim ≥ 0.42                0.32 ≤ sim < 0.42            sim < 0.32
        known person              ambiguous → retry            new visitor
        sighting + event          on a better frame            auto-enroll + event

Design decisions (and the failure they prevent)

Decision Prevents
Per-camera FaceDetector, shared thread-safe encoder cv2 input-size race between camera workers
HTTP Basic on every route, escaped dashboard output open biometric API on the LAN; stored XSS via identity labels
Percent-encoded credentials, URL built from parts @ in password silently breaking the stream (old bug)
Track-level identity, sighting cooldown one user registered per frame (old bug)
Exact IndexFlatIP on unit vectors, -1 guarded inverted L2 threshold + wrong-person metadata[-1] (old bugs)
SQLite as source of truth, index rebuilt at boot index/metadata drift, untrained-IVF crash (old bugs)
Three-zone thresholds with ambiguous no-op duplicate identities and wrong merges
One color conversion, ArcFace-native normalization off-distribution embeddings making thresholds meaningless (old bug)
All quality terms clamped to [0,1] unreachable registration threshold (old bug)
Readiness-guarded API, sinks off the hot path startup crashes, notification stalls

API

Method Path Purpose
GET / live dashboard
GET /api/health, /api/stats liveness / metrics
GET /api/cameras/{id}/stream.mjpeg annotated live stream
GET /api/cameras/{id}/frame.jpg latest annotated frame
GET /api/identities, /api/sightings, /api/events data
PATCH /api/identities/{id} rename a visitor ({"label": "Alice"})
DELETE /api/identities/{id} forget a person (embeddings removed)

Tests

pip install pytest
pytest tests -q

Configuration

Everything lives in config/default.yaml; ${VAR} placeholders resolve from the environment (.env is loaded first). Thresholds:

  • recognition.match_threshold (default 0.42): raise for fewer false matches, lower for fewer duplicates.
  • recognition.min_enroll_quality (0.65): how good a face must look (sharpness, size, lighting, frontality) before a new identity is minted.
  • tracking.min_hits_for_id (4): frames a face must persist before we spend an embedding on it — filters passers-by and phantom detections.
  • cameras[].max_width (1280): frames are downscaled at ingest — full 3MP streams waste memory and detector time.

Recognition models

The encoder picks the first usable model in models/: arcface_int8.onnx → w600k_mbf.onnx (MobileFaceNet, 13 MB, downloaded automatically) → arcface.onnx (r100, 260 MB, copied from the old project; needs ~1.5 GB free RAM to load). Pin one with recognition.model_file. Every stored embedding is tagged with the model that produced it, and only embeddings from the active model are searched — different encoders' vectors are numerically incompatible and never mix.

Description
No description provided
Readme 6 MiB
Languages
Go 60.6%
Python 19%
JavaScript 11.2%
CSS 3.5%
PLpgSQL 2.3%
Other 3.3%