Commit Graph

9 Commits

Author SHA1 Message Date
3cddd9c2e1 The install finished and then could not download a 230 KB file
With the version ceiling and the widened numpy pin in place, setup succeeded
on the Mac that found them - Python 3.14 chosen and accepted, numpy 2.5.3,
onnxruntime 1.30, faiss 1.15.1, the engine itself - and died on the last step,
fetching the YuNet model:

  ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
  certificate verify failed: unable to get local issuer certificate

A python.org macOS build ships its own OpenSSL with NO trust store, and
populates one only when somebody double-clicks Install Certificates.command in
the Python folder. Nobody installing face-recognition software has a reason to
know that exists, and the failure is forty lines of traceback about _ssl.c at
the end of a ten-minute install.

_urlopen tries the default context first and retries with certifi's bundle on
a verification failure. The order is the design:

- Default first, because on Windows and on a system or Homebrew Python the
  default context reads the machine's own certificate store, which is what
  makes a corporate proxy with its own root CA work. Replacing it
  unconditionally would break every site that has one to fix a different
  platform.
- certifi second, because it is already installed: requests is a hard
  dependency and brings it.
- URLError is re-raised untouched. "No route to host" and "no trust store" are
  different problems, and retrying the first with a different CA list only
  delays the real message.

urlretrieve had to go, since it offers no way to pass a context - exactly the
kind of rewrite that silently drops something. The `download: <label> <n>%`
lines are a contract: supervisor.go's progressRe parses them to put first-run
progress in the tray, because the API is not up yet and a shop PC showing a
stopped engine for five minutes looks broken. A test asserts them, and the
rewritten fetch was checked against the real URL: 232,589 bytes, sha256
identical to the model already on disk.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:21:41 +05:30
ff4f95c3b0 Four silent failures a demo on somebody else's Mac walked straight into
Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.

## behavision-setup chose the Python least likely to work

findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".

maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.

## numpy<2.0 was the cap; OpenCV was the hazard

Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:

  numpy 1.26 / cv2 4.11    9 passed, 0 crashed
  numpy 2.0  / cv2 4.11    8 passed, 1 crashed
  numpy 1.26 / cv2 4.14    3 passed, 6 crashed
  numpy 2.0  / cv2 4.14    2 passed, 7 crashed

numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.

It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.

opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.

## One MQTT client id for a whole shop, so two PCs fought over it

behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:

  broker connected / broker connection lost: EOF / broker connected / EOF ...

The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.

MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.

## "no such file or directory" for an engine nobody had installed

Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.

engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 17:09:25 +05:30
f61da2eeed Three states that looked like health from outside
Audited the engine for what it does when something goes wrong rather than
when it goes right. Each of these left the process healthy, the dashboard
green and the product not working.

A gallery the running encoder cannot read. Embeddings are model-tagged, so
when the fallback chain fires every vector the previous encoder wrote goes
invisible: the shop keeps its customer list and recognises nobody on it,
enrolling each regular a second time. Footfall stays correct, which is why
nothing looks wrong. The only evidence was an INFO line reading 'gallery
ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence
that states the disaster and calls it ready. Gallery.health now warns with
the count of PEOPLE lost, not vectors, and carries the same numbers to
/api/stats and /api/health, because a log line on a shop PC is read by
nobody. Proved against the real 87-embedding gallery.

Connected, and sending nothing. 'connected' meant the socket opened, so a
stream that went quiet kept it true while last_frame_age_s climbed and the
heartbeat told head office the camera was up. OpenCV breaks a blocked read
at 30s, but a camera trickling a frame every 20s never trips that and never
recovers. streaming/stalled are reported beside connected and the dashboard
says live/stalled/offline - three states because offline sends you to the
network and stalled says the camera is answering and sending nothing.

The 5-second RTSP timeout that never existed. stimeout;5000000 carried a
comment claiming it bounded a dead camera. Measured on OpenCV 4.11 /
FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with
stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so
it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither
reaches the RTSP protocol through this path; the real bound is OpenCV's own
interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source
already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause.
That matters beyond speed - the VideoCapture constructor is not
interruptible, so stop() could not cut it short and a camera removed from
head office left a daemon thread holding a socket for half a minute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-24 14:15:54 +05:30
2e60fbb57a The engine was searching an empty room fifteen times a second
Measured rather than guessed, and the first guess was wrong. Wall clock
said H.265 decode cost 58 ms a frame; cap.read() blocks until the next
frame arrives, so that was the frame interval, not work. As CPU time:
decode 3.7 ms, detection 31.0 ms - and detection ran on every frame
whether or not anything was in front of the camera, 6,649 of 8,634
frames with faces_seen 0 and active_tracks 0 throughout.

detect_threads: OpenCV spreads a small repeated job over eight threads,
costing 31.0 ms of CPU for 8.9 ms of wall. One thread costs 15.3 ms for
15.3 ms, against a 66 ms budget at 15 fps. Half the CPU for latency
nothing can notice.

motion_gate: a 160x90 greyscale absdiff, 0.1 ms against detection's 15.
Consulted only while no track is open; forced to look every
motion_max_skip frames; compared against the last frame SEARCHED so a
slow drift cannot creep under the threshold; and a threshold above this
camera's measured noise and far below a person, so anything ambiguous
detects. tests/test_motion_gate.py pins each of those rather than the
saving, including asserting the longest run of skips rather than the
total - counting the total would pass a gate that slept forty frames
and then looked forty times.

Together 80% -> 16% of a core, detection skipped on 92% of frames.
faces_seen is still 0 and the gate is not why: run directly over the
same frames the detector finds nothing at threshold 0.50 either. The
placement is the limit, as recorded; the CPU was being spent to
rediscover that fifteen times a second.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-24 13:58:05 +05:30
1607f4ce74 Find the camera on the network instead of asking for its address
The add-camera form asked for an IP address, and a shop owner does not
know their camera's IP address - it is on a sticker under the camera or
in a menu that differs by make. That field is where onboarding stopped
for anyone who was not an installer.

behavision/discover.py: one ONVIF WS-Discovery multicast (names the
camera and often its make) merged with a TCP sweep of port 554 across
the local /24 (misses nothing that streams). Stdlib only, ~4 s on the
office network, both cameras found. The add-camera sheet leads with
'Find cameras on this network'; picking a row fills the address and,
when the make is recognisable, the stream path.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-21 12:27:47 +05:30
e262fc8482 The live picture was chained to the recognition pipeline
Reported from the first Windows install: the camera feed lags. It did,
and not because of the network, the proxy or the webview.

The MJPEG stream served _annotated_jpeg - the frame the pipeline had
most recently FINISHED with, encoded after detection, quality scoring,
tracking and identification had all run on it. On a modest shop PC that
is a few frames a second, and every picture was already as old as that
processing. It looked like lag because it was lag. On the fast machine
it was developed on the pipeline kept up with the stream's own 10 fps
cap, which is why nobody here ever saw it.

Two more things compounded it. Every processed frame was JPEG-encoded
whether or not a viewer existed - CPU spent on precisely the machine
short of it. And ffmpeg ran its RTSP demuxer with default buffering,
which holds a comfortable queue of frames before handing over the first:
half a second to two seconds a live view can never recover.

Now the picture and the boxes are decoupled. latest_jpeg_since takes the
capture thread's freshest frame at the camera's own rate and draws the
boxes from the last processed frame over it - encoded on demand, per
request, so a camera nobody watches costs no encode at all. The stream
sends a frame only when the camera has a newer one, capped at 15 fps;
nothing is sent twice. Boxes older than a second are not drawn, so a
stalled pipeline cannot leave one floating over an empty spot.
_publish_annotated becomes _remember_tracks: a handful of tuples under
the lock, no copy, no encode. ffmpeg gets nobuffer / low_delay /
max_delay.

Measured on cam2's sub-stream, same machine, ten seconds each:

  before   99 frames sent,  98 distinct    9.8 new pictures/s
  after   141 frames sent, 141 distinct   14.0 new pictures/s

against a 15 fps camera, with the pipeline still processing 166 of 181
captured frames alongside - and engine CPU DOWN from 90% with no viewer
to 62% with one attached.

Engine version 1.0.0 -> 1.1.0 so a re-run of setup reinstalls it rather
than pip deciding the requirement is already satisfied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-11 16:28:05 +05:30
ffae7e45d5 Live view runs at the camera's real rate, and reports why it is MJPEG
4 fps was not "live", and it was a number I picked rather than measured.
The engine actually produces ~12 distinct frames a second, so most of it
was being left on the floor.

Now: poll a little ahead of the engine and drop frames identical to the
last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps,
20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a
picture the viewer has not seen, and the rate follows the camera instead
of a constant.

Also records why this is MJPEG rather than passing the camera's own
compressed video through, which would be smoother, cheaper and use no
CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and
BOTH are H.265, despite stream paths ending in ".264". Browsers play
H.264 everywhere and H.265 only on some platforms, so passthrough cannot
rely on it, and transcoding HEVC on the shop PC would put a video encoder
on the machine already doing the recognition.

So probe_source now reports `codec`. It decides what is possible, an
installer can usually change it, and otherwise the only way to learn it is
to read RTSP by hand - which is how this was found.

The RTSP libraries used to establish that are NOT kept: they were only
ever imported by a spike test, and two large dependencies in a shipped
binary to answer a question OpenCV already knows is a bad trade. Their
`go get` had also silently bumped the agent to go 1.25 and broken the
desktop build, which is its own argument.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 16:59:27 +05:30
c7024b57ca The shipped config could not load on a machine without a .env
config/default.yaml reads its camera's host, username and password from
${ENV}. Unset placeholders parse as YAML null, and CameraConfig had no
_normalize_blanks - the guard ApiSection and EmailSection have had all
along - so loading it raised three pydantic errors and the engine would
not start at all. On a developer's checkout .env is right there, which is
why this survived: the failing machine is every machine the product is
actually installed on, and installer/build.ps1 runs this suite, so the
Windows build would have failed on a fresh clone.

The normalisation is field-by-field, never a blanket None -> "": `webcam`
is an Optional[int] whose None means "this is not a webcam", and `tuning`
is a nested model. Sweeping either trades one validation error for
another - which it did, on the first attempt.

Second bug behind the same line: that camera entry would then have been
SEEDED into a fresh install, giving a shop a camera called cam1 that
nobody added, retrying a connection to "" forever, with the first task on
a new PC being to work out what it was. CameraConfig.addressed() says
what a camera entry needs to be one, and seed() drops the rest.

Found by cloning the repository into a temp directory and running the
tests there. Nothing in a working tree can find this class of bug.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 12:15:53 +05:30
dad04e8cda Behavision: face recognition for retail, edge to head office
Five components that ship as one product:

- behavision/  the recognition engine. RTSP ingest, YuNet detection, IoU
               tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
               FastAPI dashboard. Identity is decided once per TRACK from an
               average of at least three embeddings, never per frame.
- agent/       the Go edge agent: supervises the engine, holds a durable
               spool, and drains it to MQTT. Nothing is acked before the
               broker confirms.
- desktop/     the shop PC application (Wails + React + tray).
- server/      the cloud API, MQTT consumer, reports and assistant.
- web/         platform.loyaly.ai, the head-office app, embedded in the
               server binary.

The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.

CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 11:14:18 +05:30