Measured rather than guessed, and the first guess was wrong. Wall clock
said H.265 decode cost 58 ms a frame; cap.read() blocks until the next
frame arrives, so that was the frame interval, not work. As CPU time:
decode 3.7 ms, detection 31.0 ms - and detection ran on every frame
whether or not anything was in front of the camera, 6,649 of 8,634
frames with faces_seen 0 and active_tracks 0 throughout.
detect_threads: OpenCV spreads a small repeated job over eight threads,
costing 31.0 ms of CPU for 8.9 ms of wall. One thread costs 15.3 ms for
15.3 ms, against a 66 ms budget at 15 fps. Half the CPU for latency
nothing can notice.
motion_gate: a 160x90 greyscale absdiff, 0.1 ms against detection's 15.
Consulted only while no track is open; forced to look every
motion_max_skip frames; compared against the last frame SEARCHED so a
slow drift cannot creep under the threshold; and a threshold above this
camera's measured noise and far below a person, so anything ambiguous
detects. tests/test_motion_gate.py pins each of those rather than the
saving, including asserting the longest run of skips rather than the
total - counting the total would pass a gate that slept forty frames
and then looked forty times.
Together 80% -> 16% of a core, detection skipped on 92% of frames.
faces_seen is still 0 and the gate is not why: run directly over the
same frames the detector finds nothing at threshold 0.50 either. The
placement is the limit, as recorded; the CPU was being spent to
rediscover that fifteen times a second.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The first launch was a code box with a link under it, then an empty
Live screen with 'No cameras' in amber in a far corner, then a form
asking for an IP address, and for the first few minutes of all of it
the engine silently downloading 275 MB with nothing on screen but a
stopped-looking status. Walked in a browser with the new mock; nobody
who was not an installer would have got through it.
Now: a welcome that asks the one question a shop owner can answer -
managed from a head office, or on this PC only - with each path in a
sentence; a Getting Started checklist on Live that reads its three steps
from the engine and ticks them itself (recognition ready, camera added
and connected, camera proven by a walk-past), with the one button for
the next step, and that disappears the moment somebody is recognised;
and the model download reported as a percentage in the tray, the
sidebar and the checklist, parsed by the supervisor from the engine's
own progress lines.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The add-camera form asked for an IP address, and a shop owner does not
know their camera's IP address - it is on a sticker under the camera or
in a menu that differs by make. That field is where onboarding stopped
for anyone who was not an installer.
behavision/discover.py: one ONVIF WS-Discovery multicast (names the
camera and often its make) merged with a TCP sweep of port 554 across
the local /24 (misses nothing that streams). Stdlib only, ~4 s on the
office network, both cameras found. The add-camera sheet leads with
'Find cameras on this network'; picking a row fills the address and,
when the make is recognisable, the stream path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Brand assets in brand/ (the 512px mark, sizes for each surface, a
multi-size .ico). Windows executables carry it as a compiled-in
resource (rsrc_windows_amd64.syso from go-winres) so Explorer, the
taskbar and the installer show it; installer/build.ps1 therefore uses a
plain go build rather than wails build, which would add a second copy
and fail the link. The tray icon is the mark with a state dot over its
corner - a plain coloured circle read as a generic status light among
other icons - rendered from the embedded PNG at 32px so it survives
150% scaling. The desktop app's login, setup and sidebar marks, the
head-office web app's mark and favicon, and the engine dashboard's
favicon are the same file.
Also found while packaging: no wheel so far shipped static/, so the
engine's own dashboard at :8010 on a Windows source install would have
failed with a missing file. package-data now includes it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Reported from the first Windows install: the camera feed lags. It did,
and not because of the network, the proxy or the webview.
The MJPEG stream served _annotated_jpeg - the frame the pipeline had
most recently FINISHED with, encoded after detection, quality scoring,
tracking and identification had all run on it. On a modest shop PC that
is a few frames a second, and every picture was already as old as that
processing. It looked like lag because it was lag. On the fast machine
it was developed on the pipeline kept up with the stream's own 10 fps
cap, which is why nobody here ever saw it.
Two more things compounded it. Every processed frame was JPEG-encoded
whether or not a viewer existed - CPU spent on precisely the machine
short of it. And ffmpeg ran its RTSP demuxer with default buffering,
which holds a comfortable queue of frames before handing over the first:
half a second to two seconds a live view can never recover.
Now the picture and the boxes are decoupled. latest_jpeg_since takes the
capture thread's freshest frame at the camera's own rate and draws the
boxes from the last processed frame over it - encoded on demand, per
request, so a camera nobody watches costs no encode at all. The stream
sends a frame only when the camera has a newer one, capped at 15 fps;
nothing is sent twice. Boxes older than a second are not drawn, so a
stalled pipeline cannot leave one floating over an empty spot.
_publish_annotated becomes _remember_tracks: a handful of tuples under
the lock, no copy, no encode. ffmpeg gets nobuffer / low_delay /
max_delay.
Measured on cam2's sub-stream, same machine, ten seconds each:
before 99 frames sent, 98 distinct 9.8 new pictures/s
after 141 frames sent, 141 distinct 14.0 new pictures/s
against a 15 fps camera, with the pipeline still processing 166 of 181
captured frames alongside - and engine CPU DOWN from 90% with no viewer
to 62% with one attached.
Engine version 1.0.0 -> 1.1.0 so a re-run of setup reinstalls it rather
than pip deciding the requirement is already satisfied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
4 fps was not "live", and it was a number I picked rather than measured.
The engine actually produces ~12 distinct frames a second, so most of it
was being left on the floor.
Now: poll a little ahead of the engine and drop frames identical to the
last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps,
20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a
picture the viewer has not seen, and the rate follows the camera instead
of a constant.
Also records why this is MJPEG rather than passing the camera's own
compressed video through, which would be smoother, cheaper and use no
CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and
BOTH are H.265, despite stream paths ending in ".264". Browsers play
H.264 everywhere and H.265 only on some platforms, so passthrough cannot
rely on it, and transcoding HEVC on the shop PC would put a video encoder
on the machine already doing the recognition.
So probe_source now reports `codec`. It decides what is possible, an
installer can usually change it, and otherwise the only way to learn it is
to read RTSP by hand - which is how this was found.
The RTSP libraries used to establish that are NOT kept: they were only
ever imported by a spike test, and two large dependencies in a shipped
binary to answer a question OpenCV already knows is a bad trade. Their
`go get` had also silently bumped the agent to go 1.25 and broken the
desktop build, which is its own argument.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.
The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.
Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.
Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.
LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.
Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.
Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
config/default.yaml reads its camera's host, username and password from
${ENV}. Unset placeholders parse as YAML null, and CameraConfig had no
_normalize_blanks - the guard ApiSection and EmailSection have had all
along - so loading it raised three pydantic errors and the engine would
not start at all. On a developer's checkout .env is right there, which is
why this survived: the failing machine is every machine the product is
actually installed on, and installer/build.ps1 runs this suite, so the
Windows build would have failed on a fresh clone.
The normalisation is field-by-field, never a blanket None -> "": `webcam`
is an Optional[int] whose None means "this is not a webcam", and `tuning`
is a nested model. Sweeping either trades one validation error for
another - which it did, on the first attempt.
Second bug behind the same line: that camera entry would then have been
SEEDED into a fresh install, giving a shop a camera called cam1 that
nobody added, retrying a connection to "" forever, with the first task on
a new PC being to work out what it was. CameraConfig.addressed() says
what a camera entry needs to be one, and seed() drops the rest.
Found by cloning the repository into a temp directory and running the
tests there. Nothing in a working tree can find this class of bug.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn