Asked directly by the owner, about his own cameras, from his phone's
connection. The answer is physics and the product was not giving it.
A camera lives on the shop's LAN behind a router. 192.168.1.121 means
"something on the network I am attached to" and nothing more - from mobile
data, a hotel or head office it resolves to nobody, or to a completely
different device holding that number. There is no route in from the internet
and there must not be: an RTSP camera reachable from outside is how a shop's
cameras end up being watched by strangers.
That is why the product is split the way it is - the shop PC is the only
machine on the camera's LAN, and every other surface reaches it outbound,
which is what makes Watch live work from anywhere while nothing connects in.
What was wrong is the message. "cannot reach 192.168.1.121:554 - Operation
timed out" reads as a broken camera and sends somebody to re-type an address
and a password that were always correct. _wrong_network_hint names the cause
and separates two states that need opposite actions:
on that network -> check the camera is powered on and the address is right
somewhere else -> the COMPUTER is in the wrong place; no setting fixes it
- The local address comes from a connected UDP socket that sends nothing. It
only fixes a route so the kernel will name the source address.
- The LAN ranges are spelled out, not is_private. That property also covers
carrier-grade NAT and the documentation networks, and telling somebody who
typed 203.0.113.9 that it is "on the shop's own network" is a confident
wrong answer in the place people look first. Found by a test using that
address as its example of a PUBLIC one.
- A DNS name gets no hint: nothing can be concluded about camera.local from
the string, and guessing is the failure mode this message exists to fix.
- With no network at all it still names the cause and drops the comparison,
rather than claiming to know which network this machine is on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Audited the engine for what it does when something goes wrong rather than
when it goes right. Each of these left the process healthy, the dashboard
green and the product not working.
A gallery the running encoder cannot read. Embeddings are model-tagged, so
when the fallback chain fires every vector the previous encoder wrote goes
invisible: the shop keeps its customer list and recognises nobody on it,
enrolling each regular a second time. Footfall stays correct, which is why
nothing looks wrong. The only evidence was an INFO line reading 'gallery
ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence
that states the disaster and calls it ready. Gallery.health now warns with
the count of PEOPLE lost, not vectors, and carries the same numbers to
/api/stats and /api/health, because a log line on a shop PC is read by
nobody. Proved against the real 87-embedding gallery.
Connected, and sending nothing. 'connected' meant the socket opened, so a
stream that went quiet kept it true while last_frame_age_s climbed and the
heartbeat told head office the camera was up. OpenCV breaks a blocked read
at 30s, but a camera trickling a frame every 20s never trips that and never
recovers. streaming/stalled are reported beside connected and the dashboard
says live/stalled/offline - three states because offline sends you to the
network and stalled says the camera is answering and sending nothing.
The 5-second RTSP timeout that never existed. stimeout;5000000 carried a
comment claiming it bounded a dead camera. Measured on OpenCV 4.11 /
FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with
stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so
it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither
reaches the RTSP protocol through this path; the real bound is OpenCV's own
interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source
already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause.
That matters beyond speed - the VideoCapture constructor is not
interruptible, so stop() could not cut it short and a camera removed from
head office left a daemon thread holding a socket for half a minute.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Reported from the first Windows install: the camera feed lags. It did,
and not because of the network, the proxy or the webview.
The MJPEG stream served _annotated_jpeg - the frame the pipeline had
most recently FINISHED with, encoded after detection, quality scoring,
tracking and identification had all run on it. On a modest shop PC that
is a few frames a second, and every picture was already as old as that
processing. It looked like lag because it was lag. On the fast machine
it was developed on the pipeline kept up with the stream's own 10 fps
cap, which is why nobody here ever saw it.
Two more things compounded it. Every processed frame was JPEG-encoded
whether or not a viewer existed - CPU spent on precisely the machine
short of it. And ffmpeg ran its RTSP demuxer with default buffering,
which holds a comfortable queue of frames before handing over the first:
half a second to two seconds a live view can never recover.
Now the picture and the boxes are decoupled. latest_jpeg_since takes the
capture thread's freshest frame at the camera's own rate and draws the
boxes from the last processed frame over it - encoded on demand, per
request, so a camera nobody watches costs no encode at all. The stream
sends a frame only when the camera has a newer one, capped at 15 fps;
nothing is sent twice. Boxes older than a second are not drawn, so a
stalled pipeline cannot leave one floating over an empty spot.
_publish_annotated becomes _remember_tracks: a handful of tuples under
the lock, no copy, no encode. ffmpeg gets nobuffer / low_delay /
max_delay.
Measured on cam2's sub-stream, same machine, ten seconds each:
before 99 frames sent, 98 distinct 9.8 new pictures/s
after 141 frames sent, 141 distinct 14.0 new pictures/s
against a 15 fps camera, with the pipeline still processing 166 of 181
captured frames alongside - and engine CPU DOWN from 90% with no viewer
to 62% with one attached.
Engine version 1.0.0 -> 1.1.0 so a re-run of setup reinstalls it rather
than pip deciding the requirement is already satisfied.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
4 fps was not "live", and it was a number I picked rather than measured.
The engine actually produces ~12 distinct frames a second, so most of it
was being left on the floor.
Now: poll a little ahead of the engine and drop frames identical to the
last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps,
20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a
picture the viewer has not seen, and the rate follows the camera instead
of a constant.
Also records why this is MJPEG rather than passing the camera's own
compressed video through, which would be smoother, cheaper and use no
CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and
BOTH are H.265, despite stream paths ending in ".264". Browsers play
H.264 everywhere and H.265 only on some platforms, so passthrough cannot
rely on it, and transcoding HEVC on the shop PC would put a video encoder
on the machine already doing the recognition.
So probe_source now reports `codec`. It decides what is possible, an
installer can usually change it, and otherwise the only way to learn it is
to read RTSP by hand - which is how this was found.
The RTSP libraries used to establish that are NOT kept: they were only
ever imported by a spike test, and two large dependencies in a shipped
binary to answer a question OpenCV already knows is a bad trade. Their
`go get` had also silently bumped the agent to go 1.25 and broken the
desktop build, which is its own argument.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn