Three states that looked like health from outside

Audited the engine for what it does when something goes wrong rather than
when it goes right. Each of these left the process healthy, the dashboard
green and the product not working.

A gallery the running encoder cannot read. Embeddings are model-tagged, so
when the fallback chain fires every vector the previous encoder wrote goes
invisible: the shop keeps its customer list and recognises nobody on it,
enrolling each regular a second time. Footfall stays correct, which is why
nothing looks wrong. The only evidence was an INFO line reading 'gallery
ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence
that states the disaster and calls it ready. Gallery.health now warns with
the count of PEOPLE lost, not vectors, and carries the same numbers to
/api/stats and /api/health, because a log line on a shop PC is read by
nobody. Proved against the real 87-embedding gallery.

Connected, and sending nothing. 'connected' meant the socket opened, so a
stream that went quiet kept it true while last_frame_age_s climbed and the
heartbeat told head office the camera was up. OpenCV breaks a blocked read
at 30s, but a camera trickling a frame every 20s never trips that and never
recovers. streaming/stalled are reported beside connected and the dashboard
says live/stalled/offline - three states because offline sends you to the
network and stalled says the camera is answering and sending nothing.

The 5-second RTSP timeout that never existed. stimeout;5000000 carried a
comment claiming it bounded a dead camera. Measured on OpenCV 4.11 /
FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with
stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so
it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither
reaches the RTSP protocol through this path; the real bound is OpenCV's own
interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source
already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause.
That matters beyond speed - the VideoCapture constructor is not
interruptible, so stop() could not cut it short and a camera removed from
head office left a daemon thread holding a socket for half a minute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-24 14:15:54 +05:30
parent 2e60fbb57a
commit f61da2eeed
8 changed files with 439 additions and 5 deletions

View File

@@ -2285,6 +2285,95 @@ far away, which is the same `fraction_below_gate: 0.59` this file already
records. The placement is still the limit; the CPU was simply being spent to
discover that 15 times a second.
## Three states that looked like health from outside
Found by auditing the engine for what it does when something goes wrong,
rather than when it goes right. Each of these left the process healthy, the
dashboard green and the product not working - the class of bug this file
already calls a headcount wrong in a way nobody can detect.
`tests/test_reliability.py` covers all three.
### A gallery the running encoder cannot read
The fallback chain exists so a memory-starved box still starts, and
CLAUDE.md already warned that "on a memory-starved box the big model silently
loses the chain". What it did not say is what that **costs**: embeddings are
model-tagged, so every vector the previous encoder wrote becomes invisible.
The shop keeps its whole customer list and recognises nobody on it. Every
regular is greeted as a stranger and enrolled a second time. Footfall stays
correct, which is precisely why nothing looks wrong.
The only evidence was an INFO line reading `gallery ready: 0 embeddings
(model 'w600k_mbf') across 21 identities` - a sentence that says the disaster
and calls it ready. Run against the real 87-embedding gallery with the
fallback model forced, it now says:
```
WARNING gallery: 87 of 87 stored embeddings were written by a DIFFERENT
encoder (w600k_r50) and cannot be searched - 21 known people are
unrecognisable under the running model 'w600k_mbf'.
```
`Gallery.health` carries the same numbers to `/api/stats` and
`gallery_unreadable_embeddings` to `/api/health`, because a log line on a shop
PC is read by nobody. It travels for the same reason `fraction_below_gate`
does: beside the number it qualifies. `identities_stranded` is the figure that
matters - **people lost, not vectors** - and an empty gallery reports zero
rather than raising an alarm on a fresh install.
### Connected, and sending nothing
`connected` meant *the socket opened*. A stream that opens and then goes quiet
kept it `true` while `last_frame_age_s` climbed, so the heartbeat told head
office the camera was up. OpenCV breaks a blocked read after 30s and we
reconnect - but a camera trickling one frame every 20s never trips that at
all, so it never reconnects and never recovers.
`stalled()` and `streaming` are reported beside `connected`, and the local
dashboard now says **live / stalled / offline** rather than live / offline.
Three states because two of them need opposite actions: offline sends you to
the network, stalled says the camera is answering and sending nothing. Same
rule as `artifact` vs `no_faces` in the commissioning verdicts.
`STALL_AFTER_S = 10` is not a preference. The tracker abandons a face after
`max_misses` (25 frames, ~1.7s at 15 fps), so by 10s every track is long gone
and 150 frames are missing: whatever this is, recognition cannot use it.
### The 5-second RTSP timeout that never existed
`capture.py` set `stimeout;5000000` with a comment claiming "a 5s socket
timeout so a dead camera is noticed". Measured against this build (OpenCV
4.11, FFmpeg 7.1) on a socket that accepts the connection and then says
nothing:
```
stimeout;5000000 -> 30.0s timeout;5000000 -> 30.0s
stimeout;2000000 -> 30.5s timeout;2000000 -> 30.4s
no timeout option at all -> 30.3s
```
Identical with the option absent, under either name, so it was never honoured
through this path - `stimeout` was renamed `timeout` in FFmpeg 5.0 and neither
reaches the RTSP protocol here. The real bound is OpenCV's own interrupt
callback, a compile-time constant we do not control. Both names are still set
(harmless, and right on a build where they do work), but **nothing depends on
them**.
What replaces it is `_tcp_reachable` in `_open()` - the pre-flight
`probe_source` already used, in code we own. It matters beyond speed: the
`cv2.VideoCapture` constructor is not interruptible, so `stop()` could not cut
it short and a camera removed from head office left a daemon thread holding a
socket for half a minute. Measured:
```
unroutable address 30.3s -> 2.02s "no response from ... within 2s"
host up, port closed 30.3s -> 0.00s "cannot reach ... Connection refused"
wrong port, real cam 30.3s -> 1.01s "cannot reach ... Connection refused"
```
`last_error` is reported with the camera, because `connected: false` alone
cannot tell a wrong IP from a wrong password, and those are different jobs.
## Setting up on a new machine
1. Copy the `Behavision` folder **including `.env`** (gitignored, holds