Three states that looked like health from outside
Audited the engine for what it does when something goes wrong rather than when it goes right. Each of these left the process healthy, the dashboard green and the product not working. A gallery the running encoder cannot read. Embeddings are model-tagged, so when the fallback chain fires every vector the previous encoder wrote goes invisible: the shop keeps its customer list and recognises nobody on it, enrolling each regular a second time. Footfall stays correct, which is why nothing looks wrong. The only evidence was an INFO line reading 'gallery ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence that states the disaster and calls it ready. Gallery.health now warns with the count of PEOPLE lost, not vectors, and carries the same numbers to /api/stats and /api/health, because a log line on a shop PC is read by nobody. Proved against the real 87-embedding gallery. Connected, and sending nothing. 'connected' meant the socket opened, so a stream that went quiet kept it true while last_frame_age_s climbed and the heartbeat told head office the camera was up. OpenCV breaks a blocked read at 30s, but a camera trickling a frame every 20s never trips that and never recovers. streaming/stalled are reported beside connected and the dashboard says live/stalled/offline - three states because offline sends you to the network and stalled says the camera is answering and sending nothing. The 5-second RTSP timeout that never existed. stimeout;5000000 carried a comment claiming it bounded a dead camera. Measured on OpenCV 4.11 / FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither reaches the RTSP protocol through this path; the real bound is OpenCV's own interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause. That matters beyond speed - the VideoCapture constructor is not interruptible, so stop() could not cut it short and a camera removed from head office left a daemon thread holding a socket for half a minute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
89
CLAUDE.md
89
CLAUDE.md
@@ -2285,6 +2285,95 @@ far away, which is the same `fraction_below_gate: 0.59` this file already
|
||||
records. The placement is still the limit; the CPU was simply being spent to
|
||||
discover that 15 times a second.
|
||||
|
||||
## Three states that looked like health from outside
|
||||
|
||||
Found by auditing the engine for what it does when something goes wrong,
|
||||
rather than when it goes right. Each of these left the process healthy, the
|
||||
dashboard green and the product not working - the class of bug this file
|
||||
already calls a headcount wrong in a way nobody can detect.
|
||||
`tests/test_reliability.py` covers all three.
|
||||
|
||||
### A gallery the running encoder cannot read
|
||||
|
||||
The fallback chain exists so a memory-starved box still starts, and
|
||||
CLAUDE.md already warned that "on a memory-starved box the big model silently
|
||||
loses the chain". What it did not say is what that **costs**: embeddings are
|
||||
model-tagged, so every vector the previous encoder wrote becomes invisible.
|
||||
The shop keeps its whole customer list and recognises nobody on it. Every
|
||||
regular is greeted as a stranger and enrolled a second time. Footfall stays
|
||||
correct, which is precisely why nothing looks wrong.
|
||||
|
||||
The only evidence was an INFO line reading `gallery ready: 0 embeddings
|
||||
(model 'w600k_mbf') across 21 identities` - a sentence that says the disaster
|
||||
and calls it ready. Run against the real 87-embedding gallery with the
|
||||
fallback model forced, it now says:
|
||||
|
||||
```
|
||||
WARNING gallery: 87 of 87 stored embeddings were written by a DIFFERENT
|
||||
encoder (w600k_r50) and cannot be searched - 21 known people are
|
||||
unrecognisable under the running model 'w600k_mbf'.
|
||||
```
|
||||
|
||||
`Gallery.health` carries the same numbers to `/api/stats` and
|
||||
`gallery_unreadable_embeddings` to `/api/health`, because a log line on a shop
|
||||
PC is read by nobody. It travels for the same reason `fraction_below_gate`
|
||||
does: beside the number it qualifies. `identities_stranded` is the figure that
|
||||
matters - **people lost, not vectors** - and an empty gallery reports zero
|
||||
rather than raising an alarm on a fresh install.
|
||||
|
||||
### Connected, and sending nothing
|
||||
|
||||
`connected` meant *the socket opened*. A stream that opens and then goes quiet
|
||||
kept it `true` while `last_frame_age_s` climbed, so the heartbeat told head
|
||||
office the camera was up. OpenCV breaks a blocked read after 30s and we
|
||||
reconnect - but a camera trickling one frame every 20s never trips that at
|
||||
all, so it never reconnects and never recovers.
|
||||
|
||||
`stalled()` and `streaming` are reported beside `connected`, and the local
|
||||
dashboard now says **live / stalled / offline** rather than live / offline.
|
||||
Three states because two of them need opposite actions: offline sends you to
|
||||
the network, stalled says the camera is answering and sending nothing. Same
|
||||
rule as `artifact` vs `no_faces` in the commissioning verdicts.
|
||||
|
||||
`STALL_AFTER_S = 10` is not a preference. The tracker abandons a face after
|
||||
`max_misses` (25 frames, ~1.7s at 15 fps), so by 10s every track is long gone
|
||||
and 150 frames are missing: whatever this is, recognition cannot use it.
|
||||
|
||||
### The 5-second RTSP timeout that never existed
|
||||
|
||||
`capture.py` set `stimeout;5000000` with a comment claiming "a 5s socket
|
||||
timeout so a dead camera is noticed". Measured against this build (OpenCV
|
||||
4.11, FFmpeg 7.1) on a socket that accepts the connection and then says
|
||||
nothing:
|
||||
|
||||
```
|
||||
stimeout;5000000 -> 30.0s timeout;5000000 -> 30.0s
|
||||
stimeout;2000000 -> 30.5s timeout;2000000 -> 30.4s
|
||||
no timeout option at all -> 30.3s
|
||||
```
|
||||
|
||||
Identical with the option absent, under either name, so it was never honoured
|
||||
through this path - `stimeout` was renamed `timeout` in FFmpeg 5.0 and neither
|
||||
reaches the RTSP protocol here. The real bound is OpenCV's own interrupt
|
||||
callback, a compile-time constant we do not control. Both names are still set
|
||||
(harmless, and right on a build where they do work), but **nothing depends on
|
||||
them**.
|
||||
|
||||
What replaces it is `_tcp_reachable` in `_open()` - the pre-flight
|
||||
`probe_source` already used, in code we own. It matters beyond speed: the
|
||||
`cv2.VideoCapture` constructor is not interruptible, so `stop()` could not cut
|
||||
it short and a camera removed from head office left a daemon thread holding a
|
||||
socket for half a minute. Measured:
|
||||
|
||||
```
|
||||
unroutable address 30.3s -> 2.02s "no response from ... within 2s"
|
||||
host up, port closed 30.3s -> 0.00s "cannot reach ... Connection refused"
|
||||
wrong port, real cam 30.3s -> 1.01s "cannot reach ... Connection refused"
|
||||
```
|
||||
|
||||
`last_error` is reported with the camera, because `connected: false` alone
|
||||
cannot tell a wrong IP from a wrong password, and those are different jobs.
|
||||
|
||||
## Setting up on a new machine
|
||||
|
||||
1. Copy the `Behavision` folder **including `.env`** (gitignored, holds
|
||||
|
||||
Reference in New Issue
Block a user