Three states that looked like health from outside

Audited the engine for what it does when something goes wrong rather than
when it goes right. Each of these left the process healthy, the dashboard
green and the product not working.

A gallery the running encoder cannot read. Embeddings are model-tagged, so
when the fallback chain fires every vector the previous encoder wrote goes
invisible: the shop keeps its customer list and recognises nobody on it,
enrolling each regular a second time. Footfall stays correct, which is why
nothing looks wrong. The only evidence was an INFO line reading 'gallery
ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence
that states the disaster and calls it ready. Gallery.health now warns with
the count of PEOPLE lost, not vectors, and carries the same numbers to
/api/stats and /api/health, because a log line on a shop PC is read by
nobody. Proved against the real 87-embedding gallery.

Connected, and sending nothing. 'connected' meant the socket opened, so a
stream that went quiet kept it true while last_frame_age_s climbed and the
heartbeat told head office the camera was up. OpenCV breaks a blocked read
at 30s, but a camera trickling a frame every 20s never trips that and never
recovers. streaming/stalled are reported beside connected and the dashboard
says live/stalled/offline - three states because offline sends you to the
network and stalled says the camera is answering and sending nothing.

The 5-second RTSP timeout that never existed. stimeout;5000000 carried a
comment claiming it bounded a dead camera. Measured on OpenCV 4.11 /
FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with
stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so
it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither
reaches the RTSP protocol through this path; the real bound is OpenCV's own
interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source
already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause.
That matters beyond speed - the VideoCapture constructor is not
interruptible, so stop() could not cut it short and a camera removed from
head office left a daemon thread holding a socket for half a minute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-24 14:15:54 +05:30
parent 2e60fbb57a
commit f61da2eeed
8 changed files with 439 additions and 5 deletions

View File

@@ -281,6 +281,31 @@ class IdentityStore:
return None
return np.frombuffer(row["vector"], dtype=np.float32), float(row["quality"])
def model_counts(self) -> "dict[str, int]":
"""How many stored embeddings each encoder produced.
The gallery only ever searches vectors tagged with the *running*
encoder, so this is what says whether the rest of the gallery is
reachable at all. See `Gallery.health` for why that matters.
"""
with self._lock:
rows = self._db.execute(
"SELECT model, COUNT(*) AS n FROM embeddings "
"GROUP BY model").fetchall()
return {str(r["model"]): int(r["n"]) for r in rows}
def identities_with_model(self, model: str) -> int:
"""Identities holding at least one embedding from this encoder.
Not the same as the identity count: an identity whose only vectors
came from a previous encoder still exists, and is unrecognisable.
"""
with self._lock:
row = self._db.execute(
"SELECT COUNT(DISTINCT identity_id) AS n FROM embeddings "
"WHERE model=?", (model,)).fetchone()
return int(row["n"]) if row else 0
def embedding_owners(self, model: "str | None" = None) -> "dict[int, int]":
"""embedding_id -> identity_id, for turning index hits into identity
pairs without a round trip to SQLite per hit."""