Three states that looked like health from outside

Audited the engine for what it does when something goes wrong rather than
when it goes right. Each of these left the process healthy, the dashboard
green and the product not working.

A gallery the running encoder cannot read. Embeddings are model-tagged, so
when the fallback chain fires every vector the previous encoder wrote goes
invisible: the shop keeps its customer list and recognises nobody on it,
enrolling each regular a second time. Footfall stays correct, which is why
nothing looks wrong. The only evidence was an INFO line reading 'gallery
ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence
that states the disaster and calls it ready. Gallery.health now warns with
the count of PEOPLE lost, not vectors, and carries the same numbers to
/api/stats and /api/health, because a log line on a shop PC is read by
nobody. Proved against the real 87-embedding gallery.

Connected, and sending nothing. 'connected' meant the socket opened, so a
stream that went quiet kept it true while last_frame_age_s climbed and the
heartbeat told head office the camera was up. OpenCV breaks a blocked read
at 30s, but a camera trickling a frame every 20s never trips that and never
recovers. streaming/stalled are reported beside connected and the dashboard
says live/stalled/offline - three states because offline sends you to the
network and stalled says the camera is answering and sending nothing.

The 5-second RTSP timeout that never existed. stimeout;5000000 carried a
comment claiming it bounded a dead camera. Measured on OpenCV 4.11 /
FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with
stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so
it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither
reaches the RTSP protocol through this path; the real bound is OpenCV's own
interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source
already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause.
That matters beyond speed - the VideoCapture constructor is not
interruptible, so stop() could not cut it short and a camera removed from
head office left a daemon thread holding a socket for half a minute.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-24 14:15:54 +05:30
parent 2e60fbb57a
commit f61da2eeed
8 changed files with 439 additions and 5 deletions

View File

@@ -20,7 +20,26 @@ log = logging.getLogger(__name__)
# Set before OpenCV loads ffmpeg, which reads this once.
#
# rtsp_transport=tcp: UDP is the default and silently drops frames on lossy
# Wi-Fi. stimeout: a 5s socket timeout so a dead camera is noticed.
# Wi-Fi.
#
# The timeout here is NOT what bounds a dead camera, and the comment that
# once said it did was wrong. Measured against OpenCV 4.11 / FFmpeg 7.1 on a
# socket that accepts the connection and then says nothing:
#
# stimeout;5000000 -> 30.0s timeout;5000000 -> 30.0s
# stimeout;2000000 -> 30.5s timeout;2000000 -> 30.4s
# no timeout option at all -> 30.3s
#
# Identical with the option absent, so it is not being honoured under either
# name through this path. `stimeout` was renamed `timeout` in FFmpeg 5.0, and
# neither reaches the RTSP protocol here. What actually bounds it is
# OpenCV's own interrupt callback (30s for open, 30s for read), which is a
# compile-time constant we do not control.
#
# Both names are still set, because on a build where they DO take effect the
# shorter bound is what we want and an unrecognised option is ignored. But
# nothing may depend on it: a wrong address is caught by _tcp_reachable
# below, in code we own, in under a second.
#
# fflags=nobuffer and flags=low_delay: without them ffmpeg's RTSP demuxer
# holds a comfortable queue of frames before handing over the first, which
@@ -30,10 +49,25 @@ log = logging.getLogger(__name__)
# reorder wait for the same reason.
os.environ.setdefault(
"OPENCV_FFMPEG_CAPTURE_OPTIONS",
"rtsp_transport;tcp|stimeout;5000000|fflags;nobuffer|flags;low_delay|max_delay;200000",
"rtsp_transport;tcp|stimeout;5000000|timeout;5000000"
"|fflags;nobuffer|flags;low_delay|max_delay;200000",
)
# A stream can stay open and stop delivering. OpenCV breaks a blocked read
# after 30s and we reconnect, but for those 30s `connected` is True and the
# camera is dead — and a stream that trickles a frame every 20s never trips
# that timeout at all, so it never reconnects and never recovers either.
#
# 10s is not a preference. The tracker gives up on a face after `max_misses`
# (25 frames, ~1.7s at 15 fps), so by 10s every track is long gone and 150
# frames are missing: whatever this is, it is not something recognition can
# work with. Reported separately from `connected` because the two need
# opposite actions — one says check the network, the other says the camera
# is answering but sending nothing.
STALL_AFTER_S = 10.0
def _tcp_reachable(source: "str | int", timeout: float
) -> "tuple[bool, str]":
"""Cheap pre-flight for an rtsp:// URL. Non-URL sources pass through."""
@@ -169,6 +203,10 @@ class VideoSource(threading.Thread):
self.frames_total = 0
self.reconnects = 0
self._ever_connected = False
# Why the last open failed, in the words an installer can act on.
# Without it a camera that never connects reports only `connected:
# false`, which cannot distinguish a wrong IP from a wrong password.
self.last_error = ""
# -- public ---------------------------------------------------------
def latest(self) -> "tuple[Optional[np.ndarray], float]":
@@ -199,11 +237,23 @@ class VideoSource(threading.Thread):
def stop(self) -> None:
self._stopping.set()
def stalled(self) -> bool:
"""Open, but not delivering. See STALL_AFTER_S."""
if not self.connected or not self._frame_ts:
return False
return (time.time() - self._frame_ts) > STALL_AFTER_S
def stats(self) -> dict:
return {
"camera_id": self.camera_id,
"url": self._display_url,
"connected": self.connected,
# Connected AND delivering. `connected` alone stays true through
# a stall, so it is the wrong thing for a dashboard to colour a
# camera green on.
"streaming": self.connected and not self.stalled(),
"stalled": self.stalled(),
"last_error": self.last_error,
"frames_total": self.frames_total,
"reconnects": self.reconnects,
"last_frame_age_s": round(time.time() - self._frame_ts, 1)
@@ -260,6 +310,19 @@ class VideoSource(threading.Thread):
log.info("[%s] capture stopped", self.camera_id)
def _open(self) -> Optional[cv2.VideoCapture]:
# Pre-flight the socket, exactly as probe_source does. Without it a
# camera that is off, moved or mistyped costs 30s per attempt inside
# the VideoCapture constructor (measured; it is OpenCV's interrupt
# timeout, not ours to shorten) — and the constructor is not
# interruptible, so stop() cannot cut it short and a removed camera
# leaves a daemon thread holding a socket for half a minute. A
# refused or unroutable address answers in well under a second, which
# is also what lets the backoff below mean what it says.
reachable, why = _tcp_reachable(self._source, 2.0)
if not reachable:
log.debug("[%s] %s", self.camera_id, why)
self.last_error = why
return None
try:
if isinstance(self._source, int):
cap = cv2.VideoCapture(self._source)
@@ -268,8 +331,12 @@ class VideoSource(threading.Thread):
cap.set(cv2.CAP_PROP_BUFFERSIZE, 1)
if not cap.isOpened():
cap.release()
self.last_error = ("reachable, but the stream would not open "
"- check the path and credentials")
return None
self.last_error = ""
return cap
except cv2.error:
log.exception("[%s] VideoCapture error", self.camera_id)
self.last_error = "VideoCapture error - see the engine log"
return None