Three states that looked like health from outside
Audited the engine for what it does when something goes wrong rather than when it goes right. Each of these left the process healthy, the dashboard green and the product not working. A gallery the running encoder cannot read. Embeddings are model-tagged, so when the fallback chain fires every vector the previous encoder wrote goes invisible: the shop keeps its customer list and recognises nobody on it, enrolling each regular a second time. Footfall stays correct, which is why nothing looks wrong. The only evidence was an INFO line reading 'gallery ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence that states the disaster and calls it ready. Gallery.health now warns with the count of PEOPLE lost, not vectors, and carries the same numbers to /api/stats and /api/health, because a log line on a shop PC is read by nobody. Proved against the real 87-embedding gallery. Connected, and sending nothing. 'connected' meant the socket opened, so a stream that went quiet kept it true while last_frame_age_s climbed and the heartbeat told head office the camera was up. OpenCV breaks a blocked read at 30s, but a camera trickling a frame every 20s never trips that and never recovers. streaming/stalled are reported beside connected and the dashboard says live/stalled/offline - three states because offline sends you to the network and stalled says the camera is answering and sending nothing. The 5-second RTSP timeout that never existed. stimeout;5000000 carried a comment claiming it bounded a dead camera. Measured on OpenCV 4.11 / FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither reaches the RTSP protocol through this path; the real bound is OpenCV's own interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause. That matters beyond speed - the VideoCapture constructor is not interruptible, so stop() could not cut it short and a camera removed from head office left a daemon thread holding a socket for half a minute. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
170
tests/test_reliability.py
Normal file
170
tests/test_reliability.py
Normal file
@@ -0,0 +1,170 @@
|
||||
"""Failure modes that look like health from outside.
|
||||
|
||||
Each of these was a state the engine could be in while every existing test
|
||||
passed and the dashboard showed green. They are grouped because they share
|
||||
one property: the process is fine and the product is not working.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import time
|
||||
|
||||
import numpy as np
|
||||
import pytest
|
||||
|
||||
from behavision import capture
|
||||
from behavision.capture import VideoSource
|
||||
from behavision.config import RecognitionSection
|
||||
from behavision.gallery import Gallery, IdentityStore, VectorIndex
|
||||
from behavision.recognition import EMBEDDING_DIM
|
||||
|
||||
|
||||
def _vec(seed: int) -> np.ndarray:
|
||||
"""A distinct unit vector per seed.
|
||||
|
||||
Orthogonal per index, NOT a constant fill: a vector of all 0.3 and one of
|
||||
all 0.6 normalise to the same direction, so a fixture built that way would
|
||||
call two 'different' people one identity and prove nothing.
|
||||
"""
|
||||
v = np.zeros(EMBEDDING_DIM, dtype=np.float32)
|
||||
v[seed % EMBEDDING_DIM] = 1.0
|
||||
return v
|
||||
|
||||
|
||||
def _gallery(store: IdentityStore, model: str) -> Gallery:
|
||||
return Gallery(store, VectorIndex(EMBEDDING_DIM), RecognitionSection(),
|
||||
model_name=model)
|
||||
|
||||
|
||||
# -- the gallery the running encoder cannot read ------------------------
|
||||
|
||||
def test_matching_model_is_fully_usable(tmp_path):
|
||||
store = IdentityStore(tmp_path / "g.db")
|
||||
ident = store.create_identity("Alice")
|
||||
store.add_embedding(ident, _vec(1), 0.8, "w600k_r50")
|
||||
|
||||
health = _gallery(store, "w600k_r50").health
|
||||
assert health["usable"] == 1
|
||||
assert health["stranded"] == 0
|
||||
assert health["identities_stranded"] == 0
|
||||
|
||||
|
||||
def test_fallback_encoder_strands_the_gallery_and_says_so(tmp_path, caplog):
|
||||
"""The whole point: 2 known people, 0 recognisable, and it must be LOUD.
|
||||
|
||||
This is what a memory-starved box does when the 166 MB model loses the
|
||||
fallback chain to the 13 MB one. Footfall keeps counting, so nothing
|
||||
downstream looks wrong; every regular is simply greeted as a stranger and
|
||||
enrolled a second time.
|
||||
"""
|
||||
store = IdentityStore(tmp_path / "g.db")
|
||||
for i in (1, 2):
|
||||
ident = store.create_identity(f"Person {i}")
|
||||
store.add_embedding(ident, _vec(i), 0.8, "w600k_r50")
|
||||
|
||||
with caplog.at_level("WARNING"):
|
||||
health = _gallery(store, "w600k_mbf").health
|
||||
|
||||
assert health["usable"] == 0
|
||||
assert health["stranded"] == 2
|
||||
assert health["identities_stranded"] == 2
|
||||
assert health["other_models"] == ["w600k_r50"]
|
||||
|
||||
warning = " ".join(r.getMessage() for r in caplog.records
|
||||
if r.levelname == "WARNING")
|
||||
assert "w600k_r50" in warning and "w600k_mbf" in warning, warning
|
||||
|
||||
|
||||
def test_partially_stranded_counts_only_the_unreachable(tmp_path):
|
||||
"""A mixed gallery is the normal state after a model change, and the
|
||||
number that matters is how many people are lost, not how many vectors."""
|
||||
store = IdentityStore(tmp_path / "g.db")
|
||||
old = store.create_identity("Old")
|
||||
store.add_embedding(old, _vec(1), 0.8, "w600k_mbf")
|
||||
store.add_embedding(old, _vec(2), 0.8, "w600k_mbf")
|
||||
both = store.create_identity("Both")
|
||||
store.add_embedding(both, _vec(3), 0.8, "w600k_mbf")
|
||||
store.add_embedding(both, _vec(4), 0.8, "w600k_r50")
|
||||
|
||||
health = _gallery(store, "w600k_r50").health
|
||||
assert health["stored"] == 4
|
||||
assert health["usable"] == 1
|
||||
assert health["stranded"] == 3
|
||||
# "Both" survives the change; only "Old" is unrecognisable.
|
||||
assert health["identities_stranded"] == 1
|
||||
|
||||
|
||||
def test_empty_gallery_is_not_reported_as_stranded(tmp_path):
|
||||
"""A new install must not raise an alarm about a gallery nobody has
|
||||
filled yet - crying wolf here trains people to ignore the real one."""
|
||||
health = _gallery(IdentityStore(tmp_path / "g.db"), "w600k_r50").health
|
||||
assert health["stranded"] == 0
|
||||
assert health["identities_stranded"] == 0
|
||||
|
||||
|
||||
# -- open, but not delivering -------------------------------------------
|
||||
|
||||
def _source() -> VideoSource:
|
||||
return VideoSource("cam", "rtsp://198.51.100.9:554/x")
|
||||
|
||||
|
||||
def test_a_camera_that_never_connected_is_not_stalled():
|
||||
"""`stalled` must mean 'was working, stopped'. A camera that has never
|
||||
delivered a frame is a different fault with a different fix."""
|
||||
src = _source()
|
||||
assert src.stalled() is False
|
||||
src.connected = True
|
||||
assert src.stalled() is False, "no frame ever seen is not a stall"
|
||||
|
||||
|
||||
def test_a_fresh_frame_is_not_a_stall():
|
||||
src = _source()
|
||||
src.connected = True
|
||||
src._frame_ts = time.time()
|
||||
assert src.stalled() is False
|
||||
assert src.stats()["streaming"] is True
|
||||
|
||||
|
||||
def test_an_old_frame_on_an_open_socket_is_a_stall():
|
||||
src = _source()
|
||||
src.connected = True
|
||||
src._frame_ts = time.time() - (capture.STALL_AFTER_S + 1)
|
||||
assert src.stalled() is True
|
||||
stats = src.stats()
|
||||
# The distinction that matters: still connected, no longer streaming.
|
||||
assert stats["connected"] is True
|
||||
assert stats["streaming"] is False
|
||||
assert stats["stalled"] is True
|
||||
|
||||
|
||||
def test_a_disconnected_camera_is_reported_as_down_not_stalled():
|
||||
"""Two states, opposite actions: check the network vs. the camera is
|
||||
answering and sending nothing. They must never share a verdict."""
|
||||
src = _source()
|
||||
src.connected = False
|
||||
src._frame_ts = time.time() - 3600
|
||||
assert src.stalled() is False
|
||||
assert src.stats()["streaming"] is False
|
||||
|
||||
|
||||
# -- a wrong address must not cost 30 seconds ---------------------------
|
||||
|
||||
def test_unreachable_source_fails_fast_with_a_reason():
|
||||
"""_open() used to hand an unroutable address straight to OpenCV, which
|
||||
blocks ~30s inside the constructor and cannot be interrupted by stop().
|
||||
198.51.100.0/24 is TEST-NET-2 and routes nowhere.
|
||||
"""
|
||||
src = VideoSource("cam", "rtsp://198.51.100.9:554/x")
|
||||
started = time.time()
|
||||
assert src._open() is None
|
||||
elapsed = time.time() - started
|
||||
assert elapsed < 8.0, f"pre-flight took {elapsed:.1f}s"
|
||||
assert src.last_error, "a failed open must say why"
|
||||
assert "198.51.100.9" in src.last_error
|
||||
assert src.stats()["last_error"] == src.last_error
|
||||
|
||||
|
||||
def test_webcam_sources_skip_the_preflight():
|
||||
"""An int source is a local device with no host to reach; the check must
|
||||
pass it through rather than refuse it."""
|
||||
ok, why = capture._tcp_reachable(0, 1.0)
|
||||
assert ok is True and why == ""
|
||||
Reference in New Issue
Block a user