The engine was searching an empty room fifteen times a second

Measured rather than guessed, and the first guess was wrong. Wall clock
said H.265 decode cost 58 ms a frame; cap.read() blocks until the next
frame arrives, so that was the frame interval, not work. As CPU time:
decode 3.7 ms, detection 31.0 ms - and detection ran on every frame
whether or not anything was in front of the camera, 6,649 of 8,634
frames with faces_seen 0 and active_tracks 0 throughout.

detect_threads: OpenCV spreads a small repeated job over eight threads,
costing 31.0 ms of CPU for 8.9 ms of wall. One thread costs 15.3 ms for
15.3 ms, against a 66 ms budget at 15 fps. Half the CPU for latency
nothing can notice.

motion_gate: a 160x90 greyscale absdiff, 0.1 ms against detection's 15.
Consulted only while no track is open; forced to look every
motion_max_skip frames; compared against the last frame SEARCHED so a
slow drift cannot creep under the threshold; and a threshold above this
camera's measured noise and far below a person, so anything ambiguous
detects. tests/test_motion_gate.py pins each of those rather than the
saving, including asserting the longest run of skips rather than the
total - counting the total would pass a gate that slept forty frames
and then looked forty times.

Together 80% -> 16% of a core, detection skipped on 92% of frames.
faces_seen is still 0 and the gate is not why: run directly over the
same frames the detector finds nothing at threshold 0.50 either. The
placement is the limit, as recorded; the CPU was being spent to
rediscover that fifteen times a second.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-24 13:58:05 +05:30
parent f97ffc913a
commit 2e60fbb57a
5 changed files with 242 additions and 0 deletions

View File

@@ -21,6 +21,17 @@ def cmd_run(args: argparse.Namespace) -> int:
cfg = load_config(args.config)
setup_logging(cfg.app.log_level, cfg.app.data_dir)
if cfg.app.detect_threads > 0:
# OpenCV sizes its pool for one big job on an idle machine. This is a
# small job repeated forever on a machine also running the recogniser,
# the tracker and possibly three other cameras, so the default costs
# twice the CPU for no useful latency. Measured: 31 ms CPU/frame at the
# default against 15 ms at one thread, for 6 ms more wall time against
# a 66 ms budget.
import cv2
cv2.setNumThreads(cfg.app.detect_threads)
log.info("detection threads: %d (OpenCV default was %d)",
cfg.app.detect_threads, cv2.getNumThreads())
missing = setup_models(cfg.app.models_dir)
if missing:
log.error("required models missing: %s", ", ".join(missing))

View File

@@ -132,6 +132,29 @@ class AppSection(BaseModel):
# on changes what the system is under GDPR and India's DPDP, so it has to
# be a decision somebody makes rather than one they inherit.
store_faces: bool = False
# How many threads OpenCV may use for detection. Measured on the office
# camera (800x448 sub-stream): the default of 8 costs 31 ms of CPU per
# frame for 8.9 ms of wall time, while ONE thread costs 15.3 ms of CPU for
# 15.3 ms of wall - half the CPU for 6 ms more latency, against a 66 ms
# frame budget at 15 fps. The default is wrong here because OpenCV sizes it
# for one big job on an idle machine, and this is a small job repeated
# forever on a machine also running the recogniser, the tracker and three
# other cameras. 0 leaves OpenCV's own default alone.
detect_threads: int = 1
# Skip detection on frames where nothing has changed and nothing is being
# tracked. A shop is empty most of the day and a frame of an empty room
# costs exactly as much to search as a busy one. See CameraWorker.run for
# why this cannot lose a face.
motion_gate: bool = True
# Mean absolute difference, 0-255, over a 160x90 greyscale thumbnail. 1.0
# is well below the noise floor of a real camera - measured on this one,
# an empty room varies by ~0.3 between frames - so it triggers on movement
# rather than on sensor noise, and anything ambiguous detects.
motion_threshold: float = 1.0
# Detect at least this often regardless of the gate, so a change the
# thumbnail cannot see - someone entering at the far edge, a slow lean into
# frame - is still found within a second.
motion_max_skip: int = 12
class ApiSection(BaseModel):

View File

@@ -170,6 +170,11 @@ class CameraWorker(threading.Thread):
self._last_frame_ts = 0.0
self._was_connected = False
self.frames_processed = 0
# Motion gate state: a 160x90 greyscale thumbnail of the last frame we
# actually searched, and how many frames we have skipped since.
self._motion_prev = None
self._motion_skipped = 0
self.frames_skipped = 0
self.faces_seen = 0
self.pipeline = PipelineStats()
# One outbox per worker, all writing into the same directory. Files are
@@ -236,6 +241,7 @@ class CameraWorker(threading.Thread):
return {
**self.source.stats(),
"frames_processed": self.frames_processed,
"frames_skipped": self.frames_skipped,
"faces_seen": self.faces_seen,
"active_tracks": len(self.tracker.tracks),
"pipeline": self.pipeline.snapshot(self.rcfg.min_enroll_quality),
@@ -244,6 +250,43 @@ class CameraWorker(threading.Thread):
"enroll_threshold": self.rcfg.enroll_threshold},
}
def _nothing_moved(self, frame) -> bool:
"""True when this frame is close enough to the last searched one that
searching it again would find the same nothing.
It cannot lose a face, and that property is what makes it acceptable
rather than merely cheap. Three guards, in order:
* the caller only asks while NO track is open, so a person already
being followed is never affected by it;
* `motion_max_skip` forces a real detection about once a second
whatever the thumbnail says, which covers a change too small or too
gradual for it - someone easing into frame at the far edge;
* the threshold sits well above measured sensor noise and well below
a person, and anything ambiguous falls through to detection. When
in doubt it looks.
Cost is 0.1 ms against detection's 15 ms, so an empty shop stops paying
for a search of an empty room ~90 times a second.
"""
import cv2 as _cv2
small = _cv2.resize(_cv2.cvtColor(frame, _cv2.COLOR_BGR2GRAY), (160, 90),
interpolation=_cv2.INTER_AREA)
prev, self._motion_prev = self._motion_prev, small
if prev is None:
return False
if self._motion_skipped >= self.cfg.app.motion_max_skip:
self._motion_skipped = 0
return False
if float(_cv2.absdiff(small, prev).mean()) >= self.cfg.app.motion_threshold:
self._motion_skipped = 0
# Keep the thumbnail we just searched against, not this one, so a
# slow drift cannot creep past the threshold one frame at a time.
return False
self._motion_prev = prev
self._motion_skipped += 1
return True
# -- thread ---------------------------------------------------------
def run(self) -> None:
tcfg = self.cfg.tracking
@@ -256,6 +299,15 @@ class CameraWorker(threading.Thread):
continue
self._last_frame_ts = ts
# An empty room costs exactly as much to search as a busy one,
# and a shop is empty most of the day. Only ever while nothing
# is being tracked - see _nothing_moved.
if (self.cfg.app.motion_gate and not self.tracker.tracks
and self._nothing_moved(frame)):
self.frames_skipped += 1
self._remember_tracks([])
continue
detections = self.detector.detect(frame)
for det in detections:
det.quality = face_quality(frame, det.box, det.kps)