The engine was searching an empty room fifteen times a second

Measured rather than guessed, and the first guess was wrong. Wall clock
said H.265 decode cost 58 ms a frame; cap.read() blocks until the next
frame arrives, so that was the frame interval, not work. As CPU time:
decode 3.7 ms, detection 31.0 ms - and detection ran on every frame
whether or not anything was in front of the camera, 6,649 of 8,634
frames with faces_seen 0 and active_tracks 0 throughout.

detect_threads: OpenCV spreads a small repeated job over eight threads,
costing 31.0 ms of CPU for 8.9 ms of wall. One thread costs 15.3 ms for
15.3 ms, against a 66 ms budget at 15 fps. Half the CPU for latency
nothing can notice.

motion_gate: a 160x90 greyscale absdiff, 0.1 ms against detection's 15.
Consulted only while no track is open; forced to look every
motion_max_skip frames; compared against the last frame SEARCHED so a
slow drift cannot creep under the threshold; and a threshold above this
camera's measured noise and far below a person, so anything ambiguous
detects. tests/test_motion_gate.py pins each of those rather than the
saving, including asserting the longest run of skips rather than the
total - counting the total would pass a gate that slept forty frames
and then looked forty times.

Together 80% -> 16% of a core, detection skipped on 92% of frames.
faces_seen is still 0 and the gate is not why: run directly over the
same frames the detector finds nothing at threshold 0.50 either. The
placement is the limit, as recorded; the CPU was being spent to
rediscover that fifteen times a second.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-24 13:58:05 +05:30
parent f97ffc913a
commit 2e60fbb57a
5 changed files with 242 additions and 0 deletions

View File

@@ -2226,6 +2226,65 @@ person does not re-file them:
in the product a caller could reasonably guess wrong, and it is a route no
client app should ever call.
## CPU: the engine was searching an empty room 15 times a second
Measured on this Mac against the office camera, because "it feels hot" is not
a number. The first reading was **214% of a core**, and the first guess -
H.265 decode - was wrong. Wall-clock time said decode cost 58 ms a frame, but
`cap.read()` BLOCKS until the next frame arrives, so that number was the frame
interval, not work. Measured as CPU time instead:
```
wall/frame CPU/frame at 15 fps
H.265 decode 58.8 ms 3.7 ms 5% of a core
YuNet detect 8.9 ms 31.0 ms 47% of a core
```
Detection costs three times its wall time because OpenCV spreads it over eight
threads. Decode is nearly free. So the cost is detection, and it was running on
**every frame whether or not anything was in front of the camera** - 6,649 of
8,634 frames searched, with `faces_seen: 0` and `active_tracks: 0` throughout.
Two changes, both measured:
- **`app.detect_threads: 1`.** OpenCV sizes its pool for one big job on an idle
machine; this is a small job repeated forever on a machine also running the
recogniser, the tracker and possibly three other cameras. One thread costs
15.3 ms of CPU against the default's 31.0 ms, for 6 ms more wall time against
a 66 ms frame budget. Half the CPU, no latency that matters.
- **`app.motion_gate`.** A 160x90 greyscale thumbnail and an `absdiff`: 0.1 ms
against detection's 15 ms, ninety times cheaper. A shop is empty most of the
day and an empty room costs exactly as much to search as a busy one.
Together: **80% -> 16% of a core**, with detection skipped on 92% of frames.
Against the original main-stream reading that is 214% -> 16%.
### Why the gate cannot lose a face
Cheapness is easy; this is the part that makes it acceptable, and
`tests/test_motion_gate.py` is the argument written down rather than asserted.
- It is only consulted while **no track is open**, so a person already being
followed is never subject to it.
- `motion_max_skip` forces a real detection about once a second whatever the
thumbnail says. The test asserts the longest *run* of skips, not the total:
what matters is the worst case a person can fall into, and counting the total
would pass a gate that skipped forty frames and then looked forty times.
- The comparison is against the last frame actually **searched**, not the last
frame seen, so a slow drift accumulates and trips the gate instead of sliding
under it one frame at a time. Someone easing into view slowly would otherwise
be invisible indefinitely.
- The threshold (1.0 mean absolute difference) sits above this camera's
measured noise floor (~0.3) and far below a person. Anything ambiguous falls
through to detection: when in doubt it looks.
Verified on the live camera after the change: `faces_seen: 0` - and the gate is
not why. Running the detector directly over the same frames finds **0 faces at
threshold 0.50**, let alone 0.82. The people in view are seated, side-on and
far away, which is the same `fraction_below_gate: 0.59` this file already
records. The placement is still the limit; the CPU was simply being spent to
discover that 15 times a second.
## Setting up on a new machine
1. Copy the `Behavision` folder **including `.env`** (gitignored, holds