The engine was searching an empty room fifteen times a second
Measured rather than guessed, and the first guess was wrong. Wall clock said H.265 decode cost 58 ms a frame; cap.read() blocks until the next frame arrives, so that was the frame interval, not work. As CPU time: decode 3.7 ms, detection 31.0 ms - and detection ran on every frame whether or not anything was in front of the camera, 6,649 of 8,634 frames with faces_seen 0 and active_tracks 0 throughout. detect_threads: OpenCV spreads a small repeated job over eight threads, costing 31.0 ms of CPU for 8.9 ms of wall. One thread costs 15.3 ms for 15.3 ms, against a 66 ms budget at 15 fps. Half the CPU for latency nothing can notice. motion_gate: a 160x90 greyscale absdiff, 0.1 ms against detection's 15. Consulted only while no track is open; forced to look every motion_max_skip frames; compared against the last frame SEARCHED so a slow drift cannot creep under the threshold; and a threshold above this camera's measured noise and far below a person, so anything ambiguous detects. tests/test_motion_gate.py pins each of those rather than the saving, including asserting the longest run of skips rather than the total - counting the total would pass a gate that slept forty frames and then looked forty times. Together 80% -> 16% of a core, detection skipped on 92% of frames. faces_seen is still 0 and the gate is not why: run directly over the same frames the detector finds nothing at threshold 0.50 either. The placement is the limit, as recorded; the CPU was being spent to rediscover that fifteen times a second. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
59
CLAUDE.md
59
CLAUDE.md
@@ -2226,6 +2226,65 @@ person does not re-file them:
|
||||
in the product a caller could reasonably guess wrong, and it is a route no
|
||||
client app should ever call.
|
||||
|
||||
## CPU: the engine was searching an empty room 15 times a second
|
||||
|
||||
Measured on this Mac against the office camera, because "it feels hot" is not
|
||||
a number. The first reading was **214% of a core**, and the first guess -
|
||||
H.265 decode - was wrong. Wall-clock time said decode cost 58 ms a frame, but
|
||||
`cap.read()` BLOCKS until the next frame arrives, so that number was the frame
|
||||
interval, not work. Measured as CPU time instead:
|
||||
|
||||
```
|
||||
wall/frame CPU/frame at 15 fps
|
||||
H.265 decode 58.8 ms 3.7 ms 5% of a core
|
||||
YuNet detect 8.9 ms 31.0 ms 47% of a core
|
||||
```
|
||||
|
||||
Detection costs three times its wall time because OpenCV spreads it over eight
|
||||
threads. Decode is nearly free. So the cost is detection, and it was running on
|
||||
**every frame whether or not anything was in front of the camera** - 6,649 of
|
||||
8,634 frames searched, with `faces_seen: 0` and `active_tracks: 0` throughout.
|
||||
|
||||
Two changes, both measured:
|
||||
|
||||
- **`app.detect_threads: 1`.** OpenCV sizes its pool for one big job on an idle
|
||||
machine; this is a small job repeated forever on a machine also running the
|
||||
recogniser, the tracker and possibly three other cameras. One thread costs
|
||||
15.3 ms of CPU against the default's 31.0 ms, for 6 ms more wall time against
|
||||
a 66 ms frame budget. Half the CPU, no latency that matters.
|
||||
- **`app.motion_gate`.** A 160x90 greyscale thumbnail and an `absdiff`: 0.1 ms
|
||||
against detection's 15 ms, ninety times cheaper. A shop is empty most of the
|
||||
day and an empty room costs exactly as much to search as a busy one.
|
||||
|
||||
Together: **80% -> 16% of a core**, with detection skipped on 92% of frames.
|
||||
Against the original main-stream reading that is 214% -> 16%.
|
||||
|
||||
### Why the gate cannot lose a face
|
||||
|
||||
Cheapness is easy; this is the part that makes it acceptable, and
|
||||
`tests/test_motion_gate.py` is the argument written down rather than asserted.
|
||||
|
||||
- It is only consulted while **no track is open**, so a person already being
|
||||
followed is never subject to it.
|
||||
- `motion_max_skip` forces a real detection about once a second whatever the
|
||||
thumbnail says. The test asserts the longest *run* of skips, not the total:
|
||||
what matters is the worst case a person can fall into, and counting the total
|
||||
would pass a gate that skipped forty frames and then looked forty times.
|
||||
- The comparison is against the last frame actually **searched**, not the last
|
||||
frame seen, so a slow drift accumulates and trips the gate instead of sliding
|
||||
under it one frame at a time. Someone easing into view slowly would otherwise
|
||||
be invisible indefinitely.
|
||||
- The threshold (1.0 mean absolute difference) sits above this camera's
|
||||
measured noise floor (~0.3) and far below a person. Anything ambiguous falls
|
||||
through to detection: when in doubt it looks.
|
||||
|
||||
Verified on the live camera after the change: `faces_seen: 0` - and the gate is
|
||||
not why. Running the detector directly over the same frames finds **0 faces at
|
||||
threshold 0.50**, let alone 0.82. The people in view are seated, side-on and
|
||||
far away, which is the same `fraction_below_gate: 0.59` this file already
|
||||
records. The placement is still the limit; the CPU was simply being spent to
|
||||
discover that 15 times a second.
|
||||
|
||||
## Setting up on a new machine
|
||||
|
||||
1. Copy the `Behavision` folder **including `.env`** (gitignored, holds
|
||||
|
||||
Reference in New Issue
Block a user