The live picture was chained to the recognition pipeline

Reported from the first Windows install: the camera feed lags. It did,
and not because of the network, the proxy or the webview.

The MJPEG stream served _annotated_jpeg - the frame the pipeline had
most recently FINISHED with, encoded after detection, quality scoring,
tracking and identification had all run on it. On a modest shop PC that
is a few frames a second, and every picture was already as old as that
processing. It looked like lag because it was lag. On the fast machine
it was developed on the pipeline kept up with the stream's own 10 fps
cap, which is why nobody here ever saw it.

Two more things compounded it. Every processed frame was JPEG-encoded
whether or not a viewer existed - CPU spent on precisely the machine
short of it. And ffmpeg ran its RTSP demuxer with default buffering,
which holds a comfortable queue of frames before handing over the first:
half a second to two seconds a live view can never recover.

Now the picture and the boxes are decoupled. latest_jpeg_since takes the
capture thread's freshest frame at the camera's own rate and draws the
boxes from the last processed frame over it - encoded on demand, per
request, so a camera nobody watches costs no encode at all. The stream
sends a frame only when the camera has a newer one, capped at 15 fps;
nothing is sent twice. Boxes older than a second are not drawn, so a
stalled pipeline cannot leave one floating over an empty spot.
_publish_annotated becomes _remember_tracks: a handful of tuples under
the lock, no copy, no encode. ffmpeg gets nobuffer / low_delay /
max_delay.

Measured on cam2's sub-stream, same machine, ten seconds each:

  before   99 frames sent,  98 distinct    9.8 new pictures/s
  after   141 frames sent, 141 distinct   14.0 new pictures/s

against a 15 fps camera, with the pipeline still processing 166 of 181
captured frames alongside - and engine CPU DOWN from 90% with no viewer
to 62% with one attached.

Engine version 1.0.0 -> 1.1.0 so a re-run of setup reinstalls it rather
than pip deciding the requirement is already satisfied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-11 16:28:05 +05:30
parent b59e667a68
commit e262fc8482
6 changed files with 230 additions and 23 deletions

View File

@@ -6,6 +6,7 @@ models finish loading without a single unguarded None dereference.
from __future__ import annotations
import asyncio
import time
import logging
import secrets
from pathlib import Path
@@ -370,11 +371,23 @@ def create_app(engine: Engine) -> FastAPI:
# Stop when the camera is deleted or its worker dies - otherwise a
# removed camera leaves this generator running for the life of the
# process, holding a reference to a worker nothing else can see.
# Driven by the camera, not a timer: a frame goes out when the
# capture thread has one newer than the last one sent, so nothing
# is sent twice and nothing waits on the recognition pipeline.
# Capped at 15 fps - the office cameras' own rate - so a viewer
# never costs more encodes than the camera produces pictures.
last_ts, min_gap, sent_at = 0.0, 1.0 / 15, 0.0
while engine.workers.get(camera_id) is worker and worker.is_alive():
jpeg = worker.latest_jpeg()
if jpeg is not None:
yield boundary + jpeg + b"\r\n"
await asyncio.sleep(0.1) # ~10 fps to the browser
now = time.time()
if now - sent_at < min_gap:
await asyncio.sleep(min_gap - (now - sent_at))
continue
jpeg, ts = worker.latest_jpeg_since(last_ts)
if jpeg is None:
await asyncio.sleep(0.02)
continue
last_ts, sent_at = ts, time.time()
yield boundary + jpeg + b"\r\n"
return StreamingResponse(
generate(),