Live camera view in the app, and the green light that was lying about it

Two changes, and the second was found by verifying the first.

## Watching a camera from the app, in another building

Snapshots answer "is that camera working". They do not answer "what is
happening in my shop right now", which is what somebody who opens the app away
from the counter is asking. Head office's browser already had that answer -
LiveHub plus cameras.Live, where the shop PC asks outbound whether anybody is
watching and pushes JPEG frames for as long as somebody is - and the app could
not reach it.

cloud.CameraLive opens that feed and the app's own loopback relay re-emits it
as multipart MJPEG. That is the trick: frames arrive base64 over SSE, an <img>
cannot render that, and an <img> renders MJPEG natively - so a tile is an
ordinary <img> pointed at loopback whether the camera is in this room or
another city.

- Reconnecting happens in the relay, not the page. The server caps one push at
  five minutes, so doing it here means the <img> never sees the stream end.
- The headers are flushed before the first frame. Go writes them on the first
  body write, so without that the whole response waits for the shop PC to
  start pushing. Measured against production: 30 seconds and not even a
  Content-Type, which surfaces as the request timing out.
- One camera at a time. Watching makes a shop PC upload, so a grid that went
  live at once would put an estate's worth of cameras on the wire because
  somebody opened a page.
- live.mjpeg is behind the same per-run token as the engine routes, and a
  wrong token is a 404 that never reaches head office at all.
- CameraLive uses its own HTTP client: the shared one's 30s timeout covers the
  whole response and would sever a working view every thirty seconds - the
  trap that made the server set WriteTimeout to zero for its own SSE endpoint.

## A camera read "Connected" for 34 minutes after the shop PC went blind

Which is why the verification above looked like a failure: head office
registered the viewer and no frame ever came.

reportWith returns early when the engine is unreachable - correctly, it has
nothing to say - so the last state it sent stays in the database looking
current. Measured live: cam2 and entrance both reading Connected, in green,
with last_seen_at 34 minutes old, while the heartbeat from the same PC said
cameras_up 0 of 0. Two surfaces reading two stored fields and disagreeing.

false could not be the answer. It means "this camera is not connecting", which
sends an installer to check cabling on a camera that was working perfectly the
last time anybody could ask it. So there are four states and one function:

  connected       reported recently, and working
  not_connecting  reported recently, and the stream will not open
  waiting         no shop PC has ever reported this camera
  stale           reported once, and not lately

- Connected is CLEARED when stale or waiting. A stale true left in place stays
  available to every client reading the field directly, and leaves two fields
  on one object disagreeing - how the shops screen once came out labelled
  Working, in green, above "2 of 3 cameras not connecting".
- Computed in scanCamera, so every camera anybody reads passes through it. A
  state computed per handler is one a handler forgets, and this had already
  reached three screens.
- CameraStaleAfter is 5 minutes: five missed reports, not one. Same reasoning
  as three missed heartbeats - an indicator that cries wolf gets ignored.
- An unparseable last_seen_at is stale. It should be impossible, which is why
  it must not fall through to the state that says everything is fine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
This commit is contained in:
2026-09-30 16:48:43 +05:30
parent ecc8bbba6f
commit 48a30d97db
20 changed files with 999 additions and 79 deletions

View File

@@ -3359,3 +3359,80 @@ live one.
- **With no engine AND nobody signed in, the engine error is still the answer.**
There is nothing else to show and the person is most likely setting this PC
up; naming head office there points them at a step they have not reached.
## Watching a camera from the app, in another building
Snapshots answer *"is that camera working"*. They do not answer *"what is
happening in my shop right now"*, which is what somebody who opens the app
away from the counter is asking. Head office's browser already had the answer
— `LiveHub` plus `cameras.Live`, where the shop PC asks outbound whether
anybody is watching and pushes JPEG frames up for exactly as long as somebody
is — and the app could not reach it.
`cloud.CameraLive` opens that feed and the app's own loopback relay re-emits
it as **multipart MJPEG**, which is the whole trick: frames arrive base64 over
SSE, an `<img>` cannot render that, and an `<img>` renders MJPEG natively. So a
tile is an ordinary `<img>` pointed at loopback whether the camera is in this
room or another city, and no screen has to know which.
- **Reconnecting happens in the relay, not the page.** The server caps one push
at five minutes so a tab left open for a week cannot leave a shop uploading
for a week. Doing it here means the `<img>` never sees the stream end.
- **The headers are flushed before the first frame.** Go writes them on the
first body write, so without that the whole response — status line included —
waits for the shop PC to start pushing. Measured against production: thirty
seconds and not even a `Content-Type`, which surfaces as the *request* timing
out rather than a stream that has not painted yet.
- **One camera at a time.** Watching makes a shop PC upload, so a grid that
went live at once would put an estate's worth of cameras on the wire because
somebody opened a page. `Watch live` is per tile and toggles the previous one
off.
- **`live.mjpeg` is behind the same per-run token as the engine routes**, and a
wrong token is a 404 that never reaches head office at all. It is a live view
of a shop floor; the relay being on loopback is not on its own a control.
- **`CameraLive` uses its own HTTP client.** The shared one has a 30-second
timeout that covers the whole response and would therefore sever a working
live view every thirty seconds — the same trap that made the server set
`WriteTimeout` to zero for its own SSE endpoint.
## A camera read "Connected" for 34 minutes after the shop PC went blind
Found while verifying the live view against production, and it is the reason
that verification looked like a failure: head office registered the viewer and
no frame ever came.
`reportWith` returns early when the engine is unreachable — correctly, because
it has nothing to say — so the last state it sent **stays in the database
looking current**. Measured on the live estate: `cam2` and `entrance` both
reading **Connected**, in green, with `last_seen_at` thirty-four minutes old,
while the heartbeat from the same PC said `cameras_up: 0, cameras_total: 0`.
Two surfaces reading two stored fields and disagreeing about one fact.
`false` could not be the answer. It means *"this camera is not connecting"*,
which sends an installer to check cabling on a camera that was working
perfectly the last time anybody could ask it. So there are four states, not
three, and `api.CameraState` is the one function that decides them:
| state | meaning | what to do |
|---|---|---|
| `connected` | reported within `CameraStaleAfter`, and working | — |
| `not_connecting` | reported recently, and the stream will not open | check the address, password, cabling |
| `waiting` | no shop PC has ever reported this camera | it has not reached the PC yet |
| `stale` | reported once, and not lately | check the PC is on and Behavision is running |
- **`Connected` is CLEARED when the state is `stale` or `waiting`.** Leaving a
stale `true` in place keeps the lie available to every client that reads the
field directly — a mobile app, a script, an older desktop build — and leaves
two fields on one object disagreeing, which is exactly how the shops screen
once came out labelled **Working**, in green, above *"2 of 3 cameras not
connecting"*.
- **It is computed in `scanCamera`**, so every camera anybody reads passes
through it. A state computed per handler is a state one handler forgets, and
this one had already reached three screens.
- **`CameraStaleAfter` is 5 minutes — five missed reports, not one.** The agent
reports on a 60-second tick, so one miss is a dropped packet. Same reasoning
as a site being offline after three missed heartbeats: an indicator that
cries wolf is one people learn to ignore.
- **An unparseable `last_seen_at` is stale**, not connected. It should be
impossible, which is precisely why it must not fall through to the state that
says everything is fine.