Files
Behavision/desktop/stream_remote.go
Suriyakumarvijayanayagam 48a30d97db Live camera view in the app, and the green light that was lying about it
Two changes, and the second was found by verifying the first.

## Watching a camera from the app, in another building

Snapshots answer "is that camera working". They do not answer "what is
happening in my shop right now", which is what somebody who opens the app away
from the counter is asking. Head office's browser already had that answer -
LiveHub plus cameras.Live, where the shop PC asks outbound whether anybody is
watching and pushes JPEG frames for as long as somebody is - and the app could
not reach it.

cloud.CameraLive opens that feed and the app's own loopback relay re-emits it
as multipart MJPEG. That is the trick: frames arrive base64 over SSE, an <img>
cannot render that, and an <img> renders MJPEG natively - so a tile is an
ordinary <img> pointed at loopback whether the camera is in this room or
another city.

- Reconnecting happens in the relay, not the page. The server caps one push at
  five minutes, so doing it here means the <img> never sees the stream end.
- The headers are flushed before the first frame. Go writes them on the first
  body write, so without that the whole response waits for the shop PC to
  start pushing. Measured against production: 30 seconds and not even a
  Content-Type, which surfaces as the request timing out.
- One camera at a time. Watching makes a shop PC upload, so a grid that went
  live at once would put an estate's worth of cameras on the wire because
  somebody opened a page.
- live.mjpeg is behind the same per-run token as the engine routes, and a
  wrong token is a 404 that never reaches head office at all.
- CameraLive uses its own HTTP client: the shared one's 30s timeout covers the
  whole response and would sever a working view every thirty seconds - the
  trap that made the server set WriteTimeout to zero for its own SSE endpoint.

## A camera read "Connected" for 34 minutes after the shop PC went blind

Which is why the verification above looked like a failure: head office
registered the viewer and no frame ever came.

reportWith returns early when the engine is unreachable - correctly, it has
nothing to say - so the last state it sent stays in the database looking
current. Measured live: cam2 and entrance both reading Connected, in green,
with last_seen_at 34 minutes old, while the heartbeat from the same PC said
cameras_up 0 of 0. Two surfaces reading two stored fields and disagreeing.

false could not be the answer. It means "this camera is not connecting", which
sends an installer to check cabling on a camera that was working perfectly the
last time anybody could ask it. So there are four states and one function:

  connected       reported recently, and working
  not_connecting  reported recently, and the stream will not open
  waiting         no shop PC has ever reported this camera
  stale           reported once, and not lately

- Connected is CLEARED when stale or waiting. A stale true left in place stays
  available to every client reading the field directly, and leaves two fields
  on one object disagreeing - how the shops screen once came out labelled
  Working, in green, above "2 of 3 cameras not connecting".
- Computed in scanCamera, so every camera anybody reads passes through it. A
  state computed per handler is one a handler forgets, and this had already
  reached three screens.
- CameraStaleAfter is 5 minutes: five missed reports, not one. Same reasoning
  as three missed heartbeats - an indicator that cries wolf gets ignored.
- An unparseable last_seen_at is stale. It should be impossible, which is why
  it must not fall through to the state that says everything is fine.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-30 16:48:43 +05:30

150 lines
5.0 KiB
Go

package main
// Watching a camera in another building, from the app.
//
// The shop PC sits behind a router with no inbound route, so nothing here can
// pull its MJPEG stream - that stream is served on the shop PC's own loopback
// and always will be. Head office's LiveHub is the way round it: the agent
// asks outbound whether anyone is watching and pushes JPEG frames up for
// exactly as long as somebody is. The head-office web app already consumes
// that; this is the same feed, for the app.
//
// It arrives as base64 frames over SSE, which an <img> cannot render, so this
// re-emits them as multipart MJPEG - which an <img> renders natively, through
// the relay that already exists for the local engine. That is what keeps ONE
// code path in the screens: a tile points at a loopback URL and does not know
// or care which building the picture came from.
import (
"bufio"
"context"
"encoding/base64"
"fmt"
"net/http"
"strings"
"time"
)
// The boundary is ours to choose; it only has to be a string the JPEG bytes
// cannot contain, and a marker line never appears inside JPEG data.
const mjpegBoundary = "behavisionframe"
// A frame is base64, so ~1.33 bytes on the wire per byte of picture. The
// engine re-encodes to 640 px for the relay and those measure ~20 KB, so this
// is roughly a hundredfold headroom - large enough never to clip a real frame
// and small enough that a broken or hostile stream cannot grow this process's
// memory without bound.
const maxFrameLine = 8 << 20
func (p *streamProxy) relayRemote(w http.ResponseWriter, r *http.Request,
cameraID string, open func(context.Context, string) (*http.Response, error)) {
w.Header().Set("Content-Type", "multipart/x-mixed-replace; boundary="+mjpegBoundary)
w.Header().Set("Cache-Control", "no-store")
flusher, _ := w.(http.Flusher)
// Send the headers NOW, before any frame exists. Go writes them on the
// first body write, so without this the whole response - status line
// included - waits for the shop computer to start pushing, and a viewer
// whose camera is slow to answer sees the REQUEST time out rather than a
// stream that has not painted yet. Measured against production: 30
// seconds and not even a Content-Type.
if flusher != nil {
flusher.Flush()
}
// Reconnecting is normal, not an error. The server caps one push at five
// minutes so that a tab left open for a week cannot leave a shop
// uploading for a week - so a viewer who IS still there simply asks
// again. Doing it here rather than in the page is what lets the <img>
// survive the cap: it never sees the stream end.
sent := 0
for {
if r.Context().Err() != nil {
return
}
n, err := p.pumpRemote(w, flusher, r.Context(), cameraID, open)
sent += n
if r.Context().Err() != nil {
return
}
// Nothing was written and the attempt failed. Writing an error body
// now would be writing it into a multipart stream the <img> is
// already parsing, so the picture simply stays on whatever it last
// showed and the screen's own "not connecting" state is the report.
if err != nil && sent == 0 {
return
}
select {
case <-r.Context().Done():
return
case <-time.After(1500 * time.Millisecond):
}
}
}
// pumpRemote runs one SSE connection to exhaustion and returns how many
// frames it forwarded.
func (p *streamProxy) pumpRemote(w http.ResponseWriter, flusher http.Flusher,
ctx context.Context, cameraID string,
open func(context.Context, string) (*http.Response, error)) (int, error) {
resp, err := open(ctx, cameraID)
if err != nil {
return 0, err
}
defer resp.Body.Close()
sc := bufio.NewScanner(resp.Body)
sc.Buffer(make([]byte, 0, 64*1024), maxFrameLine)
var event, data string
frames := 0
for sc.Scan() {
line := sc.Text()
switch {
case strings.HasPrefix(line, "event: "):
event = strings.TrimSpace(line[7:])
case strings.HasPrefix(line, "data: "):
data = line[6:]
case line == "":
// End of one SSE event. `waiting` means head office has us
// registered and the shop PC has not started pushing yet - a real
// second or two while the agent is asked, and nothing to draw.
if event == "frame" && data != "" {
if err := writeMJPEGFrame(w, flusher, data); err != nil {
return frames, err // the webview went away
}
frames++
}
event, data = "", ""
}
}
return frames, sc.Err()
}
func writeMJPEGFrame(w http.ResponseWriter, flusher http.Flusher, b64 string) error {
jpg, err := base64.StdEncoding.DecodeString(b64)
if err != nil || len(jpg) == 0 {
// One malformed frame is not a reason to tear down a working view.
return nil
}
if _, err := fmt.Fprintf(w,
"--%s\r\nContent-Type: image/jpeg\r\nContent-Length: %d\r\n\r\n",
mjpegBoundary, len(jpg)); err != nil {
return err
}
if _, err := w.Write(jpg); err != nil {
return err
}
if _, err := w.Write([]byte("\r\n")); err != nil {
return err
}
// Flushed per frame. Anything held waiting for a full buffer is a tile
// that stays blank, which is indistinguishable from the view not working.
if flusher != nil {
flusher.Flush()
}
return nil
}