Two changes, and the second was found by verifying the first. ## Watching a camera from the app, in another building Snapshots answer "is that camera working". They do not answer "what is happening in my shop right now", which is what somebody who opens the app away from the counter is asking. Head office's browser already had that answer - LiveHub plus cameras.Live, where the shop PC asks outbound whether anybody is watching and pushes JPEG frames for as long as somebody is - and the app could not reach it. cloud.CameraLive opens that feed and the app's own loopback relay re-emits it as multipart MJPEG. That is the trick: frames arrive base64 over SSE, an <img> cannot render that, and an <img> renders MJPEG natively - so a tile is an ordinary <img> pointed at loopback whether the camera is in this room or another city. - Reconnecting happens in the relay, not the page. The server caps one push at five minutes, so doing it here means the <img> never sees the stream end. - The headers are flushed before the first frame. Go writes them on the first body write, so without that the whole response waits for the shop PC to start pushing. Measured against production: 30 seconds and not even a Content-Type, which surfaces as the request timing out. - One camera at a time. Watching makes a shop PC upload, so a grid that went live at once would put an estate's worth of cameras on the wire because somebody opened a page. - live.mjpeg is behind the same per-run token as the engine routes, and a wrong token is a 404 that never reaches head office at all. - CameraLive uses its own HTTP client: the shared one's 30s timeout covers the whole response and would sever a working view every thirty seconds - the trap that made the server set WriteTimeout to zero for its own SSE endpoint. ## A camera read "Connected" for 34 minutes after the shop PC went blind Which is why the verification above looked like a failure: head office registered the viewer and no frame ever came. reportWith returns early when the engine is unreachable - correctly, it has nothing to say - so the last state it sent stays in the database looking current. Measured live: cam2 and entrance both reading Connected, in green, with last_seen_at 34 minutes old, while the heartbeat from the same PC said cameras_up 0 of 0. Two surfaces reading two stored fields and disagreeing. false could not be the answer. It means "this camera is not connecting", which sends an installer to check cabling on a camera that was working perfectly the last time anybody could ask it. So there are four states and one function: connected reported recently, and working not_connecting reported recently, and the stream will not open waiting no shop PC has ever reported this camera stale reported once, and not lately - Connected is CLEARED when stale or waiting. A stale true left in place stays available to every client reading the field directly, and leaves two fields on one object disagreeing - how the shops screen once came out labelled Working, in green, above "2 of 3 cameras not connecting". - Computed in scanCamera, so every camera anybody reads passes through it. A state computed per handler is one a handler forgets, and this had already reached three screens. - CameraStaleAfter is 5 minutes: five missed reports, not one. Same reasoning as three missed heartbeats - an indicator that cries wolf gets ignored. - An unparseable last_seen_at is stale. It should be impossible, which is why it must not fall through to the state that says everything is fine. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
104 lines
3.2 KiB
Go
104 lines
3.2 KiB
Go
package main
|
|
|
|
import (
|
|
"bytes"
|
|
"context"
|
|
"net/http"
|
|
"os"
|
|
"testing"
|
|
"time"
|
|
|
|
"github.com/loyaly/behavision-desktop/internal/cloud"
|
|
)
|
|
|
|
// The whole chain against the real head office and a real shop computer:
|
|
//
|
|
// TEST_CLOUD_EMAIL=... TEST_CLOUD_PASSWORD=... \
|
|
// go test ./desktop/ -run RemoteLive -v
|
|
//
|
|
// Everything in stream_remote_test.go proves the relay against a fake that
|
|
// agrees with me. Only this proves the part that cannot be faked: that a shop
|
|
// computer behind a router with no inbound route actually pushes frames when
|
|
// asked, that they survive base64 and SSE, and that what comes out of the
|
|
// loopback relay is a multipart stream an <img> will paint.
|
|
//
|
|
// It also costs something to run, which is why it is opt-in: watching makes
|
|
// the shop computer upload for as long as the test reads.
|
|
func TestRemoteLiveFromProduction(t *testing.T) {
|
|
email, pass := os.Getenv("TEST_CLOUD_EMAIL"), os.Getenv("TEST_CLOUD_PASSWORD")
|
|
if email == "" || pass == "" {
|
|
t.Skip("set TEST_CLOUD_EMAIL and TEST_CLOUD_PASSWORD to run against production")
|
|
}
|
|
base := os.Getenv("TEST_CLOUD_URL")
|
|
if base == "" {
|
|
base = "https://mcp.loyaly.ai"
|
|
}
|
|
|
|
ctx, cancel := context.WithTimeout(context.Background(), 90*time.Second)
|
|
defer cancel()
|
|
c := cloud.New(base)
|
|
if _, err := c.Login(ctx, email, pass); err != nil {
|
|
t.Fatalf("login: %v", err)
|
|
}
|
|
|
|
cams, err := c.RemoteCameras(ctx)
|
|
if err != nil {
|
|
t.Fatalf("cameras: %v", err)
|
|
}
|
|
t.Logf("%d cameras", len(cams))
|
|
target := os.Getenv("TEST_CLOUD_CAMERA")
|
|
for _, cam := range cams {
|
|
conn := "waiting"
|
|
if cam.Connected != nil {
|
|
conn = map[bool]string{true: "connected", false: "not connecting"}[*cam.Connected]
|
|
}
|
|
t.Logf(" %-10s %-16s %-15s snapshot=%v", cam.CameraID, cam.Site, conn, cam.Snapshot.Available)
|
|
if target == "" && cam.Connected != nil && *cam.Connected {
|
|
target = cam.CameraID
|
|
}
|
|
}
|
|
if target == "" {
|
|
t.Skip("no connected camera to watch")
|
|
}
|
|
|
|
p := newStreamProxy()
|
|
if err := p.watchRemote(c.CameraLive); err != nil {
|
|
t.Fatalf("watchRemote: %v", err)
|
|
}
|
|
defer p.stop()
|
|
|
|
rctx, rcancel := context.WithTimeout(ctx, 30*time.Second)
|
|
defer rcancel()
|
|
req, _ := http.NewRequestWithContext(rctx, http.MethodGet, p.urlFor(target, "live.mjpeg"), nil)
|
|
resp, err := http.DefaultClient.Do(req)
|
|
if err != nil {
|
|
t.Fatalf("GET relay: %v", err)
|
|
}
|
|
defer resp.Body.Close()
|
|
|
|
start := time.Now()
|
|
acc, buf, frames := make([]byte, 0, 1<<20), make([]byte, 32*1024), 0
|
|
for frames < 10 {
|
|
n, rerr := resp.Body.Read(buf)
|
|
acc = append(acc, buf[:n]...)
|
|
frames = bytes.Count(acc, []byte("--"+mjpegBoundary))
|
|
if rerr != nil {
|
|
break
|
|
}
|
|
}
|
|
el := time.Since(start)
|
|
t.Logf("watching %q: %d frames, %d bytes, %.1fs (%.1f fps, %.0f KB/s)",
|
|
target, frames, len(acc), el.Seconds(),
|
|
float64(frames)/el.Seconds(), float64(len(acc))/el.Seconds()/1024)
|
|
|
|
if frames < 3 {
|
|
t.Fatalf("got %d frames from a connected camera - the shop computer is "+
|
|
"not answering head office's request to push", frames)
|
|
}
|
|
// Bytes that are actually a picture, not a framing header that happens to
|
|
// be well formed. A JPEG begins FFD8.
|
|
if !bytes.Contains(acc, []byte{0xFF, 0xD8, 0xFF}) {
|
|
t.Error("no JPEG start marker anywhere in the stream")
|
|
}
|
|
}
|