main
4 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
| ff4f95c3b0 |
Four silent failures a demo on somebody else's Mac walked straight into
Three reported from a colleague's machine, plus one the fixing uncovered. Every one produced a message that was true and useless. ## behavision-setup chose the Python least likely to work findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a floor with NO ceiling, which is exactly backwards. The newest Python on a machine is the one least likely to have binary wheels. It picked 3.14, pip found no numpy wheel for cp314, fell back to building numpy from source and produced "Unknown compiler(s)"; once the operator had installed Xcode's command line tools to get past that, ten minutes of compiling ended in "<arm_neon.h> is intended only for ARM and AArch64 targets". maxMinor refuses in one line before anything is downloaded, and "too new" is a different message from "too old" - telling somebody holding Python 3.14 that no Python was found sends them to install a newer one, which is the direction that just failed. ## numpy<2.0 was the cap; OpenCV was the hazard Widening it needed proof, and the proof found something else. Nine runs of the detector guard per combination, one machine, one sitting: numpy 1.26 / cv2 4.11 9 passed, 0 crashed numpy 2.0 / cv2 4.11 8 passed, 1 crashed numpy 1.26 / cv2 4.14 3 passed, 6 crashed numpy 2.0 / cv2 4.14 2 passed, 7 crashed numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash was test_a_shared_detector_really_does_race, which races a shared cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11 usually turned it into an exception, 4.14 usually turns it into a segfault, and 4.11 crashing once says the hazard was always there. It never reached the product - Engine._build_worker builds a detector per camera. It reached the suite: two runs in three died with no failing assertion in them. The race runs in a subprocess now, and one clean attempt proves nothing, so the premise holds if any of several attempts misbehaves. 226 passed / 2 skipped on numpy 2.0.2, five runs of five. opencv stays capped below 5: everything above was measured on 4.x, and an uncapped >=4.8.1 gives every NEW install a major release this project has never run a real camera through. ## One MQTT client id for a whole shop, so two PCs fought over it behavision-<client>-<site> is the same string on every computer claimed to one site. MQTT requires unique client ids and a broker enforces it by disconnecting the older session, so the colleague's Mac and the shop's own till took turns kicking each other off: broker connected / broker connection lost: EOF / broker connected / EOF ... The damage is not confined to the new machine. The till is the other half of that loop, so signing in on a laptop to look at the product stops a live shop delivering visits - and from each end it reads as an unstable network. MQTTClientID() appends a per-installation id, minted on first load and written back so an existing install gets one without anybody doing anything. The site stays in the name because that is what a broker log is read by. An unwritable config falls back to a per-run id rather than a shared one. ## "no such file or directory" for an engine nobody had installed Pressing Start went straight to the supervisor, which reported what exec reported: a 200-character path ending in "no such file or directory". Every word true, none of it saying "run the setup tool" - the startup path had that sentence, in a log file nobody on a shop counter opens. engineMissing() is the one function the startup path, the Start button and the status panel all consult. It also names App Translocation, which was in that path and is unguessable: macOS runs a downloaded unsigned app from a random read-only copy, so relative paths resolve inside it and an install there would not survive a restart. The product is unsigned, so that is the normal first-run state on every Mac, not an edge case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj |
|||
| 50d122e5f0 |
An unused constant was a Stop() that could hang forever
Ran staticcheck across all three Go modules for the first time. server (23k lines) and desktop came back clean. agent had seven findings, and one of them was not tidiness. `stopGrace = 10 * time.Second` was declared and wired to nothing. Stop() cancels the context, cmd.Cancel kills the process tree, and then Stop() blocks on cmd.Wait() - which, with no WaitDelay set, waits not just for the process but for every writer of its stdout pipe to close. One grandchild still holding that pipe hangs Wait, hangs Stop, and on the desktop app that is the tray's Quit never returning. The constant named the intent and nothing read it. cmd.WaitDelay = stopGrace is the line that was missing. The rest were real but small: an unused field in the live relay, an unused sleep helper in the pump, and "net/url" imported twice under two names - both genuinely used, in two functions doing the same job for the same reason, so they are unified rather than one deleted. My first pass deleted the wrong one on a bad grep and the build caught it immediately. Three findings are suppressed rather than fixed, with the reason stated: - Two "error strings should not end with punctuation". Both are multi-line messages a shop operator reads at a counter, not errors anything wraps. ST1005 exists because wrapped errors concatenate mid-sentence; stripping the full stops would run three sentences together to satisfy a rule that does not apply. - A deliberately nil context in a pump test - the point of the test is that an unconnected client does not panic. It already carried //nolint:staticcheck, which is golangci-lint's directive and staticcheck ignores, which is why it kept being reported. Also tidied agent/go.mod, which had paho and x/sys marked indirect while being imported directly. All three modules clean, all suites pass: 21 Go packages, 226 engine tests. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj |
|||
| ffae7e45d5 |
Live view runs at the camera's real rate, and reports why it is MJPEG
4 fps was not "live", and it was a number I picked rather than measured. The engine actually produces ~12 distinct frames a second, so most of it was being left on the floor. Now: poll a little ahead of the engine and drop frames identical to the last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps, 20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a picture the viewer has not seen, and the rate follows the camera instead of a constant. Also records why this is MJPEG rather than passing the camera's own compressed video through, which would be smoother, cheaper and use no CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and BOTH are H.265, despite stream paths ending in ".264". Browsers play H.264 everywhere and H.265 only on some platforms, so passthrough cannot rely on it, and transcoding HEVC on the shop PC would put a video encoder on the machine already doing the recognition. So probe_source now reports `codec`. It decides what is possible, an installer can usually change it, and otherwise the only way to learn it is to read RTSP by hand - which is how this was found. The RTSP libraries used to establish that are NOT kept: they were only ever imported by a spike test, and two large dependencies in a shipped binary to answer a question OpenCV already knows is a bad trade. Their `go get` had also silently bumped the agent to go 1.25 and broken the desktop build, which is its own argument. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn |
|||
| 18686cbceb |
Live view at head office, relayed through the agent's outbound connection
I got this wrong first time. "Head office cannot show live video cheaply" conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs WebRTC and a TURN server. The shop PC is behind a router with no inbound route, so head office cannot pull the engine's MJPEG. It can answer the agent's outbound requests, which is the shape of everything else here: the server holds a poll open, the agent asks "is anyone watching?", and pushes JPEGs up for exactly as long as somebody is. Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at 640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds - 4 fps, as configured. The UI says "about 4 frames a second" rather than letting anyone conclude the camera stutters. Nothing is uploaded when nobody is looking, which is the whole cost argument: Publish returns false once the last viewer goes, interest lapses on a timer each viewer refreshes as it reads (so a closed tab stops the upload within seconds), one push is capped at five minutes, and the UI streams one camera at a time. LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell pushes nothing because nothing may be lost; here a dropped frame is the correct outcome, so each viewer has a one-slot buffer that is overwritten - the only frame worth having is the newest, and a queue would show an ever-growing delay behind the shop instead of dropping back to live. Ownership is proved once, before anything streams: the relay is keyed on a camera id, a hub does not know whose camera it holds, and a camera id is not a secret. Verified: another tenant gets 404, no session gets 401, and an agent cannot push into another site's camera. Also fixes a bug I introduced with it - the Live button was gated on `connected`, which is head office's last report and up to two minutes stale, so it hid itself during every reconnect. "Is that camera really down?" is exactly when somebody wants to look, and a hidden control says "you cannot" where the honest answer is "here is why". Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn |