Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.
## behavision-setup chose the Python least likely to work
findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".
maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.
## numpy<2.0 was the cap; OpenCV was the hazard
Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.
It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.
opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.
## One MQTT client id for a whole shop, so two PCs fought over it
behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:
broker connected / broker connection lost: EOF / broker connected / EOF ...
The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.
MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.
## "no such file or directory" for an engine nobody had installed
Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.
engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Found by running behavision-setup against a clean state directory the
way a second machine will, which had never been done. It failed at the
first step:
Setup did not finish: no Python 3.10 or newer was found
Found, but too old: python3 3.9
on a machine that has 3.12. The search was `python3` then `python`, and
on macOS `/usr/bin/python3` is ALWAYS the Command Line Tools build -
3.9 on current macOS, below the 3.10 floor. Anything newer installs as
`python3.12`, under Homebrew, as a framework, or somewhere a GUI
application's minimal PATH never sees.
So it now tries versioned names newest-first, then the plain ones, then
the four directories macOS actually uses - and absolute candidates are
stat'd rather than passed to LookPath, which only searches PATH. It
found /Users/tenext/.local/opt/python3.12/bin/python3.12, which is
exactly the interpreter it had been ignoring.
With that, the whole install completes on a Mac for the first time:
venv, engine and dependencies, models, agent.json, and "Engine starts
and answers - verified". EXIT=0, a 298 MB runtime.
Also the last thing it prints, which is the first thing an operator
acts on. It said "Start Behavision from the Start menu" and "it appears
in the system tray; right-click there to stop it". On macOS there is no
Start menu and, deliberately, no tray at all - so the finishing message
was describing a machine the user was not sitting at, on the one step
where setup had otherwise succeeded. It now says to right-click the app
the first time because the build is not notarised, and that closing the
window stops recognition.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ran staticcheck across all three Go modules for the first time. server
(23k lines) and desktop came back clean. agent had seven findings, and
one of them was not tidiness.
`stopGrace = 10 * time.Second` was declared and wired to nothing.
Stop() cancels the context, cmd.Cancel kills the process tree, and then
Stop() blocks on cmd.Wait() - which, with no WaitDelay set, waits not
just for the process but for every writer of its stdout pipe to close.
One grandchild still holding that pipe hangs Wait, hangs Stop, and on the
desktop app that is the tray's Quit never returning. The constant named
the intent and nothing read it. cmd.WaitDelay = stopGrace is the line
that was missing.
The rest were real but small: an unused field in the live relay, an
unused sleep helper in the pump, and "net/url" imported twice under two
names - both genuinely used, in two functions doing the same job for the
same reason, so they are unified rather than one deleted. My first pass
deleted the wrong one on a bad grep and the build caught it immediately.
Three findings are suppressed rather than fixed, with the reason stated:
- Two "error strings should not end with punctuation". Both are
multi-line messages a shop operator reads at a counter, not errors
anything wraps. ST1005 exists because wrapped errors concatenate
mid-sentence; stripping the full stops would run three sentences
together to satisfy a rule that does not apply.
- A deliberately nil context in a pump test - the point of the test is
that an unconnected client does not panic. It already carried
//nolint:staticcheck, which is golangci-lint's directive and
staticcheck ignores, which is why it kept being reported.
Also tidied agent/go.mod, which had paho and x/sys marked indirect while
being imported directly.
All three modules clean, all suites pass: 21 Go packages, 226 engine
tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Cheap for a reason worth stating: the Windows package is already a SOURCE
install - a pure-Python wheel plus a setup tool that builds a venv on the
target machine, because PyInstaller cannot cross-compile. macOS needs
nothing different. Same wheel, same setup tool, a natively built .app in
place of the .exe. No frozen engine, no 200 MB, no second packaging story.
behavision-setup already handled both platforms (Scripts vs bin, the
tasklist check guarded) and cross-compiled for darwin without a change.
The one thing that did not was its advice when Python is missing: it told
everyone to tick 'Add python.exe to PATH' on a Windows installer page.
Software that does not know which machine it is running on is software
somebody stops trusting for the rest of the session.
MAC=1 opts in, so the ordinary Windows release is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Nothing ever set the child's working directory, so it took the parent's -
and an app started by double-clicking its bundle is handed "/", not
anywhere useful. On macOS the symptom was
`python: No module named behavision` repeating forever, because the dev
engine is invoked as `-m behavision` and that resolves against the
working directory.
The same app launched from a terminal inside the repo worked perfectly,
which is exactly the shape of a bug that survives every test a developer
runs. It only appeared when the app was started the way a user starts
one.
Config.EngineDir, empty meaning the install root, set by both launchers -
the desktop app and the headless agent, which had identical code and the
identical omission. It matters beyond this case: the shipped Windows
engine is a one-folder PyInstaller build whose relative paths should
resolve beside itself rather than beside Explorer's idea of a current
directory.
Verified by double-clicking the bundle with nothing in the environment:
engine up on 8010 (401, gated), w600k_r50 on CoreML, gallery 5/5
embeddings usable and none stranded, both office cameras connected and
streaming, and head office reporting cameras 2/2 one heartbeat later.
Two things that showed up while proving it, both the product being
honest rather than faults:
- The camera at .121 was genuinely unreachable for several minutes, and
last_error said so in words an installer can act on - "cannot reach
192.168.1.121:554 - No route" - rather than `connected: false`. That
field was added yesterday for precisely this.
- Head office briefly showed cameras 0/1 against a local 2/2. That is a
60-second heartbeat, not a disagreement; the next one read 2/2.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The agent read the engine's generated credential file once, at startup.
On a brand new install that file does not exist yet: the agent starts the
engine, and the engine writes its credential seconds later. So the agent
held an empty credential for the life of the process and every call it
makes - health, stats, camera sync, the embedding for a visit - came back
401, with a tray showing a red engine that was running perfectly.
Measured on a fresh state directory today: three 401s, no camera ever
reconciled, and the engine left running the YAML-seeded main stream
instead of the sub-stream head office holds. The install script hid this
on Windows because setup runs the engine once before the app starts.
config.Creds resolves lazily and re-reads on a rejection; the camera
client, the supervisor and the desktop app's engine client all retry once
when it changes. A configured BEHAVISION_API_USER is never re-read - an
operator who set one means it. Tests pin the actual first-run ordering.
Also adds demo/, a one-screen live console for showing the whole chain:
camera, the six steps with a measured camera-to-cloud latency, the
customer editable in place, and the raw JSON a phone and a dashboard
receive from production side by side.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The first launch was a code box with a link under it, then an empty
Live screen with 'No cameras' in amber in a far corner, then a form
asking for an IP address, and for the first few minutes of all of it
the engine silently downloading 275 MB with nothing on screen but a
stopped-looking status. Walked in a browser with the new mock; nobody
who was not an installer would have got through it.
Now: a welcome that asks the one question a shop owner can answer -
managed from a head office, or on this PC only - with each path in a
sentence; a Getting Started checklist on Live that reads its three steps
from the engine and ticks them itself (recognition ready, camera added
and connected, camera proven by a walk-past), with the one button for
the next step, and that disappears the moment somebody is recognised;
and the model download reported as a percentage in the tray, the
sidebar and the checklist, parsed by the supervisor from the engine's
own progress lines.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Seen on the demo PC: Start did nothing and Stop stayed grey. The
supervisor's engine had failed because a second engine already held
port 8010, and the tray reported that as nothing at all. The supervisor
now keeps the engine's last lines and turns the known ones into a
sentence - 'port 8010 is already in use - another Behavision or its
engine is still running', 'run behavision-setup again' - which the tray
and the window show. Tray clicks no longer run on the menu loop, so a
stop that waits for the process cannot make the menu look dead.
Two ways that second process came to exist are closed: setup refuses to
run while Behavision.exe or the agent is up, and the app watches
agent.json so a claim made underneath it - which rotates the API token
- is picked up instead of leaving camera sync refused until a restart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Seen on the first claimed demo install: 'session expired' on every
screen, signed in as a user from the previous demo's head office, and
'Watching 3 cameras' for a shop with one - the PC had offered its two
leftover cameras up to head office, without their passwords, so the
same lens was listed twice and one copy could never be pushed anywhere.
Claiming now clears any stored session (a new head office is a new
world), a session whose refresh fails is forgotten on disk as well as
in memory so the app returns to Login by itself, and the demo setup
removes cameras left from an earlier install before it joins the shop,
because head office is the source of truth from then on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The first demo sealed the office cameras into the package and ran the
PC standalone - a copy of the product with no head office. The bundle
can now carry an installation code instead: setup redeems it exactly as
the app's Setup screen does, the PC joins the shop, and its cameras
arrive from head office on the first sync. The demo then IS the product
- login, Loya, head office - not a local imitation of it. The code is
single-use, so one bundle is one install. release.sh ships the bundle
with DEMO_PACK=.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Brand assets in brand/ (the 512px mark, sizes for each surface, a
multi-size .ico). Windows executables carry it as a compiled-in
resource (rsrc_windows_amd64.syso from go-winres) so Explorer, the
taskbar and the installer show it; installer/build.ps1 therefore uses a
plain go build rather than wails build, which would add a second copy
and fail the link. The tray icon is the mark with a state dot over its
corner - a plain coloured circle read as a generic status light among
other icons - rendered from the embedded PNG at 32px so it survives
150% scaling. The desktop app's login, setup and sidebar marks, the
head-office web app's mark and favicon, and the engine dashboard's
favicon are the same file.
Also found while packaging: no wheel so far shipped static/, so the
engine's own dashboard at :8010 on a Windows source install would have
failed with a missing file. package-data now includes it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The server has always sent the broker's CA certificate in the enrolment
response, precisely so it never has to ship in an installer. Nothing on
the receiving end wrote it anywhere: the agent read the field under the
wrong name (ca_pem, the server says ca_cert) and the desktop app read it
correctly and dropped it. Every claimed PC therefore dialled
tls://mcp.loyaly.ai:8883 with the system trust store, the private CA
failed verification, and the agent reported 'the broker did not accept
this PC' - a TLS failure is indistinguishable from a refusal at that
layer. No real site could ever have published a visit.
Found by claiming this Mac as a real shop against production; fixed by
writing the CA to broker-ca.crt beside agent.json on both claim paths.
Verified: broker connected over TLS, camera pushed from head office,
engine streaming it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The installer runs the engine as <venv>\Scripts\python.exe, and since
Python 3.7.2 that file is a redirector that spawns the real interpreter
as a child. Stop() terminated the redirector and left the interpreter -
the process holding the cameras and the SQLite WAL - running with no
parent and nothing able to stop it. Seen on a Windows install: Quit from
the tray, and recognition still running.
The child is now started suspended, placed in a job object with
KILL_ON_JOB_CLOSE, and resumed. Cancel terminates the job, so the whole
tree goes; and the job dies with this process, so it goes even if the app
crashes. CREATE_NO_WINDOW while here: python.exe is a console program
and a GUI parent otherwise opens a black console on the shop counter.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Wanted: install it and the two office cameras are already there - but
without the release carrying their admin password where anyone with the
zip can read it. "Encode it" does not achieve that; anything the
installer can decode, anyone holding the installer can decode.
pkg/demo seals the camera list with AES-256-GCM under a key that is NOT
in the package: a 120-bit unlock code minted when the bundle is sealed,
given to whoever runs setup by voice or message, typed once. The code
is random, so it is key material directly through SHA-256; a human-
chosen passphrase would need a KDF and a dependency, 120 random bits do
not. The sealed file contains the format marker and noise. Tested: the
password and the host do not appear in it, a wrong code and a flipped
byte are both refused as ErrWrongCode, every seal differs.
behavision-demo-pack seals; it runs on the build machine and is never
shipped. The code is printed once and stored nowhere.
behavision-setup, on finding demo-cameras.enc beside the engine source,
asks for the code BEFORE the ten-minute download so a mistyped one costs
seconds, and adds the cameras at the end - through the running engine's
own Add Camera endpoint, not by writing its file. The store's save() is
what applies DPAPI to the password on Windows, so this is how the
credential ends up encrypted and machine-bound on the demo PC rather
than in cameras.json for anyone who can read ProgramData. It then marks
the PC standalone, so the app opens on Live instead of asking for an
installation code it will never get.
Which found the gap that DPAPI only works if pywin32 is importable, and
nothing had ever pulled it in - every Windows install to date would have
logged the warning and written camera passwords in the clear. Added as
a Windows-only dependency.
Verified in a clean container: a wrong code refused, the right one
unlocks two cameras, every install step passes, both cameras added
through the API, standalone set.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ran behavision-setup in a fresh Linux container: Python 3.12, nothing
else, the release contents mounted read-only the way Program Files or a
shared drive would be. It failed, and then it failed differently, and
both failures would have been the client's first experience.
1. `pip install <folder>` makes setuptools write behavision.egg-info
INTO the folder. The folder is read-only wherever a release is
sensibly unzipped, so: "could not create 'behavision.egg-info':
Read-only file system". The release now ships a wheel - pure Python,
buildable anywhere, nothing to build on the shop PC, and pip never
touches the unzipped folder. Source stays as a fallback and is copied
somewhere writable first.
2. The engine's paths.py knows two worlds - frozen (ProgramData) and a
checkout (the repo root) - and a pip-installed engine is neither. It
resolved its state root to site-packages: database there, camera
list there, and its generated API credential in a folder the app
never reads, while the app looked in ProgramData. Every call would be
401 on a stock install, with nothing in either log saying why. The
same disease as the Mac checkout two days ago, now in production
shape.
engine.ChildEnv is the one place the engine's environment is built,
used by the desktop app, the headless agent and the installer's own
smoke test. It passes BEHAVISION_DATA_DIR = this process's state
root, which paths.py honours ahead of every other rule, so the two
halves agree by construction however the engine was installed.
It also seeds config/default.yaml into the state root: a package in
site-packages has no config beside it to seed from.
Re-run on the same clean container: seven steps, all pass, models
downloaded, engine started and answered, and its data/ landed beside
agent.json - not in site-packages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The Go halves of this product cross-compile to Windows from any machine.
The engine does not: PyInstaller bundles the interpreter and the native
wheels of the machine it runs on, so a frozen engine can only be built on
Windows. That one fact was the entire reason no release had ever been
cut - two of the three binaries were ready for weeks.
behavision-setup installs the engine from source instead. It finds a
Python, builds a private virtual environment beside the database,
installs the engine into it, downloads the models, records how to start
it in the same agent.json the app reads, and then starts it and waits
for its API to answer.
That last step is the point. An installer that reports success and
leaves a shop with an engine that will not run has done worse than
failing: the failure surfaces later, to somebody who did not install it.
The trade, since whoever runs this is standing in a shop: it needs
Python and internet at install time and takes minutes, where a frozen
build needs neither. What it buys is a release that exists.
Details that are not incidental:
- `py -3` is tried before `python` on Windows. The launcher is what the
official installer puts on PATH; `python` there is often the Store
stub that prints an advert and exits 9009.
- a virtual environment, not the system Python. A shop PC may have
Python for something else, and the engine pins numpy below 2.0 -
installing that into a shared interpreter breaks the other thing
months later and silently.
- EngineExe is written absolute. The app resolves a relative one
against its install root under Program Files, where no interpreter
lives.
- pip's output is shown, not swallowed. When it fails on a proxy or a
missing build tool it says exactly what is wrong, and hiding that
leaves the operator with "setup failed" and nothing to act on.
- the console pauses before closing. Double-clicked from Explorer, a
program that finishes closes instantly and success and failure look
identical.
Verified as far as a Mac can: `pip install .` builds the wheel and
resolves every dependency, and `python -m behavision` then runs from
site-packages rather than the working directory - which is the mechanism
this depends on and had never been exercised, because the project has
only ever been run out of its own checkout.
NOT verified: any of it on Windows. Nothing here has run on the target
platform, and the `py -3` path and the ProgramData layout are exactly
where that will show.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pcn9asw19WGBfCEaHvNug6
A tenant had exactly the users somebody had created with a command on the
server. That is not a missing screen: a shop with an owner and four staff
either shared one password or raised a ticket per person, and a phone app
for the shop floor could not exist while there was one account to sign in
as.
Registration is by invitation, never open signup - the same line already
drawn around creating a company. The code carries the address and the role
and the request carries only a password, so a code that gets forwarded
cannot become somebody else's account, and a staff invitation cannot be
redeemed as an owner. Single use lives in the UPDATE and the account is
created in the same transaction.
Deactivating a member revokes their sessions in that transaction too. An
access token lives twelve hours, so without it "remove their access"
removed it sometime tomorrow. The session list and revoke that go with it
are the benefit of opaque tokens the product had been paying for and never
collecting: nothing could say what was signed in, let alone stop one.
Face images now work on a deployment with no object storage, which was
every local install and every self-hosted site - the arrivals feed said
"not storing customer photos" for every customer forever, on the screen
whose whole job is to show a face. Bounded to one row per visitor, so it
grows with the customer base and not with footfall; the bucket stays
primary wherever one exists.
Image.auth says whether a URL needs the session, because a browser img
cannot load one that does, a mobile image view can, and a webview can do
neither - the desktop client resolves those to a data URI in Go.
Found by running it, not by tests:
* UPDATE ... RETURNING gives the value AFTER the update, so the prune
read back empty keys, deleted nothing, and the table grew with
footfall exactly as if it were not there. The fake agreed with either
version; only the live Postgres test caught it.
* Trusting only the auth flag broke every shop card, because Sites.jsx
rebuilt a partial snapshot object and dropped it. A relative URL is
now sufficient on its own.
* ago() renders a future time as "just now", so a code valid for a week
read "expires just now".
Verified live against real Postgres: invite, preview, escalation refused,
register into a session, replay 404, staff forbidden, device revoked and
401 at once, last owner refused, and a 92,405-byte camera JPEG stored,
served to its owner, 401 with no session, 404 to another tenant, and
rendered in a browser.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
4 fps was not "live", and it was a number I picked rather than measured.
The engine actually produces ~12 distinct frames a second, so most of it
was being left on the floor.
Now: poll a little ahead of the engine and drop frames identical to the
last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps,
20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a
picture the viewer has not seen, and the rate follows the camera instead
of a constant.
Also records why this is MJPEG rather than passing the camera's own
compressed video through, which would be smoother, cheaper and use no
CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and
BOTH are H.265, despite stream paths ending in ".264". Browsers play
H.264 everywhere and H.265 only on some platforms, so passthrough cannot
rely on it, and transcoding HEVC on the shop PC would put a video encoder
on the machine already doing the recognition.
So probe_source now reports `codec`. It decides what is possible, an
installer can usually change it, and otherwise the only way to learn it is
to read RTSP by hand - which is how this was found.
The RTSP libraries used to establish that are NOT kept: they were only
ever imported by a spike test, and two large dependencies in a shipped
binary to answer a question OpenCV already knows is a bad trade. Their
`go get` had also silently bumped the agent to go 1.25 and broken the
desktop build, which is its own argument.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.
The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.
Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.
Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.
LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.
Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.
Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Both found by operating the stack rather than writing it: the local
processes were OOM-killed and bringing them back hit two gaps.
The headless agent had no way to be claimed at all. Bootstrap lived only
in desktop/internal/cloud, so the one configuration the agent binary
exists for - a back-office PC with no window - could only be onboarded by
hand-editing agent.json, which is the state the desktop's Setup screen was
built to end. `behavision-agent claim <code>` closes it; the CLI joins its
arguments because the code is printed in groups for reading aloud and an
operator pasting it will paste the spaces too.
Second: after the site's broker password was re-rolled, mosquitto logged
"not authorised" while the agent logged "timed out". Those need opposite
actions - re-link this PC, or go and look at the network - and paho's
SetConnectRetry collapses them, because it retries internally and the
connect token never completes. describeStall asks whether a TCP socket
opens at all, and says what is known rather than guessing at a reason the
broker never gives.
Verified end to end: minted a code from the platform as the owner,
claimed with the new command, broker connected, and the shop went to
online: true with 1/1 cameras on w600k_r50.
Also corrects this machine's memory in CLAUDE.md from 16 GB to 8 GB. It
feeds the model-fallback reasoning, and the local gallery already holds
17 embeddings tagged w600k_mbf beside 19 tagged w600k_r50 - the fallback
has silently fired before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Head office shows a camera's latest frame rather than live video, for a
reason that has not changed: the engine serves MJPEG on 127.0.0.1 on a PC
behind a shop's router with no inbound route, and relaying it needs
WebRTC/TURN. Pointing a browser straight at the shop PC is not the escape
either - the engine's API is Basic-authenticated with a credential it
generates locally and never sends anywhere, and shipping that to the
cloud so a web page could use it would put the key to the biometric API
and the live face feed in the server's database.
But that picture only worked if you had an S3 bucket. Without one,
attachSnapshots reported "This system is not storing images" for every
camera forever - on the two screens whose whole job is to show the
camera. Making them picture-led turned a missing feature into a wall of
empty tiles, on every local install and any self-hosted customer who does
not want a bucket.
migrations/009 adds camera_snapshots and the agent falls back to
PUT /api/agent/cameras/{camera}/snapshot when the presigned route answers
images_disabled - chosen by sentinel, never by matching the message, since
it picks between two routes. One row per camera is what makes this safe in
the database when face images are not: the key IS the camera, so storage
is (cameras x ~100 KB) and does not grow with footfall.
The read is session-authenticated rather than a signed link, which an
<img> cannot use - hence Shot.jsx and useAuthedImage, keyed on the URL
string rather than the snapshot object so a poll does not re-fetch 90 KB
per camera every few seconds, and revoking the object URL on cleanup.
Verified against the real office camera with no bucket configured: 90,587
bytes stored in Postgres, served as image/jpeg to a signed-in user, 401
without a session, rendered on both the Cameras and Shops cards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
`spool/` is unanchored, so it matched `agent/pkg/spool/` - the durable
queue, source code - and the repository excluded it. Cloning and building
was what found it; nothing in the working tree ever would, because the
files are right there.
`data/` and `agent.json` have the same shape and are anchored too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn