Brand assets in brand/ (the 512px mark, sizes for each surface, a
multi-size .ico). Windows executables carry it as a compiled-in
resource (rsrc_windows_amd64.syso from go-winres) so Explorer, the
taskbar and the installer show it; installer/build.ps1 therefore uses a
plain go build rather than wails build, which would add a second copy
and fail the link. The tray icon is the mark with a state dot over its
corner - a plain coloured circle read as a generic status light among
other icons - rendered from the embedded PNG at 32px so it survives
150% scaling. The desktop app's login, setup and sidebar marks, the
head-office web app's mark and favicon, and the engine dashboard's
favicon are the same file.
Also found while packaging: no wheel so far shipped static/, so the
engine's own dashboard at :8010 on a Windows source install would have
failed with a missing file. package-data now includes it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The server has always sent the broker's CA certificate in the enrolment
response, precisely so it never has to ship in an installer. Nothing on
the receiving end wrote it anywhere: the agent read the field under the
wrong name (ca_pem, the server says ca_cert) and the desktop app read it
correctly and dropped it. Every claimed PC therefore dialled
tls://mcp.loyaly.ai:8883 with the system trust store, the private CA
failed verification, and the agent reported 'the broker did not accept
this PC' - a TLS failure is indistinguishable from a refusal at that
layer. No real site could ever have published a visit.
Found by claiming this Mac as a real shop against production; fixed by
writing the CA to broker-ca.crt beside agent.json on both claim paths.
Verified: broker connected over TLS, camera pushed from head office,
engine streaming it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The installer runs the engine as <venv>\Scripts\python.exe, and since
Python 3.7.2 that file is a redirector that spawns the real interpreter
as a child. Stop() terminated the redirector and left the interpreter -
the process holding the cameras and the SQLite WAL - running with no
parent and nothing able to stop it. Seen on a Windows install: Quit from
the tray, and recognition still running.
The child is now started suspended, placed in a job object with
KILL_ON_JOB_CLOSE, and resumed. Cancel terminates the job, so the whole
tree goes; and the job dies with this process, so it goes even if the app
crashes. CREATE_NO_WINDOW while here: python.exe is a console program
and a GUI parent otherwise opens a black console on the shop counter.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Wanted: install it and the two office cameras are already there - but
without the release carrying their admin password where anyone with the
zip can read it. "Encode it" does not achieve that; anything the
installer can decode, anyone holding the installer can decode.
pkg/demo seals the camera list with AES-256-GCM under a key that is NOT
in the package: a 120-bit unlock code minted when the bundle is sealed,
given to whoever runs setup by voice or message, typed once. The code
is random, so it is key material directly through SHA-256; a human-
chosen passphrase would need a KDF and a dependency, 120 random bits do
not. The sealed file contains the format marker and noise. Tested: the
password and the host do not appear in it, a wrong code and a flipped
byte are both refused as ErrWrongCode, every seal differs.
behavision-demo-pack seals; it runs on the build machine and is never
shipped. The code is printed once and stored nowhere.
behavision-setup, on finding demo-cameras.enc beside the engine source,
asks for the code BEFORE the ten-minute download so a mistyped one costs
seconds, and adds the cameras at the end - through the running engine's
own Add Camera endpoint, not by writing its file. The store's save() is
what applies DPAPI to the password on Windows, so this is how the
credential ends up encrypted and machine-bound on the demo PC rather
than in cameras.json for anyone who can read ProgramData. It then marks
the PC standalone, so the app opens on Live instead of asking for an
installation code it will never get.
Which found the gap that DPAPI only works if pywin32 is importable, and
nothing had ever pulled it in - every Windows install to date would have
logged the warning and written camera passwords in the clear. Added as
a Windows-only dependency.
Verified in a clean container: a wrong code refused, the right one
unlocks two cameras, every install step passes, both cameras added
through the API, standalone set.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ran behavision-setup in a fresh Linux container: Python 3.12, nothing
else, the release contents mounted read-only the way Program Files or a
shared drive would be. It failed, and then it failed differently, and
both failures would have been the client's first experience.
1. `pip install <folder>` makes setuptools write behavision.egg-info
INTO the folder. The folder is read-only wherever a release is
sensibly unzipped, so: "could not create 'behavision.egg-info':
Read-only file system". The release now ships a wheel - pure Python,
buildable anywhere, nothing to build on the shop PC, and pip never
touches the unzipped folder. Source stays as a fallback and is copied
somewhere writable first.
2. The engine's paths.py knows two worlds - frozen (ProgramData) and a
checkout (the repo root) - and a pip-installed engine is neither. It
resolved its state root to site-packages: database there, camera
list there, and its generated API credential in a folder the app
never reads, while the app looked in ProgramData. Every call would be
401 on a stock install, with nothing in either log saying why. The
same disease as the Mac checkout two days ago, now in production
shape.
engine.ChildEnv is the one place the engine's environment is built,
used by the desktop app, the headless agent and the installer's own
smoke test. It passes BEHAVISION_DATA_DIR = this process's state
root, which paths.py honours ahead of every other rule, so the two
halves agree by construction however the engine was installed.
It also seeds config/default.yaml into the state root: a package in
site-packages has no config beside it to seed from.
Re-run on the same clean container: seven steps, all pass, models
downloaded, engine started and answered, and its data/ landed beside
agent.json - not in site-packages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The Go halves of this product cross-compile to Windows from any machine.
The engine does not: PyInstaller bundles the interpreter and the native
wheels of the machine it runs on, so a frozen engine can only be built on
Windows. That one fact was the entire reason no release had ever been
cut - two of the three binaries were ready for weeks.
behavision-setup installs the engine from source instead. It finds a
Python, builds a private virtual environment beside the database,
installs the engine into it, downloads the models, records how to start
it in the same agent.json the app reads, and then starts it and waits
for its API to answer.
That last step is the point. An installer that reports success and
leaves a shop with an engine that will not run has done worse than
failing: the failure surfaces later, to somebody who did not install it.
The trade, since whoever runs this is standing in a shop: it needs
Python and internet at install time and takes minutes, where a frozen
build needs neither. What it buys is a release that exists.
Details that are not incidental:
- `py -3` is tried before `python` on Windows. The launcher is what the
official installer puts on PATH; `python` there is often the Store
stub that prints an advert and exits 9009.
- a virtual environment, not the system Python. A shop PC may have
Python for something else, and the engine pins numpy below 2.0 -
installing that into a shared interpreter breaks the other thing
months later and silently.
- EngineExe is written absolute. The app resolves a relative one
against its install root under Program Files, where no interpreter
lives.
- pip's output is shown, not swallowed. When it fails on a proxy or a
missing build tool it says exactly what is wrong, and hiding that
leaves the operator with "setup failed" and nothing to act on.
- the console pauses before closing. Double-clicked from Explorer, a
program that finishes closes instantly and success and failure look
identical.
Verified as far as a Mac can: `pip install .` builds the wheel and
resolves every dependency, and `python -m behavision` then runs from
site-packages rather than the working directory - which is the mechanism
this depends on and had never been exercised, because the project has
only ever been run out of its own checkout.
NOT verified: any of it on Windows. Nothing here has run on the target
platform, and the `py -3` path and the ProgramData layout are exactly
where that will show.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pcn9asw19WGBfCEaHvNug6
A tenant had exactly the users somebody had created with a command on the
server. That is not a missing screen: a shop with an owner and four staff
either shared one password or raised a ticket per person, and a phone app
for the shop floor could not exist while there was one account to sign in
as.
Registration is by invitation, never open signup - the same line already
drawn around creating a company. The code carries the address and the role
and the request carries only a password, so a code that gets forwarded
cannot become somebody else's account, and a staff invitation cannot be
redeemed as an owner. Single use lives in the UPDATE and the account is
created in the same transaction.
Deactivating a member revokes their sessions in that transaction too. An
access token lives twelve hours, so without it "remove their access"
removed it sometime tomorrow. The session list and revoke that go with it
are the benefit of opaque tokens the product had been paying for and never
collecting: nothing could say what was signed in, let alone stop one.
Face images now work on a deployment with no object storage, which was
every local install and every self-hosted site - the arrivals feed said
"not storing customer photos" for every customer forever, on the screen
whose whole job is to show a face. Bounded to one row per visitor, so it
grows with the customer base and not with footfall; the bucket stays
primary wherever one exists.
Image.auth says whether a URL needs the session, because a browser img
cannot load one that does, a mobile image view can, and a webview can do
neither - the desktop client resolves those to a data URI in Go.
Found by running it, not by tests:
* UPDATE ... RETURNING gives the value AFTER the update, so the prune
read back empty keys, deleted nothing, and the table grew with
footfall exactly as if it were not there. The fake agreed with either
version; only the live Postgres test caught it.
* Trusting only the auth flag broke every shop card, because Sites.jsx
rebuilt a partial snapshot object and dropped it. A relative URL is
now sufficient on its own.
* ago() renders a future time as "just now", so a code valid for a week
read "expires just now".
Verified live against real Postgres: invite, preview, escalation refused,
register into a session, replay 404, staff forbidden, device revoked and
401 at once, last owner refused, and a 92,405-byte camera JPEG stored,
served to its owner, 401 with no session, 404 to another tenant, and
rendered in a browser.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
4 fps was not "live", and it was a number I picked rather than measured.
The engine actually produces ~12 distinct frames a second, so most of it
was being left on the floor.
Now: poll a little ahead of the engine and drop frames identical to the
last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps,
20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a
picture the viewer has not seen, and the rate follows the camera instead
of a constant.
Also records why this is MJPEG rather than passing the camera's own
compressed video through, which would be smoother, cheaper and use no
CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and
BOTH are H.265, despite stream paths ending in ".264". Browsers play
H.264 everywhere and H.265 only on some platforms, so passthrough cannot
rely on it, and transcoding HEVC on the shop PC would put a video encoder
on the machine already doing the recognition.
So probe_source now reports `codec`. It decides what is possible, an
installer can usually change it, and otherwise the only way to learn it is
to read RTSP by hand - which is how this was found.
The RTSP libraries used to establish that are NOT kept: they were only
ever imported by a spike test, and two large dependencies in a shipped
binary to answer a question OpenCV already knows is a bad trade. Their
`go get` had also silently bumped the agent to go 1.25 and broken the
desktop build, which is its own argument.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.
The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.
Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.
Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.
LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.
Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.
Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Both found by operating the stack rather than writing it: the local
processes were OOM-killed and bringing them back hit two gaps.
The headless agent had no way to be claimed at all. Bootstrap lived only
in desktop/internal/cloud, so the one configuration the agent binary
exists for - a back-office PC with no window - could only be onboarded by
hand-editing agent.json, which is the state the desktop's Setup screen was
built to end. `behavision-agent claim <code>` closes it; the CLI joins its
arguments because the code is printed in groups for reading aloud and an
operator pasting it will paste the spaces too.
Second: after the site's broker password was re-rolled, mosquitto logged
"not authorised" while the agent logged "timed out". Those need opposite
actions - re-link this PC, or go and look at the network - and paho's
SetConnectRetry collapses them, because it retries internally and the
connect token never completes. describeStall asks whether a TCP socket
opens at all, and says what is known rather than guessing at a reason the
broker never gives.
Verified end to end: minted a code from the platform as the owner,
claimed with the new command, broker connected, and the shop went to
online: true with 1/1 cameras on w600k_r50.
Also corrects this machine's memory in CLAUDE.md from 16 GB to 8 GB. It
feeds the model-fallback reasoning, and the local gallery already holds
17 embeddings tagged w600k_mbf beside 19 tagged w600k_r50 - the fallback
has silently fired before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Head office shows a camera's latest frame rather than live video, for a
reason that has not changed: the engine serves MJPEG on 127.0.0.1 on a PC
behind a shop's router with no inbound route, and relaying it needs
WebRTC/TURN. Pointing a browser straight at the shop PC is not the escape
either - the engine's API is Basic-authenticated with a credential it
generates locally and never sends anywhere, and shipping that to the
cloud so a web page could use it would put the key to the biometric API
and the live face feed in the server's database.
But that picture only worked if you had an S3 bucket. Without one,
attachSnapshots reported "This system is not storing images" for every
camera forever - on the two screens whose whole job is to show the
camera. Making them picture-led turned a missing feature into a wall of
empty tiles, on every local install and any self-hosted customer who does
not want a bucket.
migrations/009 adds camera_snapshots and the agent falls back to
PUT /api/agent/cameras/{camera}/snapshot when the presigned route answers
images_disabled - chosen by sentinel, never by matching the message, since
it picks between two routes. One row per camera is what makes this safe in
the database when face images are not: the key IS the camera, so storage
is (cameras x ~100 KB) and does not grow with footfall.
The read is session-authenticated rather than a signed link, which an
<img> cannot use - hence Shot.jsx and useAuthedImage, keyed on the URL
string rather than the snapshot object so a poll does not re-fetch 90 KB
per camera every few seconds, and revoking the object URL on cleanup.
Verified against the real office camera with no bucket configured: 90,587
bytes stored in Postgres, served as image/jpeg to a signed-in user, 401
without a session, rendered on both the Cameras and Shops cards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
`spool/` is unanchored, so it matched `agent/pkg/spool/` - the durable
queue, source code - and the repository excluded it. Cloning and building
was what found it; nothing in the working tree ever would, because the
files are right there.
`data/` and `agent.json` have the same shape and are anchored too.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn