Signing in on a second Mac showed "engine not reachable at
http://127.0.0.1:8010" and 0 of 0 cameras, on an account whose shops were
running and recognising people the whole time. Nothing was broken: Live() and
Cameras() read only the engine on loopback, so the app answered as though the
person had never signed in - and camera sync goes through the engine, which is
why the count was zero rather than stale.
Having no engine is a normal state. A shop PC watches cameras; an owner's
laptop, a manager's machine and a second till being set up do not, and all
three are signed in to the same estate. Both methods now fall back to head
office when loopback fails and somebody is signed in. Loopback is still tried
first: a real shop PC must never be shown a minute-old summary when the engine
two milliseconds away has the live one.
Decisions worth keeping:
- Viewing is on the snapshot, not inferred per screen. Three surfaces read it,
and a screen that computed it separately is how the shops screen once came
out labelled Working, in green, above "2 of 3 cameras not connecting".
- fraction_below_gate takes the WORST shop, never an average. 0.10 against
0.73 averages to 0.42 and hides the only shop anyone needs to visit.
- A remote camera is flagged, and Edit, Remove and Check placement are
withheld. They talk to a camera on a LAN this computer cannot reach, and a
button that cannot work is worse than one that is absent.
- connected is three states. null is "no shop computer has reported yet" and
reads as waiting; false is "Not connecting". A bare false sends somebody to
check cabling on a camera nobody has tried to reach.
- Snapshots are fetched in Go as data: URIs and cached by snapshot_at. A
webview <img> resolves a relative src against wails:// and cannot send the
bearer - the problem VisitorImage already solved - and this screen polls
every 8 seconds at ~90 KB a camera.
- With no engine AND no session, the engine error is still the answer. The
person is most likely setting this PC up.
The picture is the last snapshot and the banner says so: there is no live
video from here, because the engine's MJPEG stream is on the shop PC's
loopback behind a router with no inbound route. The LiveHub relay head office
uses is the answer to that and is a further step for this client.
Verified against production: five arrivals and two cameras parsed from the
real API. viewing_test.go covers the fallback, the worst-shop rule, the
withheld credentials and that an unchanged snapshot is fetched once across
two polls.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Nothing ever set the child's working directory, so it took the parent's -
and an app started by double-clicking its bundle is handed "/", not
anywhere useful. On macOS the symptom was
`python: No module named behavision` repeating forever, because the dev
engine is invoked as `-m behavision` and that resolves against the
working directory.
The same app launched from a terminal inside the repo worked perfectly,
which is exactly the shape of a bug that survives every test a developer
runs. It only appeared when the app was started the way a user starts
one.
Config.EngineDir, empty meaning the install root, set by both launchers -
the desktop app and the headless agent, which had identical code and the
identical omission. It matters beyond this case: the shipped Windows
engine is a one-folder PyInstaller build whose relative paths should
resolve beside itself rather than beside Explorer's idea of a current
directory.
Verified by double-clicking the bundle with nothing in the environment:
engine up on 8010 (401, gated), w600k_r50 on CoreML, gallery 5/5
embeddings usable and none stranded, both office cameras connected and
streaming, and head office reporting cameras 2/2 one heartbeat later.
Two things that showed up while proving it, both the product being
honest rather than faults:
- The camera at .121 was genuinely unreachable for several minutes, and
last_error said so in words an installer can act on - "cannot reach
192.168.1.121:554 - No route" - rather than `connected: false`. That
field was added yesterday for precisely this.
- Head office briefly showed cameras 0/1 against a local 2/2. That is a
60-second heartbeat, not a disagreement; the next one read 2/2.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The agent read the engine's generated credential file once, at startup.
On a brand new install that file does not exist yet: the agent starts the
engine, and the engine writes its credential seconds later. So the agent
held an empty credential for the life of the process and every call it
makes - health, stats, camera sync, the embedding for a visit - came back
401, with a tray showing a red engine that was running perfectly.
Measured on a fresh state directory today: three 401s, no camera ever
reconciled, and the engine left running the YAML-seeded main stream
instead of the sub-stream head office holds. The install script hid this
on Windows because setup runs the engine once before the app starts.
config.Creds resolves lazily and re-reads on a rejection; the camera
client, the supervisor and the desktop app's engine client all retry once
when it changes. A configured BEHAVISION_API_USER is never re-read - an
operator who set one means it. Tests pin the actual first-run ordering.
Also adds demo/, a one-screen live console for showing the whole chain:
camera, the six steps with a measured camera-to-cloud latency, the
customer editable in place, and the raw JSON a phone and a dashboard
receive from production side by side.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The first launch was a code box with a link under it, then an empty
Live screen with 'No cameras' in amber in a far corner, then a form
asking for an IP address, and for the first few minutes of all of it
the engine silently downloading 275 MB with nothing on screen but a
stopped-looking status. Walked in a browser with the new mock; nobody
who was not an installer would have got through it.
Now: a welcome that asks the one question a shop owner can answer -
managed from a head office, or on this PC only - with each path in a
sentence; a Getting Started checklist on Live that reads its three steps
from the engine and ticks them itself (recognition ready, camera added
and connected, camera proven by a walk-past), with the one button for
the next step, and that disappears the moment somebody is recognised;
and the model download reported as a percentage in the tray, the
sidebar and the checklist, parsed by the supervisor from the engine's
own progress lines.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The add-camera form asked for an IP address, and a shop owner does not
know their camera's IP address - it is on a sticker under the camera or
in a menu that differs by make. That field is where onboarding stopped
for anyone who was not an installer.
behavision/discover.py: one ONVIF WS-Discovery multicast (names the
camera and often its make) merged with a TCP sweep of port 554 across
the local /24 (misses nothing that streams). Stdlib only, ~4 s on the
office network, both cameras found. The add-camera sheet leads with
'Find cameras on this network'; picking a row fills the address and,
when the make is recognisable, the stream path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Seen on the demo PC: Start did nothing and Stop stayed grey. The
supervisor's engine had failed because a second engine already held
port 8010, and the tray reported that as nothing at all. The supervisor
now keeps the engine's last lines and turns the known ones into a
sentence - 'port 8010 is already in use - another Behavision or its
engine is still running', 'run behavision-setup again' - which the tray
and the window show. Tray clicks no longer run on the menu loop, so a
stop that waits for the process cannot make the menu look dead.
Two ways that second process came to exist are closed: setup refuses to
run while Behavision.exe or the agent is up, and the app watches
agent.json so a claim made underneath it - which rotates the API token
- is picked up instead of leaving camera sync refused until a restart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Seen on the first claimed demo install: 'session expired' on every
screen, signed in as a user from the previous demo's head office, and
'Watching 3 cameras' for a shop with one - the PC had offered its two
leftover cameras up to head office, without their passwords, so the
same lens was listed twice and one copy could never be pushed anywhere.
Claiming now clears any stored session (a new head office is a new
world), a session whose refresh fails is forgotten on disk as well as
in memory so the app returns to Login by itself, and the demo setup
removes cameras left from an earlier install before it joins the shop,
because head office is the source of truth from then on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ask Behavision: a panel beside any screen that talks to the head-office
assistant as the signed-in user - setup questions and 'is my shop
working' answered by the same thing, without leaving the app. The
assistant's prompt now knows how the product is set up (installation
codes, adding a camera, what a placement verdict means, the model
download on first run), so it is the help and not only the analyst. A
PC running on its own has nobody to ask and gets the essentials as text.
Cameras and Customers were still on the pre-redesign markup - the add
camera drawer ran off the right edge of the window because it used a
class the new stylesheet never sized. Both are rebuilt: cameras as
picture-led cards with connection and 'proven' as two separate claims
and a placement check laid out as the two steps it is; the customer
record as a proper sheet.
mock.js renders the app in a browser with fake bindings
(?mock=fresh|standalone|claimed, dev server only), so a screen can be
put in front of somebody without a Windows build. It is how these were
reviewed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The server has always sent the broker's CA certificate in the enrolment
response, precisely so it never has to ship in an installer. Nothing on
the receiving end wrote it anywhere: the agent read the field under the
wrong name (ca_pem, the server says ca_cert) and the desktop app read it
correctly and dropped it. Every claimed PC therefore dialled
tls://mcp.loyaly.ai:8883 with the system trust store, the private CA
failed verification, and the agent reported 'the broker did not accept
this PC' - a TLS failure is indistinguishable from a refusal at that
layer. No real site could ever have published a visit.
Found by claiming this Mac as a real shop against production; fixed by
writing the CA to broker-ca.crt beside agent.json on both claim paths.
Verified: broker connected over TLS, camera pushed from head office,
engine streaming it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The engine only ever started when somebody pressed Start. So a till
that rebooted overnight came back with the window open, the tray icon
showing, the session restored - and recognition off until a shop
assistant noticed. That is the failure the tray colours exist to catch,
and it should not be the default state every morning.
Guarded on the interpreter actually existing: on a PC where setup has
not run yet, the supervisor would loop on a missing executable with
nothing useful to say. Start and Stop remain for the case where somebody
has deliberately stopped it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ran behavision-setup in a fresh Linux container: Python 3.12, nothing
else, the release contents mounted read-only the way Program Files or a
shared drive would be. It failed, and then it failed differently, and
both failures would have been the client's first experience.
1. `pip install <folder>` makes setuptools write behavision.egg-info
INTO the folder. The folder is read-only wherever a release is
sensibly unzipped, so: "could not create 'behavision.egg-info':
Read-only file system". The release now ships a wheel - pure Python,
buildable anywhere, nothing to build on the shop PC, and pip never
touches the unzipped folder. Source stays as a fallback and is copied
somewhere writable first.
2. The engine's paths.py knows two worlds - frozen (ProgramData) and a
checkout (the repo root) - and a pip-installed engine is neither. It
resolved its state root to site-packages: database there, camera
list there, and its generated API credential in a folder the app
never reads, while the app looked in ProgramData. Every call would be
401 on a stock install, with nothing in either log saying why. The
same disease as the Mac checkout two days ago, now in production
shape.
engine.ChildEnv is the one place the engine's environment is built,
used by the desktop app, the headless agent and the installer's own
smoke test. It passes BEHAVISION_DATA_DIR = this process's state
root, which paths.py honours ahead of every other rule, so the two
halves agree by construction however the engine was installed.
It also seeds config/default.yaml into the state root: a package in
site-packages has no config beside it to seed from.
Re-run on the same clean container: seven steps, all pass, models
downloaded, engine started and answered, and its data/ landed beside
agent.json - not in site-packages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
StreamURL built http://user:pass@127.0.0.1:8010/api/cameras/<id>/
stream.mjpeg and handed it to an <img>, with a comment saying the
credentials were inline "so an <img> tag can load it".
It cannot. Chromium strips credentials from subresource URLs and has
since M59, and WebView2 is Chromium - so on the one platform this
product ships to, every camera tile on a shop counter was a broken
image. Measured against a running engine: the app's Go-side calls
returned stats and people while an <img> on that very URL failed, and
curl proved the URL answered 200. The engine was never the problem.
The password now stays on this side of the process boundary. A loopback
relay attaches Basic auth and streams the engine's bytes back
unchanged - the same reasoning Shot.jsx already follows at head office,
where an <img> equally cannot carry a session.
What the relay is careful about, since it is a door onto the biometric
API with a credential attached:
- loopback only, on a port the OS picks; a fixed one would collide
with whatever else a shop PC runs and read as "the cameras broke"
- a per-run random token in the path. The engine's own credential
exists so the live face feed is never served open; an
unauthenticated relay would hand that feed to any other process on
the PC. Compared in constant time, and a wrong one is 404, not 403
- an allow-list of stream.mjpeg and frame.jpg. Holding the token does
not reach the identity list, the gallery, or erasure
- camera ids validated, not interpolated
- every chunk flushed; a buffered MJPEG stream is a tile that never
paints, which looks identical to the bug being fixed
Two of those were written after a test failed, not before:
- `..` MATCHES the id pattern, because real camera ids contain dots.
`/api/cameras/../stream.mjpeg` is not the endpoint anyone intended.
The id can never hold a slash, so `.` and `..` are the whole
remaining traversal surface and are now refused by name.
- the serve goroutine read p.srv off the struct while stop() was
nilling it, so a quick start/stop dereferenced nil and took the
process down. Captured before launching now.
FrameURL is deliberately not added. No screen asks for a still, and a
bound method nothing calls is the same defect as a capability the UI
cannot reach, only pointing the other way.
Verified: nine unit tests, plus a live test against the real engine and
the real office camera - two MJPEG frames, 90,793 bytes, no credential
in the URL. Windows and darwin both build; vet clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pcn9asw19WGBfCEaHvNug6
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.
The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.
Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.
Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.
LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.
Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.
Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn