Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.
## behavision-setup chose the Python least likely to work
findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".
maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.
## numpy<2.0 was the cap; OpenCV was the hazard
Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.
It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.
opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.
## One MQTT client id for a whole shop, so two PCs fought over it
behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:
broker connected / broker connection lost: EOF / broker connected / EOF ...
The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.
MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.
## "no such file or directory" for an engine nobody had installed
Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.
engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Two changes, and the second was found by verifying the first.
## Watching a camera from the app, in another building
Snapshots answer "is that camera working". They do not answer "what is
happening in my shop right now", which is what somebody who opens the app away
from the counter is asking. Head office's browser already had that answer -
LiveHub plus cameras.Live, where the shop PC asks outbound whether anybody is
watching and pushes JPEG frames for as long as somebody is - and the app could
not reach it.
cloud.CameraLive opens that feed and the app's own loopback relay re-emits it
as multipart MJPEG. That is the trick: frames arrive base64 over SSE, an <img>
cannot render that, and an <img> renders MJPEG natively - so a tile is an
ordinary <img> pointed at loopback whether the camera is in this room or
another city.
- Reconnecting happens in the relay, not the page. The server caps one push at
five minutes, so doing it here means the <img> never sees the stream end.
- The headers are flushed before the first frame. Go writes them on the first
body write, so without that the whole response waits for the shop PC to
start pushing. Measured against production: 30 seconds and not even a
Content-Type, which surfaces as the request timing out.
- One camera at a time. Watching makes a shop PC upload, so a grid that went
live at once would put an estate's worth of cameras on the wire because
somebody opened a page.
- live.mjpeg is behind the same per-run token as the engine routes, and a
wrong token is a 404 that never reaches head office at all.
- CameraLive uses its own HTTP client: the shared one's 30s timeout covers the
whole response and would sever a working view every thirty seconds - the
trap that made the server set WriteTimeout to zero for its own SSE endpoint.
## A camera read "Connected" for 34 minutes after the shop PC went blind
Which is why the verification above looked like a failure: head office
registered the viewer and no frame ever came.
reportWith returns early when the engine is unreachable - correctly, it has
nothing to say - so the last state it sent stays in the database looking
current. Measured live: cam2 and entrance both reading Connected, in green,
with last_seen_at 34 minutes old, while the heartbeat from the same PC said
cameras_up 0 of 0. Two surfaces reading two stored fields and disagreeing.
false could not be the answer. It means "this camera is not connecting", which
sends an installer to check cabling on a camera that was working perfectly the
last time anybody could ask it. So there are four states and one function:
connected reported recently, and working
not_connecting reported recently, and the stream will not open
waiting no shop PC has ever reported this camera
stale reported once, and not lately
- Connected is CLEARED when stale or waiting. A stale true left in place stays
available to every client reading the field directly, and leaves two fields
on one object disagreeing - how the shops screen once came out labelled
Working, in green, above "2 of 3 cameras not connecting".
- Computed in scanCamera, so every camera anybody reads passes through it. A
state computed per handler is one a handler forgets, and this had already
reached three screens.
- CameraStaleAfter is 5 minutes: five missed reports, not one. Same reasoning
as three missed heartbeats - an indicator that cries wolf gets ignored.
- An unparseable last_seen_at is stale. It should be impossible, which is why
it must not fall through to the state that says everything is fine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Signing in on a second Mac showed "engine not reachable at
http://127.0.0.1:8010" and 0 of 0 cameras, on an account whose shops were
running and recognising people the whole time. Nothing was broken: Live() and
Cameras() read only the engine on loopback, so the app answered as though the
person had never signed in - and camera sync goes through the engine, which is
why the count was zero rather than stale.
Having no engine is a normal state. A shop PC watches cameras; an owner's
laptop, a manager's machine and a second till being set up do not, and all
three are signed in to the same estate. Both methods now fall back to head
office when loopback fails and somebody is signed in. Loopback is still tried
first: a real shop PC must never be shown a minute-old summary when the engine
two milliseconds away has the live one.
Decisions worth keeping:
- Viewing is on the snapshot, not inferred per screen. Three surfaces read it,
and a screen that computed it separately is how the shops screen once came
out labelled Working, in green, above "2 of 3 cameras not connecting".
- fraction_below_gate takes the WORST shop, never an average. 0.10 against
0.73 averages to 0.42 and hides the only shop anyone needs to visit.
- A remote camera is flagged, and Edit, Remove and Check placement are
withheld. They talk to a camera on a LAN this computer cannot reach, and a
button that cannot work is worse than one that is absent.
- connected is three states. null is "no shop computer has reported yet" and
reads as waiting; false is "Not connecting". A bare false sends somebody to
check cabling on a camera nobody has tried to reach.
- Snapshots are fetched in Go as data: URIs and cached by snapshot_at. A
webview <img> resolves a relative src against wails:// and cannot send the
bearer - the problem VisitorImage already solved - and this screen polls
every 8 seconds at ~90 KB a camera.
- With no engine AND no session, the engine error is still the answer. The
person is most likely setting this PC up.
The picture is the last snapshot and the banner says so: there is no live
video from here, because the engine's MJPEG stream is on the shop PC's
loopback behind a router with no inbound route. The LiveHub relay head office
uses is the answer to that and is a further step for this client.
Verified against production: five arrivals and two cameras parsed from the
real API. viewing_test.go covers the fallback, the worst-shop rule, the
withheld credentials and that an unchanged snapshot is fetched once across
two polls.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Reported from the Mac build: the window cannot be maximised. It is not a
Wails limitation or a WebView quirk, it is an omission with a very
specific consequence.
Wails computes zoomable INSIDE `if frontendOptions.Mac != nil`:
var fullSizeContent, hideTitleBar, zoomable, ... C.int // 0
if frontendOptions.Mac != nil {
zoomable = bool2Cint(!frontendOptions.Mac.DisableZoom)
}
and the native side then acts on the zero:
if (!zoomable && resizable) {
NSButton *button = [self.mainWindow
standardWindowButton:NSWindowZoomButton];
[button setEnabled: NO];
}
So leaving Mac unset does not mean "take the defaults" - it means the
green button is created and then explicitly disabled. There was a Windows
options block and no Mac one, which is how this survived: the platform
that was configured behaved, and the platform that was not looked broken.
Fixed by the block existing. The fields are written out rather than left
as an empty struct so it reads as a decision rather than something half
typed.
Verified at runtime rather than by reasoning about the source alone: all
three title-bar buttons report enabled=true through the accessibility
API, and the window resizes to 1440x900, the full display.
One correction to my own first check, recorded because it nearly sent me
the wrong way: querying AXFullScreenButton as an ATTRIBUTE of the window
returns "missing value" whether or not the button exists. It has to be
found by subrole among the window's buttons. The button was fine; the
question was wrong.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Nothing ever set the child's working directory, so it took the parent's -
and an app started by double-clicking its bundle is handed "/", not
anywhere useful. On macOS the symptom was
`python: No module named behavision` repeating forever, because the dev
engine is invoked as `-m behavision` and that resolves against the
working directory.
The same app launched from a terminal inside the repo worked perfectly,
which is exactly the shape of a bug that survives every test a developer
runs. It only appeared when the app was started the way a user starts
one.
Config.EngineDir, empty meaning the install root, set by both launchers -
the desktop app and the headless agent, which had identical code and the
identical omission. It matters beyond this case: the shipped Windows
engine is a one-folder PyInstaller build whose relative paths should
resolve beside itself rather than beside Explorer's idea of a current
directory.
Verified by double-clicking the bundle with nothing in the environment:
engine up on 8010 (401, gated), w600k_r50 on CoreML, gallery 5/5
embeddings usable and none stranded, both office cameras connected and
streaming, and head office reporting cameras 2/2 one heartbeat later.
Two things that showed up while proving it, both the product being
honest rather than faults:
- The camera at .121 was genuinely unreachable for several minutes, and
last_error said so in words an installer can act on - "cannot reach
192.168.1.121:554 - No route" - rather than `connected: false`. That
field was added yesterday for precisely this.
- Head office briefly showed cameras 0/1 against a local 2/2. That is a
60-second heartbeat, not a disagreement; the next one read 2/2.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Chosen deliberately as a DEVELOPER build, not a product. Indian retail
counters are Windows; shipping a Mac product means an Apple Developer
account, notarisation, a second installer format, a second frozen engine
and DPAPI having no macOS equivalent - a permanent second platform for
customers who do not have Macs. What a Mac build is worth is demoing the
desktop app on the machine it is written on, without needing the Windows
box.
It built after one missing framework (previous commit) and then crashed
within a second, twice, both times in the tray:
systray.Run SIGTRAP inside cgo. nativeLoop() takes the
macOS main run loop for itself and Wails
already has it. macOS has exactly one.
RunWithExternalLoop "NSWindow should only be instantiated on the
main thread!" - it registers in the existing
NSApplication rather than starting a second,
but still builds AppKit objects, and Wails'
OnStartup is not the main thread.
Making it work needs the status item created through a main-queue
dispatch inside Wails' lifecycle. That is real work for a build whose
purpose is a demo, so macOS has no tray and the file says so at length
rather than leaving the next person to rediscover both crashes.
The consequence is handled rather than left lying. With no tray there is
no way back from a hidden window and no way to quit, so hiding on close
would strand a running engine behind no window, no tray and no control -
force-quit or nothing. On macOS closing the window therefore quits, and
OnShutdown stops the engine. Same rule the tray's Quit already follows:
never leave it watching with no visible control. Windows is untouched,
where hiding is correct because the tray is how it comes back.
The runner is split by build tag rather than branched at runtime because
the two platforms need different systray ENTRY POINTS, not different
arguments.
Verified: 18 seconds up, zero crash markers, 88 MB resident, and an
honest "engine not installed yet" instead of a crash - against a
throwaway data dir so it claimed nothing and touched no camera. The
frozen Mac engine is deliberately not built; the app takes an engine
command from config, which is how the dev setup already points at the
venv.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Reported from the shipped Windows app. Three faults in one call, and the
first is why it failed rather than merely misbehaved.
runtime.Show is implemented by Wails as a bare mainWindow.Show(), while
runtime.WindowShow wraps the identical work in runtime.LockOSThread. Win32
window operations have to run on the thread owning the window's message
pump, and the tray's handler runs on the SYSTRAY's goroutine, which is
never that thread. An unlocked Win32 call from an arbitrary goroutine is
the bug.
Two more that would each have been enough on their own:
- Showing is not un-minimising. Hidden and minimised are different states
and Show only fixes the first, so a window the user minimised stayed
minimised.
- Windows refuses the foreground to a process that does not already hold
it, so the window came back BEHIND whatever was being looked at.
Clicking a tray icon is by definition a moment when this app is not in
front, so that is not an edge case here - it is every time. The
always-on-top flip is the ordinary way to ask, and it is why this now
runs in a goroutine rather than on the menu loop, which must not sleep.
The same four calls fix OnSecondInstanceLaunch, which had the same shape
and is reached far more often: double-clicking the desktop icon while the
app is already running.
OnBeforeClose used runtime.Hide against a reopen that used WindowShow -
different calls on Windows, one thread-locked and one not. Paired now.
And a Mac build, because the question came up and the answer turned out
to be yes. Wails' darwin frontend references UTType without linking
UniformTypeIdentifiers, so the build failed at the LINK step after
compiling everything - which reads like a broken toolchain rather than
one missing flag. There was no Mac version because of that, not because
of a design limit. darwin_link.go declares the framework in source rather
than leaving it as a CGO_LDFLAGS incantation, for the same reason
deploy.sh now finds Go itself. Verified: plain `go build` produces a
16 MB arm64 binary on this Mac, and the Windows build is unchanged.
Worth knowing for whoever edits that file: the comment directly above
`import "C"` is cgo's C preamble, not documentation. The first attempt put
the explanation there and the prose was compiled as C.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The agent read the engine's generated credential file once, at startup.
On a brand new install that file does not exist yet: the agent starts the
engine, and the engine writes its credential seconds later. So the agent
held an empty credential for the life of the process and every call it
makes - health, stats, camera sync, the embedding for a visit - came back
401, with a tray showing a red engine that was running perfectly.
Measured on a fresh state directory today: three 401s, no camera ever
reconciled, and the engine left running the YAML-seeded main stream
instead of the sub-stream head office holds. The install script hid this
on Windows because setup runs the engine once before the app starts.
config.Creds resolves lazily and re-reads on a rejection; the camera
client, the supervisor and the desktop app's engine client all retry once
when it changes. A configured BEHAVISION_API_USER is never re-read - an
operator who set one means it. Tests pin the actual first-run ordering.
Also adds demo/, a one-screen live console for showing the whole chain:
camera, the six steps with a measured camera-to-cloud latency, the
customer editable in place, and the raw JSON a phone and a dashboard
receive from production side by side.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The first launch was a code box with a link under it, then an empty
Live screen with 'No cameras' in amber in a far corner, then a form
asking for an IP address, and for the first few minutes of all of it
the engine silently downloading 275 MB with nothing on screen but a
stopped-looking status. Walked in a browser with the new mock; nobody
who was not an installer would have got through it.
Now: a welcome that asks the one question a shop owner can answer -
managed from a head office, or on this PC only - with each path in a
sentence; a Getting Started checklist on Live that reads its three steps
from the engine and ticks them itself (recognition ready, camera added
and connected, camera proven by a walk-past), with the one button for
the next step, and that disappears the moment somebody is recognised;
and the model download reported as a percentage in the tray, the
sidebar and the checklist, parsed by the supervisor from the engine's
own progress lines.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The add-camera form asked for an IP address, and a shop owner does not
know their camera's IP address - it is on a sticker under the camera or
in a menu that differs by make. That field is where onboarding stopped
for anyone who was not an installer.
behavision/discover.py: one ONVIF WS-Discovery multicast (names the
camera and often its make) merged with a TCP sweep of port 554 across
the local /24 (misses nothing that streams). Stdlib only, ~4 s on the
office network, both cameras found. The add-camera sheet leads with
'Find cameras on this network'; picking a row fills the address and,
when the make is recognisable, the stream path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Seen on the demo PC: Start did nothing and Stop stayed grey. The
supervisor's engine had failed because a second engine already held
port 8010, and the tray reported that as nothing at all. The supervisor
now keeps the engine's last lines and turns the known ones into a
sentence - 'port 8010 is already in use - another Behavision or its
engine is still running', 'run behavision-setup again' - which the tray
and the window show. Tray clicks no longer run on the menu loop, so a
stop that waits for the process cannot make the menu look dead.
Two ways that second process came to exist are closed: setup refuses to
run while Behavision.exe or the agent is up, and the app watches
agent.json so a claim made underneath it - which rotates the API token
- is picked up instead of leaving camera sync refused until a restart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Seen on the first claimed demo install: 'session expired' on every
screen, signed in as a user from the previous demo's head office, and
'Watching 3 cameras' for a shop with one - the PC had offered its two
leftover cameras up to head office, without their passwords, so the
same lens was listed twice and one copy could never be pushed anywhere.
Claiming now clears any stored session (a new head office is a new
world), a session whose refresh fails is forgotten on disk as well as
in memory so the app returns to Login by itself, and the demo setup
removes cameras left from an earlier install before it joins the shop,
because head office is the source of truth from then on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
A name, a voice and a door. The prompt now asks for a colleague on the
shop floor - answer first, one to three sentences, the shop's name and
the person's name, the one thing to do next - instead of a report with
headings. Both apps put her behind the Loyaly mark in the top-right
corner of every screen, because a buddy you have to find in a sidebar
is not around.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ask Behavision: a panel beside any screen that talks to the head-office
assistant as the signed-in user - setup questions and 'is my shop
working' answered by the same thing, without leaving the app. The
assistant's prompt now knows how the product is set up (installation
codes, adding a camera, what a placement verdict means, the model
download on first run), so it is the help and not only the analyst. A
PC running on its own has nobody to ask and gets the essentials as text.
Cameras and Customers were still on the pre-redesign markup - the add
camera drawer ran off the right edge of the window because it used a
class the new stylesheet never sized. Both are rebuilt: cameras as
picture-led cards with connection and 'proven' as two separate claims
and a placement check laid out as the two steps it is; the customer
record as a proper sheet.
mock.js renders the app in a browser with fake bindings
(?mock=fresh|standalone|claimed, dev server only), so a screen can be
put in front of somebody without a Windows build. It is how these were
reviewed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Brand assets in brand/ (the 512px mark, sizes for each surface, a
multi-size .ico). Windows executables carry it as a compiled-in
resource (rsrc_windows_amd64.syso from go-winres) so Explorer, the
taskbar and the installer show it; installer/build.ps1 therefore uses a
plain go build rather than wails build, which would add a second copy
and fail the link. The tray icon is the mark with a state dot over its
corner - a plain coloured circle read as a generic status light among
other icons - rendered from the embedded PNG at 32px so it survives
150% scaling. The desktop app's login, setup and sidebar marks, the
head-office web app's mark and favicon, and the engine dashboard's
favicon are the same file.
Also found while packaging: no wheel so far shipped static/, so the
engine's own dashboard at :8010 on a Windows source install would have
failed with a missing file. package-data now includes it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The server has always sent the broker's CA certificate in the enrolment
response, precisely so it never has to ship in an installer. Nothing on
the receiving end wrote it anywhere: the agent read the field under the
wrong name (ca_pem, the server says ca_cert) and the desktop app read it
correctly and dropped it. Every claimed PC therefore dialled
tls://mcp.loyaly.ai:8883 with the system trust store, the private CA
failed verification, and the agent reported 'the broker did not accept
this PC' - a TLS failure is indistinguishable from a refusal at that
layer. No real site could ever have published a visit.
Found by claiming this Mac as a real shop against production; fixed by
writing the CA to broker-ca.crt beside agent.json on both claim paths.
Verified: broker connected over TLS, camera pushed from head office,
engine streaming it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The window hides to the tray on close, so the natural next step for a
shop assistant is to double-click the shortcut again. That started a
second full copy of the app: a second tray icon, a second engine
supervisor on the same SQLite WAL and the same port - the start-twice
failure the agent package was built to prevent, on the one binary that
never had the guard. Seen on a Windows install as a row of tray icons.
Wails' SingleInstanceLock now hands the second launch to the first
process, which brings its window to the front.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The Live screen led with a camera tile beside the arrivals. Nobody at a
counter is watching CCTV; they are looking up at a customer and need the
name. The tile also cost CPU the recognition pipeline needs and pulled a
stream relay into the app for a picture that was decoration. Arrivals now
take the whole screen. The camera picture stays on the Cameras screen,
where it is a setup tool and not a feed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The window a shop assistant stares at all day was the weakest surface in
this system, and it looked improvised because it was: navigation drawn
with text characters (◉ ☺ ▢) that sit on the text baseline and cannot
take a stroke weight, margins set inline per screen, and four large stat
boxes dominating the page while the product's entire reason for existing
- WHO JUST WALKED IN - was a list of "person.seen" rows in the corner.
Rebuilt around the person in front of it: a counter, a cheap monitor,
somebody mid-conversation with a customer.
- ui/icons.jsx: one drawn icon set, 24-unit grid, 1.6 stroke,
currentColor, so one icon works on every surface and in every state.
- styles.css: a real system. Four-step ground→raised palette biased
blue-green (this product lives in the world of lenses), one spacing
scale, one type scale, tabular figures wherever digits are compared
or refreshed in place, and the scrollbars restyled - the default
light scrollbar on a dark panel is the loudest "web page in a frame"
tell there is.
- Live: a status strip that answers "is this working" in one line,
cameras as pictures with the caption over the image, and arrivals as
cards big enough to match against the person standing there. The
four stat boxes became a slim strip at the foot, where numbers that
nobody acts on belong.
- State is carried by shape AND colour everywhere - a pill, a dot and
an edge stripe - because this gets read from two metres away and
some operators do not see red and green apart.
- Motion only where it means something: a live camera pulses, a fresh
arrival slides in once. Nothing loops for decoration; this process
shares a CPU with recognition.
Two things fixed because the screen showed them, not because a test did:
- The sidebar read "Stopped" beside a live camera feed and a counter
ticking up, whenever the engine was running but not started BY the
app. That is the two-surfaces-disagreeing bug the tray exists to
avoid. It now reads "Running outside the app" in amber, and Start is
disabled rather than offering to launch a second engine onto one
SQLite WAL.
- The arrivals panel shrank to fit its content and left a hole beside
a tall camera tile - so the layout looked broken exactly when the
shop was quiet, which is most of the time. Both panels stretch and
scroll their own content now.
Every existing class name still resolves, so the screens not rewritten
here pick the system up unchanged. Windows and darwin build; tests pass.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The first Windows install reached the setup screen, typed an
installation code, and was told the session had expired. There was no
session. The code had been minted on a different head office, and the
server said so - 401 bad_token, "That installation code is not valid.
Ask for a new one." - and the client threw the message away, because it
mapped every 401 to the string "session expired".
A 401 on a call that carried a session is a session problem. A 401 on a
call that carried none is about the request, and the server's message is
the answer. The client now tells them apart by whether it sent a token.
Two tests, one for each side of the rule.
Also: a launcher for pointing a Windows PC at a head office on the LAN,
with the two settings that needs and a comment saying why neither is
acceptable outside a demo.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The engine only ever started when somebody pressed Start. So a till
that rebooted overnight came back with the window open, the tray icon
showing, the session restored - and recognition off until a shop
assistant noticed. That is the failure the tray colours exist to catch,
and it should not be the default state every morning.
Guarded on the interpreter actually existing: on a PC where setup has
not run yet, the supervisor would loop on a missing executable with
nothing useful to say. Start and Stop remain for the case where somebody
has deliberately stopped it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ran behavision-setup in a fresh Linux container: Python 3.12, nothing
else, the release contents mounted read-only the way Program Files or a
shared drive would be. It failed, and then it failed differently, and
both failures would have been the client's first experience.
1. `pip install <folder>` makes setuptools write behavision.egg-info
INTO the folder. The folder is read-only wherever a release is
sensibly unzipped, so: "could not create 'behavision.egg-info':
Read-only file system". The release now ships a wheel - pure Python,
buildable anywhere, nothing to build on the shop PC, and pip never
touches the unzipped folder. Source stays as a fallback and is copied
somewhere writable first.
2. The engine's paths.py knows two worlds - frozen (ProgramData) and a
checkout (the repo root) - and a pip-installed engine is neither. It
resolved its state root to site-packages: database there, camera
list there, and its generated API credential in a folder the app
never reads, while the app looked in ProgramData. Every call would be
401 on a stock install, with nothing in either log saying why. The
same disease as the Mac checkout two days ago, now in production
shape.
engine.ChildEnv is the one place the engine's environment is built,
used by the desktop app, the headless agent and the installer's own
smoke test. It passes BEHAVISION_DATA_DIR = this process's state
root, which paths.py honours ahead of every other rule, so the two
halves agree by construction however the engine was installed.
It also seeds config/default.yaml into the state root: a package in
site-packages has no config beside it to seed from.
Re-run on the same clean container: seven steps, all pass, models
downloaded, engine started and answered, and its data/ landed beside
agent.json - not in site-packages.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
StreamURL built http://user:pass@127.0.0.1:8010/api/cameras/<id>/
stream.mjpeg and handed it to an <img>, with a comment saying the
credentials were inline "so an <img> tag can load it".
It cannot. Chromium strips credentials from subresource URLs and has
since M59, and WebView2 is Chromium - so on the one platform this
product ships to, every camera tile on a shop counter was a broken
image. Measured against a running engine: the app's Go-side calls
returned stats and people while an <img> on that very URL failed, and
curl proved the URL answered 200. The engine was never the problem.
The password now stays on this side of the process boundary. A loopback
relay attaches Basic auth and streams the engine's bytes back
unchanged - the same reasoning Shot.jsx already follows at head office,
where an <img> equally cannot carry a session.
What the relay is careful about, since it is a door onto the biometric
API with a credential attached:
- loopback only, on a port the OS picks; a fixed one would collide
with whatever else a shop PC runs and read as "the cameras broke"
- a per-run random token in the path. The engine's own credential
exists so the live face feed is never served open; an
unauthenticated relay would hand that feed to any other process on
the PC. Compared in constant time, and a wrong one is 404, not 403
- an allow-list of stream.mjpeg and frame.jpg. Holding the token does
not reach the identity list, the gallery, or erasure
- camera ids validated, not interpolated
- every chunk flushed; a buffered MJPEG stream is a tile that never
paints, which looks identical to the bug being fixed
Two of those were written after a test failed, not before:
- `..` MATCHES the id pattern, because real camera ids contain dots.
`/api/cameras/../stream.mjpeg` is not the endpoint anyone intended.
The id can never hold a slash, so `.` and `..` are the whole
remaining traversal surface and are now refused by name.
- the serve goroutine read p.srv off the struct while stop() was
nilling it, so a quick start/stop dereferenced nil and took the
process down. Captured before launching now.
FrameURL is deliberately not added. No screen asks for a still, and a
bound method nothing calls is the same defect as a capability the UI
cannot reach, only pointing the other way.
Verified: nine unit tests, plus a live test against the real engine and
the real office camera - two MJPEG frames, 90,793 bytes, no credential
in the URL. Windows and darwin both build; vet clean.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pcn9asw19WGBfCEaHvNug6
`wails build` had never been executed against this project - CLAUDE.md
says so plainly - so every screen the shop floor actually touches was
unreviewed. Running it found why nobody had.
fyne.io/systray's nativeLoop must own the main thread on macOS, a Cocoa
requirement, and Wails already holds it. Starting both kills the process
with a SIGTRAP inside cgo before a single pixel is drawn. On Windows,
which is what ships, a tray on its own goroutine is fine - so the one
platform the whole team develops on was the one platform that could not
open the app, and the UI went unlooked-at as a result.
BEHAVISION_NO_TRAY runs the window without the tray, the same escape
hatch BEHAVISION_ALLOW_PLAINTEXT_MQTT already is for the broker.
Deliberately an environment variable and NOT a GOOS check: a build that
quietly drops the tray is how a shop PC ends up with no control surface
at all, and it would fail where nobody is watching. The guard is on stop()
as well, because systray.Quit() on a systray that never started is not a
no-op in v1.12.2 - it would turn closing the window into a crash on exit,
the failure most likely to be shrugged off as "it closed, fine".
go.mod gains the indirect dependencies the darwin build pulls in. No
version moved: the committed list was written by a windows-only build,
which never resolves that part of the Wails tree.
Verified: GOOS=windows build, go vet, and the agent suite all still pass,
and the packaged .app runs.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qiy5iKfz4L8S4vRaYPBdaU
Fixed on the web arrivals feed and not here, which is the failure this
codebase already warns about: two surfaces disagreeing about one fact.
Taking the first letter of each word of "Visitor 13" gives "V1" - and so
do "Visitor 10" and "Visitor 15", so three different customers wear the
same badge and it reads as the V-1 reference for a fourth.
Shows the number itself, same rule as the web app. customerRef, not
ref: React reserves that prop name and it would never arrive.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Every id in the schema is a uuid and stays one. What was wrong was
putting one in front of a person: RecordVisit named every new customer
'Visitor ' || left(id::text, 8), so the arrivals feed, the shop PC and
the mobile app all read "Visitor 3446ec35" - the string a shop assistant
reads to a colleague and types into a search box. label is a stored
column staff can overwrite and SearchVisitors matches on, so formatting
around it in a front end would have left the data wrong on three
surfaces.
Migration 012 adds a per-client visitors.number, taken from a counter on
clients with UPDATE ... RETURNING inside the visit transaction. Per
client rather than global: a global sequence would tell any customer who
signs up how many people the whole platform has ever seen, from their
own first visitor number. The backfill numbers existing rows by
first_seen_at and relabels only the eight-hex pattern the old statement
produced, so a human-typed name is never overwritten.
Three of the four things anyone addresses by URL already had a human
name and the API simply refused it - a site has a slug, a camera has the
id the engine knows it by. refs.go accepts either form anywhere an id is
taken; a uuid resolves with no lookup, so every URL a client already
stored keeps working.
- An ambiguous camera name resolves to nothing, never to a guess: two
shops may each have an "Office1" and acting on the first row would
edit the wrong shop's camera.
- 404 on a path, 400 on a query filter. /api/visits answered fine and it
was the filter that was wrong.
- site and site_id are both accepted everywhere now. They differed per
endpoint, and an unknown query parameter is silently ignored, so
getting it the wrong way round returned the whole estate.
- The search matches V-13, which is what the product now shows.
Two bugs found by running it rather than testing it:
- 'Visitor ' || $2::text beside number = $2 makes Postgres deduce two
types for one parameter and refuse the insert. It compiled and passed
every in-memory test; the first real database rejected it, along with
the existing face tests that share the path.
- The fallback avatar said "V1" for Visitor 13, Visitor 10 and Visitor
15 alike, and read as the V-1 reference for a fourth person. It shows
the number now. The prop is customerRef, not ref - React reserves
that name and it would never have arrived.
Verified on the live database and through the running API: 13 hex labels
became Visitor 1-13 in first-seen order, two typed names left alone, and
the same customer reachable by uuid, V-13 and 13.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
A tenant had exactly the users somebody had created with a command on the
server. That is not a missing screen: a shop with an owner and four staff
either shared one password or raised a ticket per person, and a phone app
for the shop floor could not exist while there was one account to sign in
as.
Registration is by invitation, never open signup - the same line already
drawn around creating a company. The code carries the address and the role
and the request carries only a password, so a code that gets forwarded
cannot become somebody else's account, and a staff invitation cannot be
redeemed as an owner. Single use lives in the UPDATE and the account is
created in the same transaction.
Deactivating a member revokes their sessions in that transaction too. An
access token lives twelve hours, so without it "remove their access"
removed it sometime tomorrow. The session list and revoke that go with it
are the benefit of opaque tokens the product had been paying for and never
collecting: nothing could say what was signed in, let alone stop one.
Face images now work on a deployment with no object storage, which was
every local install and every self-hosted site - the arrivals feed said
"not storing customer photos" for every customer forever, on the screen
whose whole job is to show a face. Bounded to one row per visitor, so it
grows with the customer base and not with footfall; the bucket stays
primary wherever one exists.
Image.auth says whether a URL needs the session, because a browser img
cannot load one that does, a mobile image view can, and a webview can do
neither - the desktop client resolves those to a data URI in Go.
Found by running it, not by tests:
* UPDATE ... RETURNING gives the value AFTER the update, so the prune
read back empty keys, deleted nothing, and the table grew with
footfall exactly as if it were not there. The fake agreed with either
version; only the live Postgres test caught it.
* Trusting only the auth flag broke every shop card, because Sites.jsx
rebuilt a partial snapshot object and dropped it. A relative URL is
now sufficient on its own.
* ago() renders a future time as "just now", so a code valid for a week
read "expires just now".
Verified live against real Postgres: invite, preview, escalation refused,
register into a session, replay 404, staff forbidden, device revoked and
401 at once, last owner refused, and a 92,405-byte camera JPEG stored,
served to its owner, 401 with no session, 404 to another tenant, and
rendered in a browser.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.
The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.
Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.
Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.
LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.
Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.
Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Migrations were run by hand and nothing recorded which had run, so
re-running the setup script against an existing database failed on the
first CREATE TABLE, and shipping a new migration gave an operator no way
to know whether an estate had it. A missed migration is not a startup
error - it is a query referencing a column that is not there, surfacing
later on whichever endpoint touches it first.
server/internal/migrate applies pending migrations at boot and refuses to
start against a schema it does not match. One transaction per file
holding both the DDL and the row that records it; an advisory lock so two
servers starting at once cannot both apply 008; checksums so an edited
migration is refused by name rather than silently skipped; numeric
ordering so 010 does not run before 009. `migrate -baseline N` adopts a
database built before any of this existed, because "the clients table
exists" does not say whether 007's index does.
Verified on the live database: adopted 001-007, applied 008.
008 adds two indexes on `purchases`, found by asking the database which
foreign keys had nothing behind them and then checking what queries the
table. The conversion report filters client_id + occurred_at, which is
exactly the estate-wide case with no site to narrow it.
run-local.sh had two bugs, both found by running it rather than reading
it: it reused a broker container whose bind mount pointed at a directory
that no longer existed, and it discarded stderr on the mosquitto_passwd
call, so under `set -e` it exited at step 5 with no output at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn