With the version ceiling and the widened numpy pin in place, setup succeeded
on the Mac that found them - Python 3.14 chosen and accepted, numpy 2.5.3,
onnxruntime 1.30, faiss 1.15.1, the engine itself - and died on the last step,
fetching the YuNet model:
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
certificate verify failed: unable to get local issuer certificate
A python.org macOS build ships its own OpenSSL with NO trust store, and
populates one only when somebody double-clicks Install Certificates.command in
the Python folder. Nobody installing face-recognition software has a reason to
know that exists, and the failure is forty lines of traceback about _ssl.c at
the end of a ten-minute install.
_urlopen tries the default context first and retries with certifi's bundle on
a verification failure. The order is the design:
- Default first, because on Windows and on a system or Homebrew Python the
default context reads the machine's own certificate store, which is what
makes a corporate proxy with its own root CA work. Replacing it
unconditionally would break every site that has one to fix a different
platform.
- certifi second, because it is already installed: requests is a hard
dependency and brings it.
- URLError is re-raised untouched. "No route to host" and "no trust store" are
different problems, and retrying the first with a different CA list only
delays the real message.
urlretrieve had to go, since it offers no way to pass a context - exactly the
kind of rewrite that silently drops something. The `download: <label> <n>%`
lines are a contract: supervisor.go's progressRe parses them to put first-run
progress in the tray, because the API is not up yet and a shop PC showing a
stopped engine for five minutes looks broken. A test asserts them, and the
rewritten fetch was checked against the real URL: 232,589 bytes, sha256
identical to the model already on disk.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
makeVenv reused any environment already on disk, whatever Python built it.
The machine that found the version bug already had a runtime built by 3.14,
left there by the run that failed - so with the ceiling in place setup would
choose a good interpreter, reach makeVenv, find the 3.14 environment, keep it,
and die in the same clang error as before.
A fix a user cannot reach because the bug's own debris is in the way is not a
fix, and it would have read as the release not working.
It now asks the interpreter inside an existing environment what it is and
rebuilds when the answer is unsupported, saying so. Rebuilding costs a
re-download of the libraries and nothing else - the models live in the state
root. An environment that cannot be asked counts as unusable too: a
half-created one answers nothing, and reusing it fails later in pip with an
error about a package rather than about the environment.
Tested against real environments rather than a fake, because what is under
test is what an interpreter on disk reports about itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.
## behavision-setup chose the Python least likely to work
findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".
maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.
## numpy<2.0 was the cap; OpenCV was the hazard
Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.
It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.
opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.
## One MQTT client id for a whole shop, so two PCs fought over it
behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:
broker connected / broker connection lost: EOF / broker connected / EOF ...
The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.
MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.
## "no such file or directory" for an engine nobody had installed
Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.
engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Two changes, and the second was found by verifying the first.
## Watching a camera from the app, in another building
Snapshots answer "is that camera working". They do not answer "what is
happening in my shop right now", which is what somebody who opens the app away
from the counter is asking. Head office's browser already had that answer -
LiveHub plus cameras.Live, where the shop PC asks outbound whether anybody is
watching and pushes JPEG frames for as long as somebody is - and the app could
not reach it.
cloud.CameraLive opens that feed and the app's own loopback relay re-emits it
as multipart MJPEG. That is the trick: frames arrive base64 over SSE, an <img>
cannot render that, and an <img> renders MJPEG natively - so a tile is an
ordinary <img> pointed at loopback whether the camera is in this room or
another city.
- Reconnecting happens in the relay, not the page. The server caps one push at
five minutes, so doing it here means the <img> never sees the stream end.
- The headers are flushed before the first frame. Go writes them on the first
body write, so without that the whole response waits for the shop PC to
start pushing. Measured against production: 30 seconds and not even a
Content-Type, which surfaces as the request timing out.
- One camera at a time. Watching makes a shop PC upload, so a grid that went
live at once would put an estate's worth of cameras on the wire because
somebody opened a page.
- live.mjpeg is behind the same per-run token as the engine routes, and a
wrong token is a 404 that never reaches head office at all.
- CameraLive uses its own HTTP client: the shared one's 30s timeout covers the
whole response and would sever a working view every thirty seconds - the
trap that made the server set WriteTimeout to zero for its own SSE endpoint.
## A camera read "Connected" for 34 minutes after the shop PC went blind
Which is why the verification above looked like a failure: head office
registered the viewer and no frame ever came.
reportWith returns early when the engine is unreachable - correctly, it has
nothing to say - so the last state it sent stays in the database looking
current. Measured live: cam2 and entrance both reading Connected, in green,
with last_seen_at 34 minutes old, while the heartbeat from the same PC said
cameras_up 0 of 0. Two surfaces reading two stored fields and disagreeing.
false could not be the answer. It means "this camera is not connecting", which
sends an installer to check cabling on a camera that was working perfectly the
last time anybody could ask it. So there are four states and one function:
connected reported recently, and working
not_connecting reported recently, and the stream will not open
waiting no shop PC has ever reported this camera
stale reported once, and not lately
- Connected is CLEARED when stale or waiting. A stale true left in place stays
available to every client reading the field directly, and leaves two fields
on one object disagreeing - how the shops screen once came out labelled
Working, in green, above "2 of 3 cameras not connecting".
- Computed in scanCamera, so every camera anybody reads passes through it. A
state computed per handler is one a handler forgets, and this had already
reached three screens.
- CameraStaleAfter is 5 minutes: five missed reports, not one. Same reasoning
as three missed heartbeats - an indicator that cries wolf gets ignored.
- An unparseable last_seen_at is stale. It should be impossible, which is why
it must not fall through to the state that says everything is fine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Signing in on a second Mac showed "engine not reachable at
http://127.0.0.1:8010" and 0 of 0 cameras, on an account whose shops were
running and recognising people the whole time. Nothing was broken: Live() and
Cameras() read only the engine on loopback, so the app answered as though the
person had never signed in - and camera sync goes through the engine, which is
why the count was zero rather than stale.
Having no engine is a normal state. A shop PC watches cameras; an owner's
laptop, a manager's machine and a second till being set up do not, and all
three are signed in to the same estate. Both methods now fall back to head
office when loopback fails and somebody is signed in. Loopback is still tried
first: a real shop PC must never be shown a minute-old summary when the engine
two milliseconds away has the live one.
Decisions worth keeping:
- Viewing is on the snapshot, not inferred per screen. Three surfaces read it,
and a screen that computed it separately is how the shops screen once came
out labelled Working, in green, above "2 of 3 cameras not connecting".
- fraction_below_gate takes the WORST shop, never an average. 0.10 against
0.73 averages to 0.42 and hides the only shop anyone needs to visit.
- A remote camera is flagged, and Edit, Remove and Check placement are
withheld. They talk to a camera on a LAN this computer cannot reach, and a
button that cannot work is worse than one that is absent.
- connected is three states. null is "no shop computer has reported yet" and
reads as waiting; false is "Not connecting". A bare false sends somebody to
check cabling on a camera nobody has tried to reach.
- Snapshots are fetched in Go as data: URIs and cached by snapshot_at. A
webview <img> resolves a relative src against wails:// and cannot send the
bearer - the problem VisitorImage already solved - and this screen polls
every 8 seconds at ~90 KB a camera.
- With no engine AND no session, the engine error is still the answer. The
person is most likely setting this PC up.
The picture is the last snapshot and the banner says so: there is no live
video from here, because the engine's MJPEG stream is on the shop PC's
loopback behind a router with no inbound route. The LiveHub relay head office
uses is the answer to that and is a further step for this client.
Verified against production: five arrivals and two cameras parsed from the
real API. viewing_test.go covers the fallback, the worst-shop rule, the
withheld credentials and that an unchanged snapshot is fetched once across
two polls.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Found by running behavision-setup against a clean state directory the
way a second machine will, which had never been done. It failed at the
first step:
Setup did not finish: no Python 3.10 or newer was found
Found, but too old: python3 3.9
on a machine that has 3.12. The search was `python3` then `python`, and
on macOS `/usr/bin/python3` is ALWAYS the Command Line Tools build -
3.9 on current macOS, below the 3.10 floor. Anything newer installs as
`python3.12`, under Homebrew, as a framework, or somewhere a GUI
application's minimal PATH never sees.
So it now tries versioned names newest-first, then the plain ones, then
the four directories macOS actually uses - and absolute candidates are
stat'd rather than passed to LookPath, which only searches PATH. It
found /Users/tenext/.local/opt/python3.12/bin/python3.12, which is
exactly the interpreter it had been ignoring.
With that, the whole install completes on a Mac for the first time:
venv, engine and dependencies, models, agent.json, and "Engine starts
and answers - verified". EXIT=0, a 298 MB runtime.
Also the last thing it prints, which is the first thing an operator
acts on. It said "Start Behavision from the Start menu" and "it appears
in the system tray; right-click there to stop it". On macOS there is no
Start menu and, deliberately, no tray at all - so the finishing message
was describing a machine the user was not sitting at, on the one step
where setup had otherwise succeeded. It now says to right-click the app
the first time because the build is not notarised, and that closing the
window stops recognition.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ran staticcheck across all three Go modules for the first time. server
(23k lines) and desktop came back clean. agent had seven findings, and
one of them was not tidiness.
`stopGrace = 10 * time.Second` was declared and wired to nothing.
Stop() cancels the context, cmd.Cancel kills the process tree, and then
Stop() blocks on cmd.Wait() - which, with no WaitDelay set, waits not
just for the process but for every writer of its stdout pipe to close.
One grandchild still holding that pipe hangs Wait, hangs Stop, and on the
desktop app that is the tray's Quit never returning. The constant named
the intent and nothing read it. cmd.WaitDelay = stopGrace is the line
that was missing.
The rest were real but small: an unused field in the live relay, an
unused sleep helper in the pump, and "net/url" imported twice under two
names - both genuinely used, in two functions doing the same job for the
same reason, so they are unified rather than one deleted. My first pass
deleted the wrong one on a bad grep and the build caught it immediately.
Three findings are suppressed rather than fixed, with the reason stated:
- Two "error strings should not end with punctuation". Both are
multi-line messages a shop operator reads at a counter, not errors
anything wraps. ST1005 exists because wrapped errors concatenate
mid-sentence; stripping the full stops would run three sentences
together to satisfy a rule that does not apply.
- A deliberately nil context in a pump test - the point of the test is
that an unconnected client does not panic. It already carried
//nolint:staticcheck, which is golangci-lint's directive and
staticcheck ignores, which is why it kept being reported.
Also tidied agent/go.mod, which had paho and x/sys marked indirect while
being imported directly.
All three modules clean, all suites pass: 21 Go packages, 226 engine
tests.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
511 lines extracted from the code at release 0.4.1 / schema 013, and
accurate for that point. The repository is nine releases and a schema
past it.
A stale document that states its own version reads as current to anyone
skimming, which is the same failure this project keeps catching
elsewhere: wrong in a way nobody can detect. So the top now lists what
it predates by name - the motion gate, Gallery.health, tenantOnly, the
password endpoint, customers and merge, the admin drill-down, sales and
dashboard, migration 014, the macOS build - and points at API.md and
CLAUDE.md, which are kept current.
No credential values in it; the matches for password/secret/token are
environment variable NAMES and package paths describing where secrets
live, which is what a dossier should say.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Cheap for a reason worth stating: the Windows package is already a SOURCE
install - a pure-Python wheel plus a setup tool that builds a venv on the
target machine, because PyInstaller cannot cross-compile. macOS needs
nothing different. Same wheel, same setup tool, a natively built .app in
place of the .exe. No frozen engine, no 200 MB, no second packaging story.
behavision-setup already handled both platforms (Scripts vs bin, the
tasklist check guarded) and cross-compiled for darwin without a change.
The one thing that did not was its advice when Python is missing: it told
everyone to tick 'Add python.exe to PATH' on a Windows installer page.
Software that does not know which machine it is running on is software
somebody stops trusting for the rest of the session.
MAC=1 opts in, so the ordinary Windows release is unchanged.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Reported from the Mac build: the window cannot be maximised. It is not a
Wails limitation or a WebView quirk, it is an omission with a very
specific consequence.
Wails computes zoomable INSIDE `if frontendOptions.Mac != nil`:
var fullSizeContent, hideTitleBar, zoomable, ... C.int // 0
if frontendOptions.Mac != nil {
zoomable = bool2Cint(!frontendOptions.Mac.DisableZoom)
}
and the native side then acts on the zero:
if (!zoomable && resizable) {
NSButton *button = [self.mainWindow
standardWindowButton:NSWindowZoomButton];
[button setEnabled: NO];
}
So leaving Mac unset does not mean "take the defaults" - it means the
green button is created and then explicitly disabled. There was a Windows
options block and no Mac one, which is how this survived: the platform
that was configured behaved, and the platform that was not looked broken.
Fixed by the block existing. The fields are written out rather than left
as an empty struct so it reads as a decision rather than something half
typed.
Verified at runtime rather than by reasoning about the source alone: all
three title-bar buttons report enabled=true through the accessibility
API, and the window resizes to 1440x900, the full display.
One correction to my own first check, recorded because it nearly sent me
the wrong way: querying AXFullScreenButton as an ATTRIBUTE of the window
returns "missing value" whether or not the button exists. It has to be
found by subrole among the window's buttons. The button was fine; the
question was wrong.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Nothing ever set the child's working directory, so it took the parent's -
and an app started by double-clicking its bundle is handed "/", not
anywhere useful. On macOS the symptom was
`python: No module named behavision` repeating forever, because the dev
engine is invoked as `-m behavision` and that resolves against the
working directory.
The same app launched from a terminal inside the repo worked perfectly,
which is exactly the shape of a bug that survives every test a developer
runs. It only appeared when the app was started the way a user starts
one.
Config.EngineDir, empty meaning the install root, set by both launchers -
the desktop app and the headless agent, which had identical code and the
identical omission. It matters beyond this case: the shipped Windows
engine is a one-folder PyInstaller build whose relative paths should
resolve beside itself rather than beside Explorer's idea of a current
directory.
Verified by double-clicking the bundle with nothing in the environment:
engine up on 8010 (401, gated), w600k_r50 on CoreML, gallery 5/5
embeddings usable and none stranded, both office cameras connected and
streaming, and head office reporting cameras 2/2 one heartbeat later.
Two things that showed up while proving it, both the product being
honest rather than faults:
- The camera at .121 was genuinely unreachable for several minutes, and
last_error said so in words an installer can act on - "cannot reach
192.168.1.121:554 - No route" - rather than `connected: false`. That
field was added yesterday for precisely this.
- Head office briefly showed cameras 0/1 against a local 2/2. That is a
60-second heartbeat, not a disagreement; the next one read 2/2.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Chosen deliberately as a DEVELOPER build, not a product. Indian retail
counters are Windows; shipping a Mac product means an Apple Developer
account, notarisation, a second installer format, a second frozen engine
and DPAPI having no macOS equivalent - a permanent second platform for
customers who do not have Macs. What a Mac build is worth is demoing the
desktop app on the machine it is written on, without needing the Windows
box.
It built after one missing framework (previous commit) and then crashed
within a second, twice, both times in the tray:
systray.Run SIGTRAP inside cgo. nativeLoop() takes the
macOS main run loop for itself and Wails
already has it. macOS has exactly one.
RunWithExternalLoop "NSWindow should only be instantiated on the
main thread!" - it registers in the existing
NSApplication rather than starting a second,
but still builds AppKit objects, and Wails'
OnStartup is not the main thread.
Making it work needs the status item created through a main-queue
dispatch inside Wails' lifecycle. That is real work for a build whose
purpose is a demo, so macOS has no tray and the file says so at length
rather than leaving the next person to rediscover both crashes.
The consequence is handled rather than left lying. With no tray there is
no way back from a hidden window and no way to quit, so hiding on close
would strand a running engine behind no window, no tray and no control -
force-quit or nothing. On macOS closing the window therefore quits, and
OnShutdown stops the engine. Same rule the tray's Quit already follows:
never leave it watching with no visible control. Windows is untouched,
where hiding is correct because the tray is how it comes back.
The runner is split by build tag rather than branched at runtime because
the two platforms need different systray ENTRY POINTS, not different
arguments.
Verified: 18 seconds up, zero crash markers, 88 MB resident, and an
honest "engine not installed yet" instead of a crash - against a
throwaway data dir so it claimed nothing and touched no camera. The
frozen Mac engine is deliberately not built; the app takes an engine
command from config, which is how the dev setup already points at the
venv.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Reported from the shipped Windows app. Three faults in one call, and the
first is why it failed rather than merely misbehaved.
runtime.Show is implemented by Wails as a bare mainWindow.Show(), while
runtime.WindowShow wraps the identical work in runtime.LockOSThread. Win32
window operations have to run on the thread owning the window's message
pump, and the tray's handler runs on the SYSTRAY's goroutine, which is
never that thread. An unlocked Win32 call from an arbitrary goroutine is
the bug.
Two more that would each have been enough on their own:
- Showing is not un-minimising. Hidden and minimised are different states
and Show only fixes the first, so a window the user minimised stayed
minimised.
- Windows refuses the foreground to a process that does not already hold
it, so the window came back BEHIND whatever was being looked at.
Clicking a tray icon is by definition a moment when this app is not in
front, so that is not an edge case here - it is every time. The
always-on-top flip is the ordinary way to ask, and it is why this now
runs in a goroutine rather than on the menu loop, which must not sleep.
The same four calls fix OnSecondInstanceLaunch, which had the same shape
and is reached far more often: double-clicking the desktop icon while the
app is already running.
OnBeforeClose used runtime.Hide against a reopen that used WindowShow -
different calls on Windows, one thread-locked and one not. Paired now.
And a Mac build, because the question came up and the answer turned out
to be yes. Wails' darwin frontend references UTType without linking
UniformTypeIdentifiers, so the build failed at the LINK step after
compiling everything - which reads like a broken toolchain rather than
one missing flag. There was no Mac version because of that, not because
of a design limit. darwin_link.go declares the framework in source rather
than leaving it as a CGO_LDFLAGS incantation, for the same reason
deploy.sh now finds Go itself. Verified: plain `go build` produces a
16 MB arm64 binary on this Mac, and the Windows build is unchanged.
Worth knowing for whoever edits that file: the comment directly above
`import "C"` is cgo's C preamble, not documentation. The first attempt put
the explanation there and the prose was compiled as C.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The merge records what it had to discard in the survivor's notes, and
search did not look there - so the value was retained and unfindable,
which answers the letter of "nothing is lost" and not the point of it. A
customer reached by their old number is exactly who somebody is looking
for when they type it.
Caught in the same patch: I wrote ESCAPE with two backslashes where the
four clauses beside it use one. In a Go raw string that is two literal
backslashes, and Postgres requires the escape to be a single character -
it would have failed the whole customer search at runtime, on a query no
in-memory test executes. All five clauses are identical now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Found by walking the scenario against production rather than by a test.
Two records, each with a phone; the survivor kept its own, and the
source's simply stopped existing. Searching for it returned nothing.
The first version's rule was "fill the survivor's blanks, never overwrite
what it has", which is right about which value WINS and said nothing
about the one that loses. One person can have two numbers, two spellings
of a name, a work address and a personal one - and a merge that quietly
deletes one is exactly the data loss this file already refuses elsewhere:
"silently turning Alice back into Visitor 3 is data loss the operator
cannot see happen."
The profile is now reconciled field by field in Go rather than in one
clever upsert, because the interesting case was never the winner. Blanks
are still filled and the survivor still keeps its own values, but every
losing value is returned in `discarded` AND appended to the survivor's
notes - the response is read once and the record is read forever.
Notes themselves are additive rather than a winner: two people writing
about one customer wrote two different true things.
mergeProfiles is pure, so the rule is asserted directly - four cases
including the ordinary one, a typed record joining a camera record with
no profile at all, which must add no noise.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
POST /api/customers and POST /api/visitors/{id}/merge. They ship together
because the first creates the need for the second: a customer typed in at
a counter has no face template, so when a camera sees that person later
the matcher has nothing to compare against and enrols them as somebody
new. That is the design working, not failing - and it means every
hand-created customer is a duplicate waiting to happen. Shipping the
create alone would manufacture duplicates into the state CLAUDE.md
already flags: "there is no merge endpoint server-side, so its
duplicates would be unrecoverable."
The number comes from clients.visitor_seq, taken exactly as RecordVisit
takes it. Two sources of visitor numbers that could disagree would be
worse than none: V-42 has to mean one person whichever way they arrived.
The label is the typed name, or "Visitor N" when they gave none - the
same string the engine writes, so a record created by hand is
indistinguishable from an enrolled one afterwards.
The merge is one transaction over FIVE tables, and the count is the
point. visits, purchases, visitor_embeddings, consents and
visitor_profiles all reference visitors ON DELETE CASCADE, so a table
this forgets to re-point is not an error - those rows are destroyed with
the source and nobody finds out until a customer's history is short.
visitor_profiles is UNIQUE on visitor_id, so the two cannot simply both
move and something has to win. Blanks on the survivor are filled from the
source and nothing it already holds is overwritten, which is exactly
right for the case this exists for: a hand-typed name and phone joining
the face that was recognised a week later.
Policies carried over from the edge gallery's merge, which had to settle
all of this once already: a human-assigned name outranks an auto
"Visitor N" whichever direction the operator merged; visit_count is
recomputed with COUNT(*) and never summed, because the stored counter may
be stale and the row count cannot be; first_seen_at takes the earlier of
the two, since it is one person and always was.
Two things that are this side's own:
- The source is deleted for real, not soft-deleted. A tombstone would
leave its number resolving to a record holding nothing, which reads as
"this customer exists and has never been here" - a worse answer than
"no such customer".
- The response names the RETIRED reference. Staff write V-42 on cards and
read it aloud; a merge that does not say which one stopped working
leaves somebody to discover it at a counter.
Manager and above, not staff. Apart from erasure this is the only
irreversible operation on a customer: two people welded together cannot
be separated, because nothing records which visit came from whom. It logs
at WARNING and writes an audit row for the same reason.
Also fixed while here: two s.Log.Printf calls - one of them mine, from
the password endpoint - that would panic on a nil logger. The package has
a nil-guarded s.logf and those were the only two not using it. The
password one sat in an error path no test reaches, which is exactly where
that bug waits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
POST /api/auth/password. The cost of its absence was measured today
rather than argued: rotating three production accounts took a shell on
the host, three round trips, and briefly left a PLATFORM ADMIN - the
account that reads every company on the estate - with the password
PASTE_IT_HERE, because a placeholder in a pasted command was taken
literally and there was no way to correct it from the product.
A manager could always reset somebody ELSE's password. A platform admin
could be reset by nobody: they have no client, so the team routes are
not theirs, and `provision user` on the host was the only route. For
software that puts accounts on shop-floor PCs and staff phones, this is
not a feature - it is what makes every other credential decision
recoverable.
Three decisions:
- **authed, not tenantOnly.** A session is not a company's data, and the
account with no company is precisely the one that had no route. Scoping
this by client would have reproduced the hole it exists to close, which
is also why SetUserPassword is not scoped by client the way
ResetMemberPassword beside it is. The user id comes from the verified
session, never the request, so there is nothing to point at anyone else.
- **The current password is required.** An access token lives twelve
hours and travels on devices that get lost and shared; without this a
stolen one owns the account permanently instead of until it expires.
- **Every OTHER session is revoked, and the caller's is kept.** Somebody
changing their password because they believe it is known must not have
to wonder whether the device that already had it is still signed in -
and must not be signed out of the one in their hand while dealing with
it. A failure there is logged, not returned: the password IS changed by
then, and reporting an error would send them to retry with a current
password that no longer exists.
The suite's login() helper fatals on anything but 200, which is right
everywhere else and useless here - half of what these tests assert is
that a password has STOPPED working. loginCode() returns the status.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
"When was this customer last in" and "did they buy" are one question
staff ask in one breath, and answering it meant two calls and a join in
the client. Each visit row carries purchases, spend and currency.
LATERAL, not a join onto purchases. A plain join returns the visit TWICE
when it holds two sales, which would make a customer look like they came
more often than they did - a wrong number of exactly the kind this
product is otherwise careful about, arrived at by adding a feature.
Mixed currencies on one visit report the count and NO figure. Adding
rupees to dollars produces something that looks like money and is not,
and the sales still happened, so the count is the honest part to keep.
A purchase with no visit_id is deliberately absent: it belongs to the
customer rather than to a moment, and GET /api/sales?customer=V-42 lists
it. The two surfaces together cover every sale exactly once.
Both properties are asserted in the LIVE store tests, because both live
in the SQL. An in-memory fake asserting that a LATERAL does not duplicate
a row would only be checking the fake.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Audited before its first run, because deploy.sh taught us what a script
nobody has executed contains.
PowerShell's $ErrorActionPreference = "Stop" governs PowerShell errors. A
native .exe returning non-zero is not one, so `npm ci` and `npm run build`
were unchecked and the script sailed past them. That matters here more
than anywhere else: a built frontend/dist is COMMITTED to this repository
so `go build` type-checks without npm, which means a silently failed npm
build leaves the old one in place and it embeds perfectly. The output is
an installer that builds, installs, opens and shows a stale UI, with
nothing anywhere saying so - the silent-wrong outcome, reached through
the single most likely failure on a fresh Windows box.
A Run() helper now throws on any non-zero native exit, across nine call
sites: venv, both pip installs, pytest, pyinstaller, npm ci, npm build,
both go builds, and the frozen engine's own smoke test.
The pip installs were also piped to Out-Null, so a failure there produced
no output AND no stop. run-local.sh has already been caught making
exactly that mistake, where it "exited at step 5 with no output at all -
the single hardest failure to diagnose, and it took three runs to find".
Not worth repeating in a script that runs on a machine nobody is sitting
at.
Two smaller ones from the same read:
- frontend\dist\index.html is deleted before npm runs, and its absence
afterwards is an error. Checking the exit code is not enough when the
artefact it was meant to produce is already sitting there from git.
- `go build -o dist\...` does not create its target directory, and dist\
is gitignored. It exists on a fresh clone only because PyInstaller ran
first and made it - an ordering dependency nothing stated. Stated now,
and created explicitly.
None of this has been run on Windows. It cannot be from here - PyInstaller
freezes the interpreter and native wheels of the machine it runs on. What
this buys is that the first Windows run fails for a real reason rather
than for a bug in the script.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The admin drill-down, the sales reads and the dashboard summary, each
with the shape production actually returns - copied from live responses
rather than written from the structs, because that is the difference
between documentation and a guess.
Three things stated because a client would otherwise get them wrong:
admin camera rows are a DIFFERENT shape from GET /api/cameras and carry
no host, port, path, username or has_password; a sale with no visitor is
listed rather than joined away, so this agrees with the conversion report
over the same rows; and the sales list has no cursor, with the reason,
because purchases has no monotonic column and a cursor would imply a
delivery guarantee it cannot make.
Also the boundary the work exposed: 'authed' meant any signed-in user,
and a platform admin is signed in with no company at all. That now has a
sentence and a code (403 not_a_tenant_account) instead of being a 500
nobody had called.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
/api/visits, /api/cameras, /api/sites, /api/visitors and
/api/reports/footfall, all in production, all before today's work. A
platform admin is defined by having NO client, and every tenant query
scopes on client_id = $1::uuid - so the empty string reaches Postgres as
''::uuid, which is a cast ERROR rather than an empty result. Found by
calling them while verifying the new routes, which have the same shape
and were failing the same way.
tenantOnly is the guard, beside adminOnly and for the opposite audience.
Per-query casts would have been the wrong fix twice over: it is a fix the
next query forgets, and the next query would then 500 in production
exactly as these did.
403, not adminOnly's 404, because the two hide opposite things. A tenant
must not learn a platform surface exists. A platform admin already knows
the tenant surface does - they are reading its data through /api/admin -
so nothing is concealed by pretending otherwise, and the refusal names
the route to use instead. "Forbidden" alone sends somebody hunting a
permissions problem that does not exist.
/api/auth/* stays on plain authed: a session is not a company's data, and
signing out or revoking a lost device must keep working for an account
with no tenant.
The fake could not have caught this either - it compares client ids as
strings and is perfectly content with "". The test asserts the contract
(403 and a message naming /api/admin) and a third case that matters more
than either: an ordinary tenant user still reaches all of it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
/api/admin/clients/not-a-uuid/sites returned 500. `c.id = $1::uuid` makes
Postgres cast the path segment, and casting a malformed string - or the
empty one the shape check handed back in its place - is an ERROR, not a
miss. `c.id::text = $1` cannot fail: an id that is not a uuid matches
nothing, which is the 404 a wrong URL should get.
The two sibling resolvers were already written this way and correctly
404'd the same input. I applied the rule to two of three places, which is
the shape of a rule that holds until somebody adds the next write path.
The shape check is gone with it - it existed only to produce the empty
string that then broke the cast.
The in-memory fake could not have caught this and did not: it resolves a
merchant with a map lookup, so every handler test passed, including the
one named for the case. That test stays, because 404-not-500 is still the
contract, but the property belongs to Postgres - so
api_admin_monitor_live_test.go asserts it where it lives, over every
free-text identifier these queries take. It skips without
TEST_DATABASE_URL, like the rest of the live store tests.
Also in deploy.sh, found by reading its own output: step 3 reported the
WRONG backup. `ls | tail -1` sorts alphabetically, so pre-...-demo-12
sorts before pre-...-demo-6 and it printed a dump from four days earlier.
A deploy that names the wrong safety net is worse than one that names
none, because that is the file somebody reaches for at the worst possible
moment. It echoes the filename it just wrote, and refuses to continue on
an empty one - pipefail catches a failing pg_dump, but a zero-byte gzip
would still have satisfied it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
"go: command not found". Go sits in a directory the operator's .zprofile
adds and a script does not inherit, so the very first step of the deploy
failed for a reason having nothing to do with the deploy. Found the only
way it could be - by somebody running it - and a deploy that needs the
operator to fix their environment before it works is a deploy that gets
skipped, which is the failure this script exists to end.
It now looks in the three places Go actually lands and says so plainly if
it finds none.
Step 7 also verified five routes and none of them were the nine that
shipped in the last two commits. It checks all of them now, and treats
401 as a PASS on purpose: an unauthenticated call to a route that exists
is refused, while a route the binary never registered is a 404. That
makes this step prove the ROUTING rather than the auth - which is
precisely what a deploy gets wrong, and what otherwise surfaces weeks
later as a console reporting "Backend integration required" against an
API that had already shipped. A missing route now fails the deploy loudly
instead of printing a number nobody reads.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Three routes over data the server already stores.
GET /api/sales and /api/sales/{id}. The purchases table has existed
since the conversion report did, and nothing could read a row of it - so
"revenue was 41,000 last week" was a number that could not be checked
against a till. The list carries the customer reference the product
actually shows people (V-42) beside the uuid, and a sale with NO
customer is listed rather than joined away: an unidentified walk-in is
still revenue, and an inner join would make this disagree with the
conversion report computed over the same rows.
No cursor, deliberately. A keyset cursor needs a monotonic
server-assigned column and purchases has none; ordering by
(occurred_at, id) with a random uuid tie-break is exactly the shape that
silently dropped four of six simultaneous visits from the arrivals feed
before visits.seq existed. Offering one here would imply a delivery
guarantee this table cannot make, so the list is bounded by the date
window and a limit - which is how a sales list is browsed anyway.
GET /api/dashboard/summary. Four calls a client had to make and then
combine, which is how the desktop Footfall screen once produced its
headline by adding the daily bars up: silently too high, because a
customer who came twice is one person and two bucket-visitors. The
combining happens here, against Footfall and SiteHealth rather than new
SQL - a second definition of "unique visitor" or of "online" drifts, and
a home screen that disagrees with the report it links to is the one
nobody trusts afterwards. fraction_below_gate travels with the count for
the same reason it does everywhere else: it is what says whether the
headcount is a number or a floor.
Today is cut in the shop's timezone. In the one market this ships to,
UTC is five and a half hours wrong.
An unknown shop filter is a 400, not an ignored parameter. This API has
already been bitten once by a silently ignored filter handing back the
whole estate, which is a wrong number nobody would question.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Six read-only routes: merchant detail, its shops, one shop, its cameras,
one camera, and the platform totals. The console drills down
merchant -> store -> camera and every level below the first showed
'Backend integration required'.
They cannot be the tenant routes, and the reason is structural rather
than incidental. Every tenant handler derives the client from the
SESSION - that is what makes cross-tenant access impossible rather than
merely disallowed - and a platform admin has no client at all. The three
workarounds each make it worse: passing a company id to a tenant route
puts a caller-chosen tenant back in the one place this system refuses to
take one, filtering the estate in the browser ships every merchant's
data to render one, and signing in as the owner audits the wrong person.
So the tenant STORE functions are reused with an explicit client id -
they already take one - and the scoping the tenant handlers get from the
session is done in the handler instead.
AdminCamera is a separate type from Camera, for the same reason
AgentCamera is. It cannot carry host, port, path, username or
has_password. A tenant seeing those for their own camera is correct; a
platform admin browsing another company's estate is a different
question, and an RTSP host with a username beside it is most of a live
path into a customer's camera. Blanking fields on a shared struct leaves
'remember to redact, on every path, forever' as the only thing
preventing a leak. The test asserts on the raw JSON, because decoding
into the struct would discard exactly what it is looking for.
An unowned site is 404, never an empty list. The tenant resolver returns
a uuid untouched and lets client_id = downstream scope it, which is
sound only because that id comes from a session; here the caller names
both halves, so an unowned uuid would reach a query that quietly returns
nothing - 'this shop has no cameras' when the truth is 'not your shop'.
Both resolvers check the whole chain in one statement.
Two things the in-memory fake could not have caught, so neither was left
to it. The fake ignored clientID in SiteHealth and Cameras, which would
have made every cross-merchant test pass while returning another
company's shops; it is client-aware now for these paths. And the SQL was
written to make the documented $2-deduced-as-two-types bug impossible
rather than to be caught by a database later: id::text = $2 in place
of id = $2::uuid, one type per parameter, which also turns a malformed
path segment into the 404 it should be instead of a cast error.
Every read below the merchant list writes an audit row naming the admin
and the merchant - an admin is the one account for which nothing else
here leaves a trace. The counts-only summary does not: a console
refreshes it on a timer, and logging that buries the reads worth
finding. A suspended merchant stays readable, because that is precisely
what an admin opens the console to look at.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Audited the engine for what it does when something goes wrong rather than
when it goes right. Each of these left the process healthy, the dashboard
green and the product not working.
A gallery the running encoder cannot read. Embeddings are model-tagged, so
when the fallback chain fires every vector the previous encoder wrote goes
invisible: the shop keeps its customer list and recognises nobody on it,
enrolling each regular a second time. Footfall stays correct, which is why
nothing looks wrong. The only evidence was an INFO line reading 'gallery
ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence
that states the disaster and calls it ready. Gallery.health now warns with
the count of PEOPLE lost, not vectors, and carries the same numbers to
/api/stats and /api/health, because a log line on a shop PC is read by
nobody. Proved against the real 87-embedding gallery.
Connected, and sending nothing. 'connected' meant the socket opened, so a
stream that went quiet kept it true while last_frame_age_s climbed and the
heartbeat told head office the camera was up. OpenCV breaks a blocked read
at 30s, but a camera trickling a frame every 20s never trips that and never
recovers. streaming/stalled are reported beside connected and the dashboard
says live/stalled/offline - three states because offline sends you to the
network and stalled says the camera is answering and sending nothing.
The 5-second RTSP timeout that never existed. stimeout;5000000 carried a
comment claiming it bounded a dead camera. Measured on OpenCV 4.11 /
FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with
stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so
it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither
reaches the RTSP protocol through this path; the real bound is OpenCV's own
interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source
already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause.
That matters beyond speed - the VideoCapture constructor is not
interruptible, so stop() could not cut it short and a camera removed from
head office left a daemon thread holding a socket for half a minute.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Measured rather than guessed, and the first guess was wrong. Wall clock
said H.265 decode cost 58 ms a frame; cap.read() blocks until the next
frame arrives, so that was the frame interval, not work. As CPU time:
decode 3.7 ms, detection 31.0 ms - and detection ran on every frame
whether or not anything was in front of the camera, 6,649 of 8,634
frames with faces_seen 0 and active_tracks 0 throughout.
detect_threads: OpenCV spreads a small repeated job over eight threads,
costing 31.0 ms of CPU for 8.9 ms of wall. One thread costs 15.3 ms for
15.3 ms, against a 66 ms budget at 15 fps. Half the CPU for latency
nothing can notice.
motion_gate: a 160x90 greyscale absdiff, 0.1 ms against detection's 15.
Consulted only while no track is open; forced to look every
motion_max_skip frames; compared against the last frame SEARCHED so a
slow drift cannot creep under the threshold; and a threshold above this
camera's measured noise and far below a person, so anything ambiguous
detects. tests/test_motion_gate.py pins each of those rather than the
saving, including asserting the longest run of skips rather than the
total - counting the total would pass a gate that slept forty frames
and then looked forty times.
Together 80% -> 16% of a core, detection skipped on 92% of frames.
faces_seen is still 0 and the gate is not why: run directly over the
same frames the detector finds nothing at threshold 0.50 either. The
placement is the limit, as recorded; the CPU was being spent to
rediscover that fifteen times a second.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The camera used an MJPEG stream through the proxy and the engine's
stats under names it does not use (frames/faces rather than
frames_processed/faces_seen), so the picture was blank and the counter
read zero. The engine re-serves its latest frame until the pipeline
produces a new one, so a polled still is the same picture with none of
the multipart fragility - which matters when the audience is in the
room.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Every other thing in this product a person refers to already had a
readable reference: a shop is chennai, a camera cam1, a customer V-42, a
person their email. An audit of every list response found exactly one
gap, and it was the row people look at most - the arrivals feed showed a
visit as 36 hex characters.
012 argued no route takes a visit id so none was needed. That is true of
routing and false of everything else: it is what the feed shows, what a
support conversation quotes, and what somebody reading an API response
judges the product by.
Migration 014 mirrors the visitor scheme exactly - per client, so it
discloses no platform-wide volume, and beside the uuid rather than
instead of it. A stored counter is affordable on the busiest table
because visits from one tenant are already serialised by the consumer's
SetOrderMatters(true), so it adds no contention that was not already
there. A derived reference was the alternative and does not work:
several people through one door share occurred_at to the microsecond,
which is the collision 004 exists to handle.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The agent read the engine's generated credential file once, at startup.
On a brand new install that file does not exist yet: the agent starts the
engine, and the engine writes its credential seconds later. So the agent
held an empty credential for the life of the process and every call it
makes - health, stats, camera sync, the embedding for a visit - came back
401, with a tray showing a red engine that was running perfectly.
Measured on a fresh state directory today: three 401s, no camera ever
reconciled, and the engine left running the YAML-seeded main stream
instead of the sub-stream head office holds. The install script hid this
on Windows because setup runs the engine once before the app starts.
config.Creds resolves lazily and re-reads on a rejection; the camera
client, the supervisor and the desktop app's engine client all retry once
when it changes. A configured BEHAVISION_API_USER is never re-read - an
operator who set one means it. Tests pin the actual first-run ordering.
Also adds demo/, a one-screen live console for showing the whole chain:
camera, the six steps with a measured camera-to-cloud latency, the
customer editable in place, and the raw JSON a phone and a dashboard
receive from production side by side.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
GET /api/cameras read only site_id, while every other filtered endpoint
takes both spellings through siteParam. So ?site=chennai was not a
filter at all but an unknown query parameter, silently ignored, and the
caller got every camera in the tenant believing it had one shop's.
Found by using it: a setup script saw another shop's cameras, concluded
three shops already had theirs and created none; then a delete aimed at
a test shop removed the live Coimbatore entrance camera, which had to be
restored. This is exactly the hazard already recorded for site vs
site_id - the note existed, the handler was simply missed.
One line to fix, and a test that asserts the whole class rather than
this one route: both spellings must narrow, and only an absent filter
may return more than one shop.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Production had 1,211 events accepted and six recognised customers from
the office cameras - the first time the whole chain has carried a real
person, and the project had never been able to claim it. Repeat
sightings score 0.44-0.72, a distribution the match threshold sits
clearly below, on the head-height camera this file has recommended since
August. fraction_below_gate is still 0.59, so the visit count is a floor
and the report says so beside it.
The face-image chain was exercised on production as a shop PC does it -
upload URL, PUT to object storage, anonymous read refused 403. Every
server link holds; the only reason a customer has no photo is
app.store_faces being false by default, which is a data-protection
decision rather than a gap.
Sixteen mobile-API checks pass as a staff account. Three apparent bugs
were test errors and are written down so nobody re-files them, along
with the one field name a caller could guess wrong (site_token).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The first launch was a code box with a link under it, then an empty
Live screen with 'No cameras' in amber in a far corner, then a form
asking for an IP address, and for the first few minutes of all of it
the engine silently downloading 275 MB with nothing on screen but a
stopped-looking status. Walked in a browser with the new mock; nobody
who was not an installer would have got through it.
Now: a welcome that asks the one question a shop owner can answer -
managed from a head office, or on this PC only - with each path in a
sentence; a Getting Started checklist on Live that reads its three steps
from the engine and ticks them itself (recognition ready, camera added
and connected, camera proven by a walk-past), with the one button for
the next step, and that disappears the moment somebody is recognised;
and the model download reported as a percentage in the tray, the
sidebar and the checklist, parsed by the supervisor from the engine's
own progress lines.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The add-camera form asked for an IP address, and a shop owner does not
know their camera's IP address - it is on a sticker under the camera or
in a menu that differs by make. That field is where onboarding stopped
for anyone who was not an installer.
behavision/discover.py: one ONVIF WS-Discovery multicast (names the
camera and often its make) merged with a TCP sweep of port 554 across
the local /24 (misses nothing that streams). Stdlib only, ~4 s on the
office network, both cameras found. The add-camera sheet leads with
'Find cameras on this network'; picking a row fills the address and,
when the make is recognisable, the stream path.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The display name was always meant to be editable and the slug frozen;
until now neither had a way in. PATCH /api/sites/{site} takes a name
and a timezone (manager and above), DELETE removes an empty shop
(owner). The shop drawer in head office gets both, with the short name
shown read-only and the reason beside it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Seen on the demo PC: Start did nothing and Stop stayed grey. The
supervisor's engine had failed because a second engine already held
port 8010, and the tray reported that as nothing at all. The supervisor
now keeps the engine's last lines and turns the known ones into a
sentence - 'port 8010 is already in use - another Behavision or its
engine is still running', 'run behavision-setup again' - which the tray
and the window show. Tray clicks no longer run on the menu loop, so a
stop that waits for the process cannot make the menu look dead.
Two ways that second process came to exist are closed: setup refuses to
run while Behavision.exe or the agent is up, and the app watches
agent.json so a claim made underneath it - which rotates the API token
- is picked up instead of leaving camera sync refused until a restart.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Seen on the first claimed demo install: 'session expired' on every
screen, signed in as a user from the previous demo's head office, and
'Watching 3 cameras' for a shop with one - the PC had offered its two
leftover cameras up to head office, without their passwords, so the
same lens was listed twice and one copy could never be pushed anywhere.
Claiming now clears any stored session (a new head office is a new
world), a session whose refresh fails is forgotten on disk as well as
in memory so the app returns to Login by itself, and the demo setup
removes cameras left from an earlier install before it joins the shop,
because head office is the source of truth from then on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
A plain go build of a Wails app starts, shows 'Wails applications will
not build without the correct build tags' and exits. That is what the
first Windows install of v0.4.4-demo saw. -tags desktop,production is
what wails build passes; both build paths pass it now.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The first demo sealed the office cameras into the package and ran the
PC standalone - a copy of the product with no head office. The bundle
can now carry an installation code instead: setup redeems it exactly as
the app's Setup screen does, the PC joins the shop, and its cameras
arrive from head office on the first sync. The demo then IS the product
- login, Loya, head office - not a local imitation of it. The code is
single-use, so one bundle is one install. release.sh ships the bundle
with DEMO_PACK=.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
A name, a voice and a door. The prompt now asks for a colleague on the
shop floor - answer first, one to three sentences, the shop's name and
the person's name, the one thing to do next - instead of a report with
headings. Both apps put her behind the Loyaly mark in the top-right
corner of every screen, because a buddy you have to find in a sidebar
is not around.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The key-needs-a-workspace error arrived as 'must include the
anthropic-workspace-id header' and was reported as a bare 500 instead of
503 assistant_misconfigured naming the variable. Match the header name,
not the sentence around it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Ask Behavision: a panel beside any screen that talks to the head-office
assistant as the signed-in user - setup questions and 'is my shop
working' answered by the same thing, without leaving the app. The
assistant's prompt now knows how the product is set up (installation
codes, adding a camera, what a placement verdict means, the model
download on first run), so it is the help and not only the analyst. A
PC running on its own has nobody to ask and gets the essentials as text.
Cameras and Customers were still on the pre-redesign markup - the add
camera drawer ran off the right edge of the window because it used a
class the new stylesheet never sized. Both are rebuilt: cameras as
picture-led cards with connection and 'proven' as two separate claims
and a placement check laid out as the two steps it is; the customer
record as a proper sheet.
mock.js renders the app in a browser with fake bindings
(?mock=fresh|standalone|claimed, dev server only), so a screen can be
put in front of somebody without a Windows build. It is how these were
reviewed.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Brand assets in brand/ (the 512px mark, sizes for each surface, a
multi-size .ico). Windows executables carry it as a compiled-in
resource (rsrc_windows_amd64.syso from go-winres) so Explorer, the
taskbar and the installer show it; installer/build.ps1 therefore uses a
plain go build rather than wails build, which would add a second copy
and fail the link. The tray icon is the mark with a state dot over its
corner - a plain coloured circle read as a generic status light among
other icons - rendered from the embedded PNG at 32px so it survives
150% scaling. The desktop app's login, setup and sidebar marks, the
head-office web app's mark and favicon, and the engine dashboard's
favicon are the same file.
Also found while packaging: no wheel so far shipped static/, so the
engine's own dashboard at :8010 on a Windows source install would have
failed with a missing file. package-data now includes it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The previous releases were assembled by hand. This builds the Windows
zip from a clean tree - desktop app, agent, setup tool cross-compiled
here, the engine as a pure-Python wheel with its source beside it - tags,
and publishes to Gitea with notes from a reviewed file.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj