Audited the engine for what it does when something goes wrong rather than
when it goes right. Each of these left the process healthy, the dashboard
green and the product not working.
A gallery the running encoder cannot read. Embeddings are model-tagged, so
when the fallback chain fires every vector the previous encoder wrote goes
invisible: the shop keeps its customer list and recognises nobody on it,
enrolling each regular a second time. Footfall stays correct, which is why
nothing looks wrong. The only evidence was an INFO line reading 'gallery
ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence
that states the disaster and calls it ready. Gallery.health now warns with
the count of PEOPLE lost, not vectors, and carries the same numbers to
/api/stats and /api/health, because a log line on a shop PC is read by
nobody. Proved against the real 87-embedding gallery.
Connected, and sending nothing. 'connected' meant the socket opened, so a
stream that went quiet kept it true while last_frame_age_s climbed and the
heartbeat told head office the camera was up. OpenCV breaks a blocked read
at 30s, but a camera trickling a frame every 20s never trips that and never
recovers. streaming/stalled are reported beside connected and the dashboard
says live/stalled/offline - three states because offline sends you to the
network and stalled says the camera is answering and sending nothing.
The 5-second RTSP timeout that never existed. stimeout;5000000 carried a
comment claiming it bounded a dead camera. Measured on OpenCV 4.11 /
FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with
stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so
it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither
reaches the RTSP protocol through this path; the real bound is OpenCV's own
interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source
already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause.
That matters beyond speed - the VideoCapture constructor is not
interruptible, so stop() could not cut it short and a camera removed from
head office left a daemon thread holding a socket for half a minute.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Measured rather than guessed, and the first guess was wrong. Wall clock
said H.265 decode cost 58 ms a frame; cap.read() blocks until the next
frame arrives, so that was the frame interval, not work. As CPU time:
decode 3.7 ms, detection 31.0 ms - and detection ran on every frame
whether or not anything was in front of the camera, 6,649 of 8,634
frames with faces_seen 0 and active_tracks 0 throughout.
detect_threads: OpenCV spreads a small repeated job over eight threads,
costing 31.0 ms of CPU for 8.9 ms of wall. One thread costs 15.3 ms for
15.3 ms, against a 66 ms budget at 15 fps. Half the CPU for latency
nothing can notice.
motion_gate: a 160x90 greyscale absdiff, 0.1 ms against detection's 15.
Consulted only while no track is open; forced to look every
motion_max_skip frames; compared against the last frame SEARCHED so a
slow drift cannot creep under the threshold; and a threshold above this
camera's measured noise and far below a person, so anything ambiguous
detects. tests/test_motion_gate.py pins each of those rather than the
saving, including asserting the longest run of skips rather than the
total - counting the total would pass a gate that slept forty frames
and then looked forty times.
Together 80% -> 16% of a core, detection skipped on 92% of frames.
faces_seen is still 0 and the gate is not why: run directly over the
same frames the detector finds nothing at threshold 0.50 either. The
placement is the limit, as recorded; the CPU was being spent to
rediscover that fifteen times a second.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Production had 1,211 events accepted and six recognised customers from
the office cameras - the first time the whole chain has carried a real
person, and the project had never been able to claim it. Repeat
sightings score 0.44-0.72, a distribution the match threshold sits
clearly below, on the head-height camera this file has recommended since
August. fraction_below_gate is still 0.59, so the visit count is a floor
and the report says so beside it.
The face-image chain was exercised on production as a shop PC does it -
upload URL, PUT to object storage, anonymous read refused 403. Every
server link holds; the only reason a customer has no photo is
app.store_faces being false by default, which is a data-protection
decision rather than a gap.
Sixteen mobile-API checks pass as a staff account. Three apparent bugs
were test errors and are written down so nobody re-files them, along
with the one field name a caller could guess wrong (site_token).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The last step of onboarding that needed a shell: provision site printed
a broker password and a person typed it into Mosquitto's passwd file on
the host - mounted read-only in the container, so the first attempt
failed silently and the password was re-rolled. No tenant could open a
second branch without us.
The server now drives Mosquitto's dynamic-security plugin over its own
broker login: POST /api/sites (owner) writes the row and the sealed
password, registers the login and a per-site role with literal topics
(the 2.0 plugin does not substitute %u - measured), and removes the row
again if the broker refuses, so a shop cannot exist in the database and
not on the broker. provision site goes through the same path. The
head-office Shops screen gets 'Open a new shop'.
broker-init converts the existing passwd file into the plugin's store
with every hash intact - PBKDF2-SHA512 both sides - so the cutover
re-claims no shop PC. Rehearsed locally: old logins keep working,
isolation holds, the health probe works, and a PC claiming a shop opened
through the API connects as that shop. run-local.sh now brings the
broker up the same way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
012 turned three descriptive columns into identifiers other systems
keep: in agent.json on a shop counter, in a saved URL, in a scheduled
report. All three were already treated as stable and none of it was
enforced.
- clients.slug is an MQTT topic segment the broker ACL is written
against. Rename one and that tenant's whole estate is silently refused
by the broker, with no way to tell the agents.
- sites.slug is what a shop PC calls itself - agent.json holds
"site_id": "chennai", never the uuid. A rename orphans the PC from the
shop it is standing in.
- site_cameras.camera_id lands in visits.camera_id, which is text and
not a foreign key. A rename orphans every visit already attributed to
the old name: the footfall is still there and no longer joins to a
camera. This was half-enforced in handleUpdateCamera and nowhere else,
which is the shape of a rule that holds until somebody adds a second
write path.
- visitors.number is assigned once from the tenant's counter and read
back as V-42.
A trigger, not a CHECK: a CHECK cannot see the old row and the rule is
about the transition. The DISPLAY name is deliberately not frozen -
"TeNext Chennai", "Front door" - it is what a person reads, nothing keys
on it, and a system that cannot fix a typo in a shop's name has confused
the two.
Also records why the uuid stays where a slug would do. The length was
never the problem; needing it was, and that is fixed. Replacing it would
touch eight foreign keys on a live database to shorten a field clients
are already told not to use, and a sequential id would make any future
tenancy hole walkable by counting. It is NOT because ids must be minted
offline - sites, visitors and visits are all created server-side with a
database in hand, and claiming otherwise would defend the status quo
rather than explain it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Asked of the row the feed actually returns.
site_id had a reference all along and the feed was not sending it. A
client could read the shop's NAME off an arrival and still had no way to
ask for that shop except by uuid - the exact gap the reference scheme
exists to close. site_slug now travels with it.
visit_id stays a uuid and needs no reference: no route takes it, it is a
key a client de-duplicates on because delivery is at-least-once, and
nobody says a visit id out loud.
The uuid in a face URL must STAY random. visit_faces.id is
gen_random_uuid() and a derived or sequential one would let somebody
walk a shop's customers by date - the same reason bucket keys are random
rather than derived from the event id. A readable identifier is right
for a customer and wrong for the thing that points at their photograph.
And seq is now json:"-". visits.seq is a plain bigserial, so it counts
every visit on the PLATFORM, and shipping it put the total footfall of
every customer we have on every row of every tenant's feed - the same
German-tank estimate that decided visitors.number had to be per client.
It was a convenience for "have I fallen behind", nothing ever read it,
and the cursor answers that without disclosing a number. The SSE event
id was never the raw value; it has always been the opaque cursor.
The one test that broke was reading seq back off the wire to assert the
cursor pointed at the last row of a burst. It asserts against the seeded
position now: the property is unchanged, and the test can no longer see
what a client cannot.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Every id in the schema is a uuid and stays one. What was wrong was
putting one in front of a person: RecordVisit named every new customer
'Visitor ' || left(id::text, 8), so the arrivals feed, the shop PC and
the mobile app all read "Visitor 3446ec35" - the string a shop assistant
reads to a colleague and types into a search box. label is a stored
column staff can overwrite and SearchVisitors matches on, so formatting
around it in a front end would have left the data wrong on three
surfaces.
Migration 012 adds a per-client visitors.number, taken from a counter on
clients with UPDATE ... RETURNING inside the visit transaction. Per
client rather than global: a global sequence would tell any customer who
signs up how many people the whole platform has ever seen, from their
own first visitor number. The backfill numbers existing rows by
first_seen_at and relabels only the eight-hex pattern the old statement
produced, so a human-typed name is never overwritten.
Three of the four things anyone addresses by URL already had a human
name and the API simply refused it - a site has a slug, a camera has the
id the engine knows it by. refs.go accepts either form anywhere an id is
taken; a uuid resolves with no lookup, so every URL a client already
stored keeps working.
- An ambiguous camera name resolves to nothing, never to a guess: two
shops may each have an "Office1" and acting on the first row would
edit the wrong shop's camera.
- 404 on a path, 400 on a query filter. /api/visits answered fine and it
was the filter that was wrong.
- site and site_id are both accepted everywhere now. They differed per
endpoint, and an unknown query parameter is silently ignored, so
getting it the wrong way round returned the whole estate.
- The search matches V-13, which is what the product now shows.
Two bugs found by running it rather than testing it:
- 'Visitor ' || $2::text beside number = $2 makes Postgres deduce two
types for one parameter and refuse the insert. It compiled and passed
every in-memory test; the first real database rejected it, along with
the existing face tests that share the path.
- The fallback avatar said "V1" for Visitor 13, Visitor 10 and Visitor
15 alike, and read as the V-1 reference for a fourth person. It shows
the number now. The prop is customerRef, not ref - React reserves
that name and it would never have arrived.
Verified on the live database and through the running API: 13 hex labels
became Visitor 1-13 in first-seen order, two typed names left alone, and
the same customer reachable by uuid, V-13 and 13.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
A tenant had exactly the users somebody had created with a command on the
server. That is not a missing screen: a shop with an owner and four staff
either shared one password or raised a ticket per person, and a phone app
for the shop floor could not exist while there was one account to sign in
as.
Registration is by invitation, never open signup - the same line already
drawn around creating a company. The code carries the address and the role
and the request carries only a password, so a code that gets forwarded
cannot become somebody else's account, and a staff invitation cannot be
redeemed as an owner. Single use lives in the UPDATE and the account is
created in the same transaction.
Deactivating a member revokes their sessions in that transaction too. An
access token lives twelve hours, so without it "remove their access"
removed it sometime tomorrow. The session list and revoke that go with it
are the benefit of opaque tokens the product had been paying for and never
collecting: nothing could say what was signed in, let alone stop one.
Face images now work on a deployment with no object storage, which was
every local install and every self-hosted site - the arrivals feed said
"not storing customer photos" for every customer forever, on the screen
whose whole job is to show a face. Bounded to one row per visitor, so it
grows with the customer base and not with footfall; the bucket stays
primary wherever one exists.
Image.auth says whether a URL needs the session, because a browser img
cannot load one that does, a mobile image view can, and a webview can do
neither - the desktop client resolves those to a data URI in Go.
Found by running it, not by tests:
* UPDATE ... RETURNING gives the value AFTER the update, so the prune
read back empty keys, deleted nothing, and the table grew with
footfall exactly as if it were not there. The fake agreed with either
version; only the live Postgres test caught it.
* Trusting only the auth flag broke every shop card, because Sites.jsx
rebuilt a partial snapshot object and dropped it. A relative URL is
now sufficient on its own.
* ago() renders a future time as "just now", so a code valid for a week
read "expires just now".
Verified live against real Postgres: invite, preview, escalation refused,
register into a session, replay 404, staff forbidden, device revoked and
401 at once, last owner refused, and a 92,405-byte camera JPEG stored,
served to its owner, 401 with no session, 404 to another tenant, and
rendered in a browser.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
4 fps was not "live", and it was a number I picked rather than measured.
The engine actually produces ~12 distinct frames a second, so most of it
was being left on the floor.
Now: poll a little ahead of the engine and drop frames identical to the
last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps,
20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a
picture the viewer has not seen, and the rate follows the camera instead
of a constant.
Also records why this is MJPEG rather than passing the camera's own
compressed video through, which would be smoother, cheaper and use no
CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and
BOTH are H.265, despite stream paths ending in ".264". Browsers play
H.264 everywhere and H.265 only on some platforms, so passthrough cannot
rely on it, and transcoding HEVC on the shop PC would put a video encoder
on the machine already doing the recognition.
So probe_source now reports `codec`. It decides what is possible, an
installer can usually change it, and otherwise the only way to learn it is
to read RTSP by hand - which is how this was found.
The RTSP libraries used to establish that are NOT kept: they were only
ever imported by a spike test, and two large dependencies in a shipped
binary to answer a question OpenCV already knows is a bad trade. Their
`go get` had also silently bumped the agent to go 1.25 and broken the
desktop build, which is its own argument.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.
The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.
Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.
Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.
LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.
Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.
Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Both found by operating the stack rather than writing it: the local
processes were OOM-killed and bringing them back hit two gaps.
The headless agent had no way to be claimed at all. Bootstrap lived only
in desktop/internal/cloud, so the one configuration the agent binary
exists for - a back-office PC with no window - could only be onboarded by
hand-editing agent.json, which is the state the desktop's Setup screen was
built to end. `behavision-agent claim <code>` closes it; the CLI joins its
arguments because the code is printed in groups for reading aloud and an
operator pasting it will paste the spaces too.
Second: after the site's broker password was re-rolled, mosquitto logged
"not authorised" while the agent logged "timed out". Those need opposite
actions - re-link this PC, or go and look at the network - and paho's
SetConnectRetry collapses them, because it retries internally and the
connect token never completes. describeStall asks whether a TCP socket
opens at all, and says what is known rather than guessing at a reason the
broker never gives.
Verified end to end: minted a code from the platform as the owner,
claimed with the new command, broker connected, and the shop went to
online: true with 1/1 cameras on w600k_r50.
Also corrects this machine's memory in CLAUDE.md from 16 GB to 8 GB. It
feeds the model-fallback reasoning, and the local gallery already holds
17 embeddings tagged w600k_mbf beside 19 tagged w600k_r50 - the fallback
has silently fired before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Head office shows a camera's latest frame rather than live video, for a
reason that has not changed: the engine serves MJPEG on 127.0.0.1 on a PC
behind a shop's router with no inbound route, and relaying it needs
WebRTC/TURN. Pointing a browser straight at the shop PC is not the escape
either - the engine's API is Basic-authenticated with a credential it
generates locally and never sends anywhere, and shipping that to the
cloud so a web page could use it would put the key to the biometric API
and the live face feed in the server's database.
But that picture only worked if you had an S3 bucket. Without one,
attachSnapshots reported "This system is not storing images" for every
camera forever - on the two screens whose whole job is to show the
camera. Making them picture-led turned a missing feature into a wall of
empty tiles, on every local install and any self-hosted customer who does
not want a bucket.
migrations/009 adds camera_snapshots and the agent falls back to
PUT /api/agent/cameras/{camera}/snapshot when the presigned route answers
images_disabled - chosen by sentinel, never by matching the message, since
it picks between two routes. One row per camera is what makes this safe in
the database when face images are not: the key IS the camera, so storage
is (cameras x ~100 KB) and does not grow with footfall.
The read is session-authenticated rather than a signed link, which an
<img> cannot use - hence Shot.jsx and useAuthedImage, keyed on the URL
string rather than the snapshot object so a poll does not re-fetch 90 KB
per camera every few seconds, and revoking the object URL on cleanup.
Verified against the real office camera with no bucket configured: 90,587
bytes stored in Postgres, served as image/jpeg to a signed-in user, 401
without a session, rendered on both the Cameras and Shops cards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Migrations were run by hand and nothing recorded which had run, so
re-running the setup script against an existing database failed on the
first CREATE TABLE, and shipping a new migration gave an operator no way
to know whether an estate had it. A missed migration is not a startup
error - it is a query referencing a column that is not there, surfacing
later on whichever endpoint touches it first.
server/internal/migrate applies pending migrations at boot and refuses to
start against a schema it does not match. One transaction per file
holding both the DDL and the row that records it; an advisory lock so two
servers starting at once cannot both apply 008; checksums so an edited
migration is refused by name rather than silently skipped; numeric
ordering so 010 does not run before 009. `migrate -baseline N` adopts a
database built before any of this existed, because "the clients table
exists" does not say whether 007's index does.
Verified on the live database: adopted 001-007, applied 008.
008 adds two indexes on `purchases`, found by asking the database which
foreign keys had nothing behind them and then checking what queries the
table. The conversion report filters client_id + occurred_at, which is
exactly the estate-wide case with no site to narrow it.
run-local.sh had two bugs, both found by running it rather than reading
it: it reused a broker container whose bind mount pointed at a directory
that no longer existed, and it discarded stderr on the mosquitto_passwd
call, so under `set -e` it exited at step 5 with no output at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn