Asked after the mobile-internet question: could a VPN let the office cameras
be shown in the demo. Three jobs get confused there and only one needs one.
Seeing the estate from anywhere already works and needs nothing - viewer mode
plus LiveHub is exactly that, outbound, no installation and no credential.
Demonstrating recognition is better done on the demo machine's own camera.
`webcam: 0` picks a capture index instead of building an RTSP URL and the
engine has supported it since the first version: CameraStore round-trips it,
source() returns the index, safe_url() reports webcam:0, and
POST /api/cameras {"id":"laptop","webcam":0} has always worked. No screen
offered it - the same gap this repo already records for the customer record
and per-camera tuning. It is now an option in the make picker, and it is the
strongest demo available: real faces, in the room, depending on no network.
A demo pointed at a camera in another building depends on two internet
connections and a tunnel staying up while somebody is talking.
The address and the index are alternatives, not extras: source() takes the
webcam first, so a half-typed host left behind would make the saved camera
describe two things and use one. The scan, the path and the camera password
are hidden for a local camera because none of them mean anything.
A tunnel is still right for one case - running the engine on a remote machine
against the office's own cameras - and still wrong for the product: it is
per-site infrastructure on every shop PC, and it gives head office
network-level access into a customer's LAN, where today we can read a
camera's picture and nothing else.
Also: installing httpx took the suite from 239 passed to 271. The HTTP tests
importorskip it so a bare checkout runs, which means the number at the bottom
of a run is not the number of tests that exist. Added to the dev extra.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Asked directly by the owner, about his own cameras, from his phone's
connection. The answer is physics and the product was not giving it.
A camera lives on the shop's LAN behind a router. 192.168.1.121 means
"something on the network I am attached to" and nothing more - from mobile
data, a hotel or head office it resolves to nobody, or to a completely
different device holding that number. There is no route in from the internet
and there must not be: an RTSP camera reachable from outside is how a shop's
cameras end up being watched by strangers.
That is why the product is split the way it is - the shop PC is the only
machine on the camera's LAN, and every other surface reaches it outbound,
which is what makes Watch live work from anywhere while nothing connects in.
What was wrong is the message. "cannot reach 192.168.1.121:554 - Operation
timed out" reads as a broken camera and sends somebody to re-type an address
and a password that were always correct. _wrong_network_hint names the cause
and separates two states that need opposite actions:
on that network -> check the camera is powered on and the address is right
somewhere else -> the COMPUTER is in the wrong place; no setting fixes it
- The local address comes from a connected UDP socket that sends nothing. It
only fixes a route so the kernel will name the source address.
- The LAN ranges are spelled out, not is_private. That property also covers
carrier-grade NAT and the documentation networks, and telling somebody who
typed 203.0.113.9 that it is "on the shop's own network" is a confident
wrong answer in the place people look first. Found by a test using that
address as its example of a PUBLIC one.
- A DNS name gets no hint: nothing can be concluded about camera.local from
the string, and guessing is the failure mode this message exists to fix.
- With no network at all it still names the cause and drops the comparison,
rather than claiming to know which network this machine is on.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The certificate fix shipped and failed on the machine it was written for, with
the exact traceback it was meant to prevent. The retry was written
except ssl.SSLCertVerificationError:
and urllib never raises that from urlopen. It catches it and re-raises
urllib.error.URLError(err), carrying the original on .reason. So the except
matched nothing, ever, and the fallback could not fire.
The unit test passed throughout, because the stub it used raised the bare SSL
error - a shape real urllib never produces. That is the lesson: a fake that
agrees with the author is worse than no test, because it converts an untested
path into a tested-looking one. This file already says that about
UPDATE ... RETURNING and about the in-memory API fake, and it got written
again anyway.
_is_cert_failure checks the exception and its .reason, and the tests now raise
URLError(SSLCertVerificationError(...)) - what the traceback actually shows. A
plain URLError is re-raised untouched, and a test asserts no second attempt is
made for one.
Beside the stubs there is now a real reproduction, opt-in behind
BEHAVISION_NETWORK_TESTS=1. Python's default context honours SSL_CERT_FILE, so
an empty file gives a context that trusts nobody - the python.org condition
exactly - while certifi is loaded by path and is unaffected. It skips rather
than passes where it cannot reproduce that, and the difference is measured:
macOS Command Line Tools LibreSSL 2.8.3 128 CAs with an empty CA file
python.org / pyenv build OpenSSL 3.5.8 0 CAs -> reproduces it
Checked for teeth by putting the shipped except back: both the corrected stub
test and the live one fail, and pass again when it is restored.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The SSL fix shipped and did not reach the machine it was written for. That
log said so, one line above the tick:
behavision is already installed with the same version as the provided
wheel. Use --force-reinstall to force an installation of the wheel.
[ok] Engine and dependencies installed
and the traceback below it still pointed at model_assets.py line 59,
urllib.request.urlretrieve - code the fix had deleted.
Two frozen literals caused it: version = "1.1.0" in pyproject.toml and
__version__ = "1.0.0" in behavision/__init__.py. They disagreed with each
other and neither tracked a release, so every release built
behavision-1.1.0-py3-none-any.whl and pip install --upgrade on a machine that
already had 1.1.0 is a no-op. The comment beside that call claimed the
opposite.
The shape of the damage is what makes it bad. The Go binaries - app, agent,
setup tool - are rebuilt every release and updated normally. So a shop PC ran
a current app supervising an engine several releases old, and nothing said
which: /api/health reported the model, the paths, the cameras and the gallery,
and no version at all.
- One version, in the package, read by pyproject through
[tool.setuptools.dynamic]. In a checkout it reads 0.0.0+dev: a
plausible-looking number on a developer's /api/health is worse than none.
- pip install --force-reinstall --no-deps <wheel>, after the ordinary
--upgrade. --upgrade settles the dependencies; the second call guarantees our
own code is the code in the folder. --no-deps keeps it cheap - forcing the
dependencies too would re-download ~300 MB every run. A rebuild at an
unchanged version is the ordinary case while developing, so this must not
rely on the version moving.
- release.sh stamps the tag: v0.5.6-demo -> 0.5.6+demo, valid PEP 440. Into a
copy of the line, reverted in a trap, so the tree is never left dirty.
/api/health reports version now. Without it there is no way to tell a shop PC
three releases behind from a current one, which is how this survived several
releases.
Reproduced end to end before fixing, against real wheels on Python 3.12: two
builds of the same version, --upgrade leaves the old code in place and prints
the same sentence the colleague's Mac printed, --force-reinstall --no-deps
replaces it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
With the version ceiling and the widened numpy pin in place, setup succeeded
on the Mac that found them - Python 3.14 chosen and accepted, numpy 2.5.3,
onnxruntime 1.30, faiss 1.15.1, the engine itself - and died on the last step,
fetching the YuNet model:
ssl.SSLCertVerificationError: [SSL: CERTIFICATE_VERIFY_FAILED]
certificate verify failed: unable to get local issuer certificate
A python.org macOS build ships its own OpenSSL with NO trust store, and
populates one only when somebody double-clicks Install Certificates.command in
the Python folder. Nobody installing face-recognition software has a reason to
know that exists, and the failure is forty lines of traceback about _ssl.c at
the end of a ten-minute install.
_urlopen tries the default context first and retries with certifi's bundle on
a verification failure. The order is the design:
- Default first, because on Windows and on a system or Homebrew Python the
default context reads the machine's own certificate store, which is what
makes a corporate proxy with its own root CA work. Replacing it
unconditionally would break every site that has one to fix a different
platform.
- certifi second, because it is already installed: requests is a hard
dependency and brings it.
- URLError is re-raised untouched. "No route to host" and "no trust store" are
different problems, and retrying the first with a different CA list only
delays the real message.
urlretrieve had to go, since it offers no way to pass a context - exactly the
kind of rewrite that silently drops something. The `download: <label> <n>%`
lines are a contract: supervisor.go's progressRe parses them to put first-run
progress in the tray, because the API is not up yet and a shop PC showing a
stopped engine for five minutes looks broken. A test asserts them, and the
rewritten fetch was checked against the real URL: 232,589 bytes, sha256
identical to the model already on disk.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
makeVenv reused any environment already on disk, whatever Python built it.
The machine that found the version bug already had a runtime built by 3.14,
left there by the run that failed - so with the ceiling in place setup would
choose a good interpreter, reach makeVenv, find the 3.14 environment, keep it,
and die in the same clang error as before.
A fix a user cannot reach because the bug's own debris is in the way is not a
fix, and it would have read as the release not working.
It now asks the interpreter inside an existing environment what it is and
rebuilds when the answer is unsupported, saying so. Rebuilding costs a
re-download of the libraries and nothing else - the models live in the state
root. An environment that cannot be asked counts as unusable too: a
half-created one answers nothing, and reusing it fails later in pip with an
error about a package rather than about the environment.
Tested against real environments rather than a fake, because what is under
test is what an interpreter on disk reports about itself.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Three reported from a colleague's machine, plus one the fixing uncovered.
Every one produced a message that was true and useless.
## behavision-setup chose the Python least likely to work
findPython walked 3.14, 3.13, 3.12, 3.11, 3.10 and took the first hit - a
floor with NO ceiling, which is exactly backwards. The newest Python on a
machine is the one least likely to have binary wheels. It picked 3.14, pip
found no numpy wheel for cp314, fell back to building numpy from source and
produced "Unknown compiler(s)"; once the operator had installed Xcode's
command line tools to get past that, ten minutes of compiling ended in
"<arm_neon.h> is intended only for ARM and AArch64 targets".
maxMinor refuses in one line before anything is downloaded, and "too new" is
a different message from "too old" - telling somebody holding Python 3.14
that no Python was found sends them to install a newer one, which is the
direction that just failed.
## numpy<2.0 was the cap; OpenCV was the hazard
Widening it needed proof, and the proof found something else. Nine runs of
the detector guard per combination, one machine, one sitting:
numpy 1.26 / cv2 4.11 9 passed, 0 crashed
numpy 2.0 / cv2 4.11 8 passed, 1 crashed
numpy 1.26 / cv2 4.14 3 passed, 6 crashed
numpy 2.0 / cv2 4.14 2 passed, 7 crashed
numpy is not the variable; OpenCV is - the third row is numpy 1.26. The crash
was test_a_shared_detector_really_does_race, which races a shared
cv2.FaceDetectorYN on purpose. That is undefined behaviour in C++: 4.11
usually turned it into an exception, 4.14 usually turns it into a segfault,
and 4.11 crashing once says the hazard was always there.
It never reached the product - Engine._build_worker builds a detector per
camera. It reached the suite: two runs in three died with no failing
assertion in them. The race runs in a subprocess now, and one clean attempt
proves nothing, so the premise holds if any of several attempts misbehaves.
226 passed / 2 skipped on numpy 2.0.2, five runs of five.
opencv stays capped below 5: everything above was measured on 4.x, and an
uncapped >=4.8.1 gives every NEW install a major release this project has
never run a real camera through.
## One MQTT client id for a whole shop, so two PCs fought over it
behavision-<client>-<site> is the same string on every computer claimed to one
site. MQTT requires unique client ids and a broker enforces it by
disconnecting the older session, so the colleague's Mac and the shop's own
till took turns kicking each other off:
broker connected / broker connection lost: EOF / broker connected / EOF ...
The damage is not confined to the new machine. The till is the other half of
that loop, so signing in on a laptop to look at the product stops a live shop
delivering visits - and from each end it reads as an unstable network.
MQTTClientID() appends a per-installation id, minted on first load and written
back so an existing install gets one without anybody doing anything. The site
stays in the name because that is what a broker log is read by. An unwritable
config falls back to a per-run id rather than a shared one.
## "no such file or directory" for an engine nobody had installed
Pressing Start went straight to the supervisor, which reported what exec
reported: a 200-character path ending in "no such file or directory". Every
word true, none of it saying "run the setup tool" - the startup path had that
sentence, in a log file nobody on a shop counter opens.
engineMissing() is the one function the startup path, the Start button and the
status panel all consult. It also names App Translocation, which was in that
path and is unguessable: macOS runs a downloaded unsigned app from a random
read-only copy, so relative paths resolve inside it and an install there would
not survive a restart. The product is unsigned, so that is the normal
first-run state on every Mac, not an edge case.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Two changes, and the second was found by verifying the first.
## Watching a camera from the app, in another building
Snapshots answer "is that camera working". They do not answer "what is
happening in my shop right now", which is what somebody who opens the app away
from the counter is asking. Head office's browser already had that answer -
LiveHub plus cameras.Live, where the shop PC asks outbound whether anybody is
watching and pushes JPEG frames for as long as somebody is - and the app could
not reach it.
cloud.CameraLive opens that feed and the app's own loopback relay re-emits it
as multipart MJPEG. That is the trick: frames arrive base64 over SSE, an <img>
cannot render that, and an <img> renders MJPEG natively - so a tile is an
ordinary <img> pointed at loopback whether the camera is in this room or
another city.
- Reconnecting happens in the relay, not the page. The server caps one push at
five minutes, so doing it here means the <img> never sees the stream end.
- The headers are flushed before the first frame. Go writes them on the first
body write, so without that the whole response waits for the shop PC to
start pushing. Measured against production: 30 seconds and not even a
Content-Type, which surfaces as the request timing out.
- One camera at a time. Watching makes a shop PC upload, so a grid that went
live at once would put an estate's worth of cameras on the wire because
somebody opened a page.
- live.mjpeg is behind the same per-run token as the engine routes, and a
wrong token is a 404 that never reaches head office at all.
- CameraLive uses its own HTTP client: the shared one's 30s timeout covers the
whole response and would sever a working view every thirty seconds - the
trap that made the server set WriteTimeout to zero for its own SSE endpoint.
## A camera read "Connected" for 34 minutes after the shop PC went blind
Which is why the verification above looked like a failure: head office
registered the viewer and no frame ever came.
reportWith returns early when the engine is unreachable - correctly, it has
nothing to say - so the last state it sent stays in the database looking
current. Measured live: cam2 and entrance both reading Connected, in green,
with last_seen_at 34 minutes old, while the heartbeat from the same PC said
cameras_up 0 of 0. Two surfaces reading two stored fields and disagreeing.
false could not be the answer. It means "this camera is not connecting", which
sends an installer to check cabling on a camera that was working perfectly the
last time anybody could ask it. So there are four states and one function:
connected reported recently, and working
not_connecting reported recently, and the stream will not open
waiting no shop PC has ever reported this camera
stale reported once, and not lately
- Connected is CLEARED when stale or waiting. A stale true left in place stays
available to every client reading the field directly, and leaves two fields
on one object disagreeing - how the shops screen once came out labelled
Working, in green, above "2 of 3 cameras not connecting".
- Computed in scanCamera, so every camera anybody reads passes through it. A
state computed per handler is one a handler forgets, and this had already
reached three screens.
- CameraStaleAfter is 5 minutes: five missed reports, not one. Same reasoning
as three missed heartbeats - an indicator that cries wolf gets ignored.
- An unparseable last_seen_at is stale. It should be impossible, which is why
it must not fall through to the state that says everything is fine.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Signing in on a second Mac showed "engine not reachable at
http://127.0.0.1:8010" and 0 of 0 cameras, on an account whose shops were
running and recognising people the whole time. Nothing was broken: Live() and
Cameras() read only the engine on loopback, so the app answered as though the
person had never signed in - and camera sync goes through the engine, which is
why the count was zero rather than stale.
Having no engine is a normal state. A shop PC watches cameras; an owner's
laptop, a manager's machine and a second till being set up do not, and all
three are signed in to the same estate. Both methods now fall back to head
office when loopback fails and somebody is signed in. Loopback is still tried
first: a real shop PC must never be shown a minute-old summary when the engine
two milliseconds away has the live one.
Decisions worth keeping:
- Viewing is on the snapshot, not inferred per screen. Three surfaces read it,
and a screen that computed it separately is how the shops screen once came
out labelled Working, in green, above "2 of 3 cameras not connecting".
- fraction_below_gate takes the WORST shop, never an average. 0.10 against
0.73 averages to 0.42 and hides the only shop anyone needs to visit.
- A remote camera is flagged, and Edit, Remove and Check placement are
withheld. They talk to a camera on a LAN this computer cannot reach, and a
button that cannot work is worse than one that is absent.
- connected is three states. null is "no shop computer has reported yet" and
reads as waiting; false is "Not connecting". A bare false sends somebody to
check cabling on a camera nobody has tried to reach.
- Snapshots are fetched in Go as data: URIs and cached by snapshot_at. A
webview <img> resolves a relative src against wails:// and cannot send the
bearer - the problem VisitorImage already solved - and this screen polls
every 8 seconds at ~90 KB a camera.
- With no engine AND no session, the engine error is still the answer. The
person is most likely setting this PC up.
The picture is the last snapshot and the banner says so: there is no live
video from here, because the engine's MJPEG stream is on the shop PC's
loopback behind a router with no inbound route. The LiveHub relay head office
uses is the answer to that and is a further step for this client.
Verified against production: five arrivals and two cameras parsed from the
real API. viewing_test.go covers the fallback, the worst-shop rule, the
withheld credentials and that an unchanged snapshot is fetched once across
two polls.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Audited the engine for what it does when something goes wrong rather than
when it goes right. Each of these left the process healthy, the dashboard
green and the product not working.
A gallery the running encoder cannot read. Embeddings are model-tagged, so
when the fallback chain fires every vector the previous encoder wrote goes
invisible: the shop keeps its customer list and recognises nobody on it,
enrolling each regular a second time. Footfall stays correct, which is why
nothing looks wrong. The only evidence was an INFO line reading 'gallery
ready: 0 embeddings (model w600k_mbf) across 21 identities' - a sentence
that states the disaster and calls it ready. Gallery.health now warns with
the count of PEOPLE lost, not vectors, and carries the same numbers to
/api/stats and /api/health, because a log line on a shop PC is read by
nobody. Proved against the real 87-embedding gallery.
Connected, and sending nothing. 'connected' meant the socket opened, so a
stream that went quiet kept it true while last_frame_age_s climbed and the
heartbeat told head office the camera was up. OpenCV breaks a blocked read
at 30s, but a camera trickling a frame every 20s never trips that and never
recovers. streaming/stalled are reported beside connected and the dashboard
says live/stalled/offline - three states because offline sends you to the
network and stalled says the camera is answering and sending nothing.
The 5-second RTSP timeout that never existed. stimeout;5000000 carried a
comment claiming it bounded a dead camera. Measured on OpenCV 4.11 /
FFmpeg 7.1 against a socket that accepts and then says nothing: 30.0s with
stimeout, 30.0s with timeout, 30.3s with no option at all - identical, so
it was never honoured. stimeout became timeout in FFmpeg 5.0 and neither
reaches the RTSP protocol through this path; the real bound is OpenCV's own
interrupt constant. Replaced by the _tcp_reachable pre-flight probe_source
already used, in code we own: 30.3s -> 0.00-2.02s, each naming its cause.
That matters beyond speed - the VideoCapture constructor is not
interruptible, so stop() could not cut it short and a camera removed from
head office left a daemon thread holding a socket for half a minute.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Measured rather than guessed, and the first guess was wrong. Wall clock
said H.265 decode cost 58 ms a frame; cap.read() blocks until the next
frame arrives, so that was the frame interval, not work. As CPU time:
decode 3.7 ms, detection 31.0 ms - and detection ran on every frame
whether or not anything was in front of the camera, 6,649 of 8,634
frames with faces_seen 0 and active_tracks 0 throughout.
detect_threads: OpenCV spreads a small repeated job over eight threads,
costing 31.0 ms of CPU for 8.9 ms of wall. One thread costs 15.3 ms for
15.3 ms, against a 66 ms budget at 15 fps. Half the CPU for latency
nothing can notice.
motion_gate: a 160x90 greyscale absdiff, 0.1 ms against detection's 15.
Consulted only while no track is open; forced to look every
motion_max_skip frames; compared against the last frame SEARCHED so a
slow drift cannot creep under the threshold; and a threshold above this
camera's measured noise and far below a person, so anything ambiguous
detects. tests/test_motion_gate.py pins each of those rather than the
saving, including asserting the longest run of skips rather than the
total - counting the total would pass a gate that slept forty frames
and then looked forty times.
Together 80% -> 16% of a core, detection skipped on 92% of frames.
faces_seen is still 0 and the gate is not why: run directly over the
same frames the detector finds nothing at threshold 0.50 either. The
placement is the limit, as recorded; the CPU was being spent to
rediscover that fifteen times a second.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Production had 1,211 events accepted and six recognised customers from
the office cameras - the first time the whole chain has carried a real
person, and the project had never been able to claim it. Repeat
sightings score 0.44-0.72, a distribution the match threshold sits
clearly below, on the head-height camera this file has recommended since
August. fraction_below_gate is still 0.59, so the visit count is a floor
and the report says so beside it.
The face-image chain was exercised on production as a shop PC does it -
upload URL, PUT to object storage, anonymous read refused 403. Every
server link holds; the only reason a customer has no photo is
app.store_faces being false by default, which is a data-protection
decision rather than a gap.
Sixteen mobile-API checks pass as a staff account. Three apparent bugs
were test errors and are written down so nobody re-files them, along
with the one field name a caller could guess wrong (site_token).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The last step of onboarding that needed a shell: provision site printed
a broker password and a person typed it into Mosquitto's passwd file on
the host - mounted read-only in the container, so the first attempt
failed silently and the password was re-rolled. No tenant could open a
second branch without us.
The server now drives Mosquitto's dynamic-security plugin over its own
broker login: POST /api/sites (owner) writes the row and the sealed
password, registers the login and a per-site role with literal topics
(the 2.0 plugin does not substitute %u - measured), and removes the row
again if the broker refuses, so a shop cannot exist in the database and
not on the broker. provision site goes through the same path. The
head-office Shops screen gets 'Open a new shop'.
broker-init converts the existing passwd file into the plugin's store
with every hash intact - PBKDF2-SHA512 both sides - so the cutover
re-claims no shop PC. Rehearsed locally: old logins keep working,
isolation holds, the health probe works, and a PC claiming a shop opened
through the API connects as that shop. run-local.sh now brings the
broker up the same way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
012 turned three descriptive columns into identifiers other systems
keep: in agent.json on a shop counter, in a saved URL, in a scheduled
report. All three were already treated as stable and none of it was
enforced.
- clients.slug is an MQTT topic segment the broker ACL is written
against. Rename one and that tenant's whole estate is silently refused
by the broker, with no way to tell the agents.
- sites.slug is what a shop PC calls itself - agent.json holds
"site_id": "chennai", never the uuid. A rename orphans the PC from the
shop it is standing in.
- site_cameras.camera_id lands in visits.camera_id, which is text and
not a foreign key. A rename orphans every visit already attributed to
the old name: the footfall is still there and no longer joins to a
camera. This was half-enforced in handleUpdateCamera and nowhere else,
which is the shape of a rule that holds until somebody adds a second
write path.
- visitors.number is assigned once from the tenant's counter and read
back as V-42.
A trigger, not a CHECK: a CHECK cannot see the old row and the rule is
about the transition. The DISPLAY name is deliberately not frozen -
"TeNext Chennai", "Front door" - it is what a person reads, nothing keys
on it, and a system that cannot fix a typo in a shop's name has confused
the two.
Also records why the uuid stays where a slug would do. The length was
never the problem; needing it was, and that is fixed. Replacing it would
touch eight foreign keys on a live database to shorten a field clients
are already told not to use, and a sequential id would make any future
tenancy hole walkable by counting. It is NOT because ids must be minted
offline - sites, visitors and visits are all created server-side with a
database in hand, and claiming otherwise would defend the status quo
rather than explain it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Asked of the row the feed actually returns.
site_id had a reference all along and the feed was not sending it. A
client could read the shop's NAME off an arrival and still had no way to
ask for that shop except by uuid - the exact gap the reference scheme
exists to close. site_slug now travels with it.
visit_id stays a uuid and needs no reference: no route takes it, it is a
key a client de-duplicates on because delivery is at-least-once, and
nobody says a visit id out loud.
The uuid in a face URL must STAY random. visit_faces.id is
gen_random_uuid() and a derived or sequential one would let somebody
walk a shop's customers by date - the same reason bucket keys are random
rather than derived from the event id. A readable identifier is right
for a customer and wrong for the thing that points at their photograph.
And seq is now json:"-". visits.seq is a plain bigserial, so it counts
every visit on the PLATFORM, and shipping it put the total footfall of
every customer we have on every row of every tenant's feed - the same
German-tank estimate that decided visitors.number had to be per client.
It was a convenience for "have I fallen behind", nothing ever read it,
and the cursor answers that without disclosing a number. The SSE event
id was never the raw value; it has always been the opaque cursor.
The one test that broke was reading seq back off the wire to assert the
cursor pointed at the last row of a burst. It asserts against the seeded
position now: the property is unchanged, and the test can no longer see
what a client cannot.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Every id in the schema is a uuid and stays one. What was wrong was
putting one in front of a person: RecordVisit named every new customer
'Visitor ' || left(id::text, 8), so the arrivals feed, the shop PC and
the mobile app all read "Visitor 3446ec35" - the string a shop assistant
reads to a colleague and types into a search box. label is a stored
column staff can overwrite and SearchVisitors matches on, so formatting
around it in a front end would have left the data wrong on three
surfaces.
Migration 012 adds a per-client visitors.number, taken from a counter on
clients with UPDATE ... RETURNING inside the visit transaction. Per
client rather than global: a global sequence would tell any customer who
signs up how many people the whole platform has ever seen, from their
own first visitor number. The backfill numbers existing rows by
first_seen_at and relabels only the eight-hex pattern the old statement
produced, so a human-typed name is never overwritten.
Three of the four things anyone addresses by URL already had a human
name and the API simply refused it - a site has a slug, a camera has the
id the engine knows it by. refs.go accepts either form anywhere an id is
taken; a uuid resolves with no lookup, so every URL a client already
stored keeps working.
- An ambiguous camera name resolves to nothing, never to a guess: two
shops may each have an "Office1" and acting on the first row would
edit the wrong shop's camera.
- 404 on a path, 400 on a query filter. /api/visits answered fine and it
was the filter that was wrong.
- site and site_id are both accepted everywhere now. They differed per
endpoint, and an unknown query parameter is silently ignored, so
getting it the wrong way round returned the whole estate.
- The search matches V-13, which is what the product now shows.
Two bugs found by running it rather than testing it:
- 'Visitor ' || $2::text beside number = $2 makes Postgres deduce two
types for one parameter and refuse the insert. It compiled and passed
every in-memory test; the first real database rejected it, along with
the existing face tests that share the path.
- The fallback avatar said "V1" for Visitor 13, Visitor 10 and Visitor
15 alike, and read as the V-1 reference for a fourth person. It shows
the number now. The prop is customerRef, not ref - React reserves
that name and it would never have arrived.
Verified on the live database and through the running API: 13 hex labels
became Visitor 1-13 in first-seen order, two typed names left alone, and
the same customer reachable by uuid, V-13 and 13.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
A tenant had exactly the users somebody had created with a command on the
server. That is not a missing screen: a shop with an owner and four staff
either shared one password or raised a ticket per person, and a phone app
for the shop floor could not exist while there was one account to sign in
as.
Registration is by invitation, never open signup - the same line already
drawn around creating a company. The code carries the address and the role
and the request carries only a password, so a code that gets forwarded
cannot become somebody else's account, and a staff invitation cannot be
redeemed as an owner. Single use lives in the UPDATE and the account is
created in the same transaction.
Deactivating a member revokes their sessions in that transaction too. An
access token lives twelve hours, so without it "remove their access"
removed it sometime tomorrow. The session list and revoke that go with it
are the benefit of opaque tokens the product had been paying for and never
collecting: nothing could say what was signed in, let alone stop one.
Face images now work on a deployment with no object storage, which was
every local install and every self-hosted site - the arrivals feed said
"not storing customer photos" for every customer forever, on the screen
whose whole job is to show a face. Bounded to one row per visitor, so it
grows with the customer base and not with footfall; the bucket stays
primary wherever one exists.
Image.auth says whether a URL needs the session, because a browser img
cannot load one that does, a mobile image view can, and a webview can do
neither - the desktop client resolves those to a data URI in Go.
Found by running it, not by tests:
* UPDATE ... RETURNING gives the value AFTER the update, so the prune
read back empty keys, deleted nothing, and the table grew with
footfall exactly as if it were not there. The fake agreed with either
version; only the live Postgres test caught it.
* Trusting only the auth flag broke every shop card, because Sites.jsx
rebuilt a partial snapshot object and dropped it. A relative URL is
now sufficient on its own.
* ago() renders a future time as "just now", so a code valid for a week
read "expires just now".
Verified live against real Postgres: invite, preview, escalation refused,
register into a session, replay 404, staff forbidden, device revoked and
401 at once, last owner refused, and a 92,405-byte camera JPEG stored,
served to its owner, 401 with no session, 404 to another tenant, and
rendered in a browser.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
4 fps was not "live", and it was a number I picked rather than measured.
The engine actually produces ~12 distinct frames a second, so most of it
was being left on the floor.
Now: poll a little ahead of the engine and drop frames identical to the
last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps,
20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a
picture the viewer has not seen, and the rate follows the camera instead
of a constant.
Also records why this is MJPEG rather than passing the camera's own
compressed video through, which would be smoother, cheaper and use no
CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and
BOTH are H.265, despite stream paths ending in ".264". Browsers play
H.264 everywhere and H.265 only on some platforms, so passthrough cannot
rely on it, and transcoding HEVC on the shop PC would put a video encoder
on the machine already doing the recognition.
So probe_source now reports `codec`. It decides what is possible, an
installer can usually change it, and otherwise the only way to learn it is
to read RTSP by hand - which is how this was found.
The RTSP libraries used to establish that are NOT kept: they were only
ever imported by a spike test, and two large dependencies in a shipped
binary to answer a question OpenCV already knows is a bad trade. Their
`go get` had also silently bumped the agent to go 1.25 and broken the
desktop build, which is its own argument.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.
The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.
Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.
Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.
LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.
Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.
Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Both found by operating the stack rather than writing it: the local
processes were OOM-killed and bringing them back hit two gaps.
The headless agent had no way to be claimed at all. Bootstrap lived only
in desktop/internal/cloud, so the one configuration the agent binary
exists for - a back-office PC with no window - could only be onboarded by
hand-editing agent.json, which is the state the desktop's Setup screen was
built to end. `behavision-agent claim <code>` closes it; the CLI joins its
arguments because the code is printed in groups for reading aloud and an
operator pasting it will paste the spaces too.
Second: after the site's broker password was re-rolled, mosquitto logged
"not authorised" while the agent logged "timed out". Those need opposite
actions - re-link this PC, or go and look at the network - and paho's
SetConnectRetry collapses them, because it retries internally and the
connect token never completes. describeStall asks whether a TCP socket
opens at all, and says what is known rather than guessing at a reason the
broker never gives.
Verified end to end: minted a code from the platform as the owner,
claimed with the new command, broker connected, and the shop went to
online: true with 1/1 cameras on w600k_r50.
Also corrects this machine's memory in CLAUDE.md from 16 GB to 8 GB. It
feeds the model-fallback reasoning, and the local gallery already holds
17 embeddings tagged w600k_mbf beside 19 tagged w600k_r50 - the fallback
has silently fired before.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Head office shows a camera's latest frame rather than live video, for a
reason that has not changed: the engine serves MJPEG on 127.0.0.1 on a PC
behind a shop's router with no inbound route, and relaying it needs
WebRTC/TURN. Pointing a browser straight at the shop PC is not the escape
either - the engine's API is Basic-authenticated with a credential it
generates locally and never sends anywhere, and shipping that to the
cloud so a web page could use it would put the key to the biometric API
and the live face feed in the server's database.
But that picture only worked if you had an S3 bucket. Without one,
attachSnapshots reported "This system is not storing images" for every
camera forever - on the two screens whose whole job is to show the
camera. Making them picture-led turned a missing feature into a wall of
empty tiles, on every local install and any self-hosted customer who does
not want a bucket.
migrations/009 adds camera_snapshots and the agent falls back to
PUT /api/agent/cameras/{camera}/snapshot when the presigned route answers
images_disabled - chosen by sentinel, never by matching the message, since
it picks between two routes. One row per camera is what makes this safe in
the database when face images are not: the key IS the camera, so storage
is (cameras x ~100 KB) and does not grow with footfall.
The read is session-authenticated rather than a signed link, which an
<img> cannot use - hence Shot.jsx and useAuthedImage, keyed on the URL
string rather than the snapshot object so a poll does not re-fetch 90 KB
per camera every few seconds, and revoking the object URL on cleanup.
Verified against the real office camera with no bucket configured: 90,587
bytes stored in Postgres, served as image/jpeg to a signed-in user, 401
without a session, rendered on both the Cameras and Shops cards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Migrations were run by hand and nothing recorded which had run, so
re-running the setup script against an existing database failed on the
first CREATE TABLE, and shipping a new migration gave an operator no way
to know whether an estate had it. A missed migration is not a startup
error - it is a query referencing a column that is not there, surfacing
later on whichever endpoint touches it first.
server/internal/migrate applies pending migrations at boot and refuses to
start against a schema it does not match. One transaction per file
holding both the DDL and the row that records it; an advisory lock so two
servers starting at once cannot both apply 008; checksums so an edited
migration is refused by name rather than silently skipped; numeric
ordering so 010 does not run before 009. `migrate -baseline N` adopts a
database built before any of this existed, because "the clients table
exists" does not say whether 007's index does.
Verified on the live database: adopted 001-007, applied 008.
008 adds two indexes on `purchases`, found by asking the database which
foreign keys had nothing behind them and then checking what queries the
table. The conversion report filters client_id + occurred_at, which is
exactly the estate-wide case with no site to narrow it.
run-local.sh had two bugs, both found by running it rather than reading
it: it reused a broker container whose bind mount pointed at a directory
that no longer existed, and it discarded stderr on the mosquitto_passwd
call, so under `set -e` it exited at step 5 with no output at all.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn