Commit Graph

33 Commits

Author SHA1 Message Date
a74cb899b4 server/deploy.sh: build here, back up, migrate, switch, verify
Production ran code from 31 August and answered 404 to most of the API
the merchant and mobile clients are written against. The script builds
the web app into a static linux binary on the developer machine (the
host has 3.6 GB shared with other services and must not compile), backs
the database up, runs the migrations with the new binary while the old
server still serves so a failure stops with nothing changed, switches,
and proves the routes over the public URL. API.md now names the API host
correctly: mcp.loyaly.ai, not the console's platform.loyaly.ai.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-18 11:37:47 +05:30
8c88aad06e Stopping the engine on Windows stops the whole engine
The installer runs the engine as <venv>\Scripts\python.exe, and since
Python 3.7.2 that file is a redirector that spawns the real interpreter
as a child. Stop() terminated the redirector and left the interpreter -
the process holding the cameras and the SQLite WAL - running with no
parent and nothing able to stop it. Seen on a Windows install: Quit from
the tray, and recognition still running.

The child is now started suspended, placed in a job object with
KILL_ON_JOB_CLOSE, and resumed. Cancel terminates the job, so the whole
tree goes; and the job dies with this process, so it goes even if the app
crashes. CREATE_NO_WINDOW while here: python.exe is a console program
and a GUI parent otherwise opens a black console on the shop counter.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-18 11:06:49 +05:30
50a843ce46 The shop app starts once, and a second launch just shows the window
The window hides to the tray on close, so the natural next step for a
shop assistant is to double-click the shortcut again. That started a
second full copy of the app: a second tray icon, a second engine
supervisor on the same SQLite WAL and the same port - the start-twice
failure the agent package was built to prevent, on the one binary that
never had the guard. Seen on a Windows install as a row of tray icons.
Wails' SingleInstanceLock now hands the second launch to the first
process, which brings its window to the front.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-18 11:01:19 +05:30
979aa77cda The shop screen no longer shows camera video
The Live screen led with a camera tile beside the arrivals. Nobody at a
counter is watching CCTV; they are looking up at a customer and need the
name. The tile also cost CPU the recognition pipeline needs and pulled a
stream relay into the app for a picture that was decoration. Arrivals now
take the whole screen. The camera picture stays on the Cameras screen,
where it is a setup tool and not a feed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-18 10:53:32 +05:30
3d3775c8be The shop app looks like a product now, not a prototype
The window a shop assistant stares at all day was the weakest surface in
this system, and it looked improvised because it was: navigation drawn
with text characters (◉ ☺ ▢) that sit on the text baseline and cannot
take a stroke weight, margins set inline per screen, and four large stat
boxes dominating the page while the product's entire reason for existing
- WHO JUST WALKED IN - was a list of "person.seen" rows in the corner.

Rebuilt around the person in front of it: a counter, a cheap monitor,
somebody mid-conversation with a customer.

  - ui/icons.jsx: one drawn icon set, 24-unit grid, 1.6 stroke,
    currentColor, so one icon works on every surface and in every state.
  - styles.css: a real system. Four-step ground→raised palette biased
    blue-green (this product lives in the world of lenses), one spacing
    scale, one type scale, tabular figures wherever digits are compared
    or refreshed in place, and the scrollbars restyled - the default
    light scrollbar on a dark panel is the loudest "web page in a frame"
    tell there is.
  - Live: a status strip that answers "is this working" in one line,
    cameras as pictures with the caption over the image, and arrivals as
    cards big enough to match against the person standing there. The
    four stat boxes became a slim strip at the foot, where numbers that
    nobody acts on belong.
  - State is carried by shape AND colour everywhere - a pill, a dot and
    an edge stripe - because this gets read from two metres away and
    some operators do not see red and green apart.
  - Motion only where it means something: a live camera pulses, a fresh
    arrival slides in once. Nothing loops for decoration; this process
    shares a CPU with recognition.

Two things fixed because the screen showed them, not because a test did:

  - The sidebar read "Stopped" beside a live camera feed and a counter
    ticking up, whenever the engine was running but not started BY the
    app. That is the two-surfaces-disagreeing bug the tray exists to
    avoid. It now reads "Running outside the app" in amber, and Start is
    disabled rather than offering to launch a second engine onto one
    SQLite WAL.
  - The arrivals panel shrank to fit its content and left a hole beside
    a tall camera tile - so the layout looked broken exactly when the
    shop was quiet, which is most of the time. Both panels stretch and
    scroll their own content now.

Every existing class name still resolves, so the screens not rewritten
here pick the system up unchanged. Windows and darwin build; tests pass.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-15 11:06:57 +05:30
a1fe0942e2 docs: the architecture overview, as a file
Nine diagrams, one HTML file, no dependencies beyond web fonts that
fall back to system faces offline. The same document is published as
an artifact; this is the copy that ships with the repository.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-11 17:12:50 +05:30
e262fc8482 The live picture was chained to the recognition pipeline
Reported from the first Windows install: the camera feed lags. It did,
and not because of the network, the proxy or the webview.

The MJPEG stream served _annotated_jpeg - the frame the pipeline had
most recently FINISHED with, encoded after detection, quality scoring,
tracking and identification had all run on it. On a modest shop PC that
is a few frames a second, and every picture was already as old as that
processing. It looked like lag because it was lag. On the fast machine
it was developed on the pipeline kept up with the stream's own 10 fps
cap, which is why nobody here ever saw it.

Two more things compounded it. Every processed frame was JPEG-encoded
whether or not a viewer existed - CPU spent on precisely the machine
short of it. And ffmpeg ran its RTSP demuxer with default buffering,
which holds a comfortable queue of frames before handing over the first:
half a second to two seconds a live view can never recover.

Now the picture and the boxes are decoupled. latest_jpeg_since takes the
capture thread's freshest frame at the camera's own rate and draws the
boxes from the last processed frame over it - encoded on demand, per
request, so a camera nobody watches costs no encode at all. The stream
sends a frame only when the camera has a newer one, capped at 15 fps;
nothing is sent twice. Boxes older than a second are not drawn, so a
stalled pipeline cannot leave one floating over an empty spot.
_publish_annotated becomes _remember_tracks: a handful of tuples under
the lock, no copy, no encode. ffmpeg gets nobuffer / low_delay /
max_delay.

Measured on cam2's sub-stream, same machine, ten seconds each:

  before   99 frames sent,  98 distinct    9.8 new pictures/s
  after   141 frames sent, 141 distinct   14.0 new pictures/s

against a 15 fps camera, with the pipeline still processing 166 of 181
captured frames alongside - and engine CPU DOWN from 90% with no viewer
to 62% with one attached.

Engine version 1.0.0 -> 1.1.0 so a re-run of setup reinstalls it rather
than pip deciding the requirement is already satisfied.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
v0.3.2 v0.4.1-demo
2026-09-11 16:28:05 +05:30
b59e667a68 A demo release with the office cameras sealed inside it
Wanted: install it and the two office cameras are already there - but
without the release carrying their admin password where anyone with the
zip can read it. "Encode it" does not achieve that; anything the
installer can decode, anyone holding the installer can decode.

pkg/demo seals the camera list with AES-256-GCM under a key that is NOT
in the package: a 120-bit unlock code minted when the bundle is sealed,
given to whoever runs setup by voice or message, typed once. The code
is random, so it is key material directly through SHA-256; a human-
chosen passphrase would need a KDF and a dependency, 120 random bits do
not. The sealed file contains the format marker and noise. Tested: the
password and the host do not appear in it, a wrong code and a flipped
byte are both refused as ErrWrongCode, every seal differs.

behavision-demo-pack seals; it runs on the build machine and is never
shipped. The code is printed once and stored nowhere.

behavision-setup, on finding demo-cameras.enc beside the engine source,
asks for the code BEFORE the ten-minute download so a mistyped one costs
seconds, and adds the cameras at the end - through the running engine's
own Add Camera endpoint, not by writing its file. The store's save() is
what applies DPAPI to the password on Windows, so this is how the
credential ends up encrypted and machine-bound on the demo PC rather
than in cameras.json for anyone who can read ProgramData. It then marks
the PC standalone, so the app opens on Live instead of asking for an
installation code it will never get.

Which found the gap that DPAPI only works if pywin32 is importable, and
nothing had ever pulled it in - every Windows install to date would have
logged the warning and written camera passwords in the clear. Added as
a Windows-only dependency.

Verified in a clean container: a wrong code refused, the right one
unlocks two cameras, every install step passes, both cameras added
through the API, standalone set.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
v0.4.0-demo
2026-09-11 16:07:49 +05:30
70c447873d "Session expired" on a screen where nobody had signed in
The first Windows install reached the setup screen, typed an
installation code, and was told the session had expired. There was no
session. The code had been minted on a different head office, and the
server said so - 401 bad_token, "That installation code is not valid.
Ask for a new one." - and the client threw the message away, because it
mapped every 401 to the string "session expired".

A 401 on a call that carried a session is a session problem. A 401 on a
call that carried none is about the request, and the server's message is
the answer. The client now tells them apart by whether it sent a token.
Two tests, one for each side of the rule.

Also: a launcher for pointing a Windows PC at a head office on the LAN,
with the two settings that needs and a comment saying why neither is
acceptable outside a demo.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-11 15:31:55 +05:30
719ba2c7f5 Recognition starts with the app, not with a button
The engine only ever started when somebody pressed Start. So a till
that rebooted overnight came back with the window open, the tray icon
showing, the session restored - and recognition off until a shop
assistant noticed. That is the failure the tray colours exist to catch,
and it should not be the default state every morning.

Guarded on the interpreter actually existing: on a PC where setup has
not run yet, the supervisor would loop on a missing executable with
nothing useful to say. Start and Stop remain for the case where somebody
has deliberately stopped it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-11 12:49:56 +05:30
5e1dcf7050 INSTALL.txt lives in the repo, not only inside a zip
The v0.3.0 release carried it and the repository did not, so rebuilding
the release from a clean state produced an empty file where the shop
operator's instructions should be. Caught by checking the byte count
before uploading, which is not a process.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
v0.3.1
2026-09-11 12:27:54 +05:30
92573e9067 The installer, run on a clean machine, found two bugs in itself
Ran behavision-setup in a fresh Linux container: Python 3.12, nothing
else, the release contents mounted read-only the way Program Files or a
shared drive would be. It failed, and then it failed differently, and
both failures would have been the client's first experience.

1. `pip install <folder>` makes setuptools write behavision.egg-info
   INTO the folder. The folder is read-only wherever a release is
   sensibly unzipped, so: "could not create 'behavision.egg-info':
   Read-only file system". The release now ships a wheel - pure Python,
   buildable anywhere, nothing to build on the shop PC, and pip never
   touches the unzipped folder. Source stays as a fallback and is copied
   somewhere writable first.

2. The engine's paths.py knows two worlds - frozen (ProgramData) and a
   checkout (the repo root) - and a pip-installed engine is neither. It
   resolved its state root to site-packages: database there, camera
   list there, and its generated API credential in a folder the app
   never reads, while the app looked in ProgramData. Every call would be
   401 on a stock install, with nothing in either log saying why. The
   same disease as the Mac checkout two days ago, now in production
   shape.

   engine.ChildEnv is the one place the engine's environment is built,
   used by the desktop app, the headless agent and the installer's own
   smoke test. It passes BEHAVISION_DATA_DIR = this process's state
   root, which paths.py honours ahead of every other rule, so the two
   halves agree by construction however the engine was installed.

   It also seeds config/default.yaml into the state root: a package in
   site-packages has no config beside it to seed from.

Re-run on the same clean container: seven steps, all pass, models
downloaded, engine started and answered, and its data/ landed beside
agent.json - not in site-packages.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-11 12:26:54 +05:30
92b12bcb1c A merchant can create a salesperson's login and hand it over
The flow this product is sold on is three tiers: the platform admin
registers a merchant, the merchant registers their sales staff, the
staff sign in on a phone. Tier 1 handed the new owner a password. Tier 2
could not - a manager could only mint an invitation code, which the
salesperson had to redeem themselves, on their own phone, choosing their
own password. Good practice, and no use to a manager setting somebody up
before their first shift with a card and a pen.

POST /api/team/members mirrors POST /api/admin/clients: generated
password unless one is given, returned exactly once, bcrypt-hashed on
the way in and not recoverable after. Same permission shape as an
invitation - manager and above, only an owner mints an owner, admin
refused - so a manager cannot do through one door what they are refused
at the other. The invitation path stays; it is the better one whenever
the salesperson has their phone.

POST /api/team/{id}/password is the everyday case on a shop floor:
they forgot it. It sets a new one AND revokes every session they hold,
in one transaction, because the other reason a manager resets a
password is a lost phone, and a reset that left that phone signed in
would look complete while fixing nothing. Tenant-scoped in the UPDATE
itself; another company's user id is 404, never 403. No self-service
and no reset-by-email, deliberately: a floor account often has no
mailbox anyone checks, and the person who can vouch for the salesperson
standing in front of them is their manager.

RandomPassword moves from a private helper in the store to auth, so the
admin path, the merchant path and the reset all mint the same 80-bit
credential - rather than someone later writing a shorter one for the
"less important" account.

Verified: eight handler tests, and two against a real Postgres for the
things a fake cannot see - the RETURNING list scans on a row with no
last_login_at, the tenant scope holds, and the sessions row is actually
revoked. The tenant cleanup from yesterday held throughout.

API.md now documents the chain with both paths, and the note saying a
merchant could not create a login directly is gone because it is no
longer true.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-11 12:12:54 +05:30
c50a74de47 The onboarding chain, as a chain
Admin creates the merchant, merchant invites the staff, staff redeem
the code on a phone. Every endpoint for it already existed and was
already documented - scattered across four sections in the order the
server groups them, not the order a person meets them.

Now one section, in tier order, each step with the request that makes
it and the response it hands to the next tier: the owner password shown
once, the invitation code shown once, the session returned by register
so a new salesperson is never sent to a login form. The status codes
were checked against the handlers: all three creations are 201.

Three absences named rather than left to be found: a merchant cannot
create a staff login directly (invitation only, on purpose); there is
no mobile app in this repository, only the API it will call; and an
admin cannot reset an owner's password or suspend a merchant over HTTP.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-11 11:54:47 +05:30
0a423ed8cc API.md documented 27 routes; the server has 48
A mobile developer builds against this file, so a gap in it is a gap in
the app. Checked route by route against the mux: nineteen routes had no
entry at all, including the ENTIRE platform-admin surface, adding and
checking cameras, issuing shop-PC installation codes, the assistant, and
the face bytes endpoint. Most of what was documented had no response
shape - a client had to guess the field names for shops, cameras, team,
customers, history and both reports.

Every shape here is now taken from the server's own types, and the
uncertain claims were checked against the handlers rather than written
from memory: check requests return 202, history is newest first, the
visitor list is most-recently-seen first and excludes the erased, an
admin slug is derived from the company name when omitted.

Restructured by audience, because "who may call this" was scattered:

  - three callers named up front - merchant, platform admin, shop PC -
    and what each one signs in with and sees
  - the three merchant roles and what each adds, taken from
    CanWriteProfiles / CanManageSites rather than paraphrased
  - a permission matrix: every route and the least role that may call it
  - quick starts for the three clients that will actually be written:
    a floor app for staff, a console for owners, and admin
  - /api/agent/* listed once as "not for you", so nobody wonders

The prose that explained WHY - refresh rules, the cursor, photos as data
not errors, the report arithmetic - is kept; that is the part a client
developer cannot get from the code.

Also recorded plainly: the admin API is two endpoints. There is no way
to suspend a company, delete one, or reset an owner's password over
HTTP. Written down rather than left for someone to discover.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
2026-09-11 11:31:44 +05:30
22196ab9ba Ship the engine from source, so a release can be built anywhere
The Go halves of this product cross-compile to Windows from any machine.
The engine does not: PyInstaller bundles the interpreter and the native
wheels of the machine it runs on, so a frozen engine can only be built on
Windows. That one fact was the entire reason no release had ever been
cut - two of the three binaries were ready for weeks.

behavision-setup installs the engine from source instead. It finds a
Python, builds a private virtual environment beside the database,
installs the engine into it, downloads the models, records how to start
it in the same agent.json the app reads, and then starts it and waits
for its API to answer.

That last step is the point. An installer that reports success and
leaves a shop with an engine that will not run has done worse than
failing: the failure surfaces later, to somebody who did not install it.

The trade, since whoever runs this is standing in a shop: it needs
Python and internet at install time and takes minutes, where a frozen
build needs neither. What it buys is a release that exists.

Details that are not incidental:

  - `py -3` is tried before `python` on Windows. The launcher is what the
    official installer puts on PATH; `python` there is often the Store
    stub that prints an advert and exits 9009.
  - a virtual environment, not the system Python. A shop PC may have
    Python for something else, and the engine pins numpy below 2.0 -
    installing that into a shared interpreter breaks the other thing
    months later and silently.
  - EngineExe is written absolute. The app resolves a relative one
    against its install root under Program Files, where no interpreter
    lives.
  - pip's output is shown, not swallowed. When it fails on a proxy or a
    missing build tool it says exactly what is wrong, and hiding that
    leaves the operator with "setup failed" and nothing to act on.
  - the console pauses before closing. Double-clicked from Explorer, a
    program that finishes closes instantly and success and failure look
    identical.

Verified as far as a Mac can: `pip install .` builds the wheel and
resolves every dependency, and `python -m behavision` then runs from
site-packages rather than the working directory - which is the mechanism
this depends on and had never been exercised, because the project has
only ever been run out of its own checkout.

NOT verified: any of it on Windows. Nothing here has run on the target
platform, and the `py -3` path and the ProgramData layout are exactly
where that will show.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pcn9asw19WGBfCEaHvNug6
v0.3.0
2026-09-10 20:21:32 +05:30
5e544eee3d The camera tiles put a password in the page, and loaded nothing
StreamURL built http://user:pass@127.0.0.1:8010/api/cameras/<id>/
stream.mjpeg and handed it to an <img>, with a comment saying the
credentials were inline "so an <img> tag can load it".

It cannot. Chromium strips credentials from subresource URLs and has
since M59, and WebView2 is Chromium - so on the one platform this
product ships to, every camera tile on a shop counter was a broken
image. Measured against a running engine: the app's Go-side calls
returned stats and people while an <img> on that very URL failed, and
curl proved the URL answered 200. The engine was never the problem.

The password now stays on this side of the process boundary. A loopback
relay attaches Basic auth and streams the engine's bytes back
unchanged - the same reasoning Shot.jsx already follows at head office,
where an <img> equally cannot carry a session.

What the relay is careful about, since it is a door onto the biometric
API with a credential attached:

  - loopback only, on a port the OS picks; a fixed one would collide
    with whatever else a shop PC runs and read as "the cameras broke"
  - a per-run random token in the path. The engine's own credential
    exists so the live face feed is never served open; an
    unauthenticated relay would hand that feed to any other process on
    the PC. Compared in constant time, and a wrong one is 404, not 403
  - an allow-list of stream.mjpeg and frame.jpg. Holding the token does
    not reach the identity list, the gallery, or erasure
  - camera ids validated, not interpolated
  - every chunk flushed; a buffered MJPEG stream is a tile that never
    paints, which looks identical to the bug being fixed

Two of those were written after a test failed, not before:

  - `..` MATCHES the id pattern, because real camera ids contain dots.
    `/api/cameras/../stream.mjpeg` is not the endpoint anyone intended.
    The id can never hold a slash, so `.` and `..` are the whole
    remaining traversal surface and are now refused by name.
  - the serve goroutine read p.srv off the struct while stop() was
    nilling it, so a quick start/stop dereferenced nil and took the
    process down. Captured before launching now.

FrameURL is deliberately not added. No screen asks for a still, and a
bound method nothing calls is the same defect as a capability the UI
cannot reach, only pointing the other way.

Verified: nine unit tests, plus a live test against the real engine and
the real office camera - two MJPEG frames, 90,793 bytes, no credential
in the URL. Windows and darwin both build; vet clean.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Pcn9asw19WGBfCEaHvNug6
2026-09-10 19:53:28 +05:30
9521cb986b The shop PC's UI had never once been run
`wails build` had never been executed against this project - CLAUDE.md
says so plainly - so every screen the shop floor actually touches was
unreviewed. Running it found why nobody had.

fyne.io/systray's nativeLoop must own the main thread on macOS, a Cocoa
requirement, and Wails already holds it. Starting both kills the process
with a SIGTRAP inside cgo before a single pixel is drawn. On Windows,
which is what ships, a tray on its own goroutine is fine - so the one
platform the whole team develops on was the one platform that could not
open the app, and the UI went unlooked-at as a result.

BEHAVISION_NO_TRAY runs the window without the tray, the same escape
hatch BEHAVISION_ALLOW_PLAINTEXT_MQTT already is for the broker.
Deliberately an environment variable and NOT a GOOS check: a build that
quietly drops the tray is how a shop PC ends up with no control surface
at all, and it would fail where nobody is watching. The guard is on stop()
as well, because systray.Quit() on a systray that never started is not a
no-op in v1.12.2 - it would turn closing the window into a crash on exit,
the failure most likely to be shrugged off as "it closed, fine".

go.mod gains the indirect dependencies the darwin build pulls in. No
version moved: the committed list was written by a windows-only build,
which never resolves that part of the Wails tree.

Verified: GOOS=windows build, go vet, and the agent suite all still pass,
and the packaged .app runs.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qiy5iKfz4L8S4vRaYPBdaU
2026-09-09 13:06:51 +05:30
30e01765ae The live tests seeded a tenant per run and never took it back
Each live store test makes its own client - deliberately, so they can
run in any order and so the isolation assertions have a real neighbour
to be isolated from - and none of them removed it afterwards. The dev
database had reached 242 abandoned tenants against the one real
company.

That is not untidy, it is a broken screen. The platform admin's
Companies view lists every client, so the real company sat under pages
of `walk1788761685056287000`, which is the first thing anyone opening
tenant administration would see.

dropTenant registers the cleanup against the CLIENT rather than each
table: every foreign key onto clients is ON DELETE CASCADE, so one
delete takes the sites, visitors, visits, face images, embeddings,
cameras and agents with it. A per-table list would rot the first time a
migration adds a table, and it would rot silently - the same shape as
the leak it replaces.

A failed cleanup calls t.Errorf rather than being ignored. A tenant
left behind is precisely what this exists to prevent, and swallowing
the error would let the leak come back with nothing to show for it.

Verified against the live database: three consecutive runs of the store
suite leave clients, sites and visits unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qiy5iKfz4L8S4vRaYPBdaU
2026-09-09 12:43:21 +05:30
ee9e8b80b7 A backup of .env is still a copy of the camera password
Editing .env leaves .env.bak-<timestamp> beside it, and only the
anchored /.env pattern was ignored - so the backup showed up as an
untracked file holding the RTSP password in plaintext, one `git add -A`
away from being committed. The pattern that protects the original has to
protect its copies.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Qiy5iKfz4L8S4vRaYPBdaU
v0.2.0
2026-09-09 11:56:15 +05:30
3598d8e9c0 The shop PC's avatar had the same V1 collision
Fixed on the web arrivals feed and not here, which is the failure this
codebase already warns about: two surfaces disagreeing about one fact.
Taking the first letter of each word of "Visitor 13" gives "V1" - and so
do "Visitor 10" and "Visitor 15", so three different customers wear the
same badge and it reads as the V-1 reference for a fourth.

Shows the number itself, same rule as the web app. customerRef, not
ref: React reserves that prop name and it would never arrive.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-07 12:45:38 +05:30
ce0223006b References are immutable, because clients now store them
012 turned three descriptive columns into identifiers other systems
keep: in agent.json on a shop counter, in a saved URL, in a scheduled
report. All three were already treated as stable and none of it was
enforced.

- clients.slug is an MQTT topic segment the broker ACL is written
  against. Rename one and that tenant's whole estate is silently refused
  by the broker, with no way to tell the agents.
- sites.slug is what a shop PC calls itself - agent.json holds
  "site_id": "chennai", never the uuid. A rename orphans the PC from the
  shop it is standing in.
- site_cameras.camera_id lands in visits.camera_id, which is text and
  not a foreign key. A rename orphans every visit already attributed to
  the old name: the footfall is still there and no longer joins to a
  camera. This was half-enforced in handleUpdateCamera and nowhere else,
  which is the shape of a rule that holds until somebody adds a second
  write path.
- visitors.number is assigned once from the tenant's counter and read
  back as V-42.

A trigger, not a CHECK: a CHECK cannot see the old row and the rule is
about the transition. The DISPLAY name is deliberately not frozen -
"TeNext Chennai", "Front door" - it is what a person reads, nothing keys
on it, and a system that cannot fix a typo in a shop's name has confused
the two.

Also records why the uuid stays where a slug would do. The length was
never the problem; needing it was, and that is fixed. Replacing it would
touch eight foreign keys on a live database to shorten a field clients
are already told not to use, and a sequential id would make any future
tenancy hole walkable by counting. It is NOT because ids must be minted
offline - sites, visitors and visits are all created server-side with a
database in hand, and claiming otherwise would defend the status quo
rather than explain it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-07 12:17:49 +05:30
08873f4a67 Three uuids on one arrival, three different answers
Asked of the row the feed actually returns.

site_id had a reference all along and the feed was not sending it. A
client could read the shop's NAME off an arrival and still had no way to
ask for that shop except by uuid - the exact gap the reference scheme
exists to close. site_slug now travels with it.

visit_id stays a uuid and needs no reference: no route takes it, it is a
key a client de-duplicates on because delivery is at-least-once, and
nobody says a visit id out loud.

The uuid in a face URL must STAY random. visit_faces.id is
gen_random_uuid() and a derived or sequential one would let somebody
walk a shop's customers by date - the same reason bucket keys are random
rather than derived from the event id. A readable identifier is right
for a customer and wrong for the thing that points at their photograph.

And seq is now json:"-". visits.seq is a plain bigserial, so it counts
every visit on the PLATFORM, and shipping it put the total footfall of
every customer we have on every row of every tenant's feed - the same
German-tank estimate that decided visitors.number had to be per client.
It was a convenience for "have I fallen behind", nothing ever read it,
and the cursor answers that without disclosing a number. The SSE event
id was never the raw value; it has always been the opaque cursor.

The one test that broke was reading seq back off the wire to assert the
cursor pointed at the last row of a burst. It asserts against the seeded
position now: the property is unchanged, and the test can no longer see
what a client cannot.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-07 12:12:48 +05:30
9182f70442 A customer number people can say out loud
Every id in the schema is a uuid and stays one. What was wrong was
putting one in front of a person: RecordVisit named every new customer
'Visitor ' || left(id::text, 8), so the arrivals feed, the shop PC and
the mobile app all read "Visitor 3446ec35" - the string a shop assistant
reads to a colleague and types into a search box. label is a stored
column staff can overwrite and SearchVisitors matches on, so formatting
around it in a front end would have left the data wrong on three
surfaces.

Migration 012 adds a per-client visitors.number, taken from a counter on
clients with UPDATE ... RETURNING inside the visit transaction. Per
client rather than global: a global sequence would tell any customer who
signs up how many people the whole platform has ever seen, from their
own first visitor number. The backfill numbers existing rows by
first_seen_at and relabels only the eight-hex pattern the old statement
produced, so a human-typed name is never overwritten.

Three of the four things anyone addresses by URL already had a human
name and the API simply refused it - a site has a slug, a camera has the
id the engine knows it by. refs.go accepts either form anywhere an id is
taken; a uuid resolves with no lookup, so every URL a client already
stored keeps working.

- An ambiguous camera name resolves to nothing, never to a guess: two
  shops may each have an "Office1" and acting on the first row would
  edit the wrong shop's camera.
- 404 on a path, 400 on a query filter. /api/visits answered fine and it
  was the filter that was wrong.
- site and site_id are both accepted everywhere now. They differed per
  endpoint, and an unknown query parameter is silently ignored, so
  getting it the wrong way round returned the whole estate.
- The search matches V-13, which is what the product now shows.

Two bugs found by running it rather than testing it:

- 'Visitor ' || $2::text beside number = $2 makes Postgres deduce two
  types for one parameter and refuse the insert. It compiled and passed
  every in-memory test; the first real database rejected it, along with
  the existing face tests that share the path.
- The fallback avatar said "V1" for Visitor 13, Visitor 10 and Visitor
  15 alike, and read as the V-1 reference for a fourth person. It shows
  the number now. The prop is customerRef, not ref - React reserves
  that name and it would never have arrived.

Verified on the live database and through the running API: 13 hex labels
became Visitor 1-13 in first-seen order, two typed names left alone, and
the same customer reachable by uuid, V-13 and 13.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-07 11:52:32 +05:30
3f9fb33b24 Accounts people can create, and photos on a server with no bucket
A tenant had exactly the users somebody had created with a command on the
server. That is not a missing screen: a shop with an owner and four staff
either shared one password or raised a ticket per person, and a phone app
for the shop floor could not exist while there was one account to sign in
as.

Registration is by invitation, never open signup - the same line already
drawn around creating a company. The code carries the address and the role
and the request carries only a password, so a code that gets forwarded
cannot become somebody else's account, and a staff invitation cannot be
redeemed as an owner. Single use lives in the UPDATE and the account is
created in the same transaction.

Deactivating a member revokes their sessions in that transaction too. An
access token lives twelve hours, so without it "remove their access"
removed it sometime tomorrow. The session list and revoke that go with it
are the benefit of opaque tokens the product had been paying for and never
collecting: nothing could say what was signed in, let alone stop one.

Face images now work on a deployment with no object storage, which was
every local install and every self-hosted site - the arrivals feed said
"not storing customer photos" for every customer forever, on the screen
whose whole job is to show a face. Bounded to one row per visitor, so it
grows with the customer base and not with footfall; the bucket stays
primary wherever one exists.

Image.auth says whether a URL needs the session, because a browser img
cannot load one that does, a mobile image view can, and a webview can do
neither - the desktop client resolves those to a data URI in Go.

Found by running it, not by tests:

  * UPDATE ... RETURNING gives the value AFTER the update, so the prune
    read back empty keys, deleted nothing, and the table grew with
    footfall exactly as if it were not there. The fake agreed with either
    version; only the live Postgres test caught it.
  * Trusting only the auth flag broke every shop card, because Sites.jsx
    rebuilt a partial snapshot object and dropped it. A relative URL is
    now sufficient on its own.
  * ago() renders a future time as "just now", so a code valid for a week
    read "expires just now".

Verified live against real Postgres: invite, preview, escalation refused,
register into a session, replay 404, staff forbidden, device revoked and
401 at once, last owner refused, and a 92,405-byte camera JPEG stored,
served to its owner, 401 with no session, 404 to another tenant, and
rendered in a browser.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-05 11:45:42 +05:30
ffae7e45d5 Live view runs at the camera's real rate, and reports why it is MJPEG
4 fps was not "live", and it was a number I picked rather than measured.
The engine actually produces ~12 distinct frames a second, so most of it
was being left on the floor.

Now: poll a little ahead of the engine and drop frames identical to the
last one by hash. Measured end to end - 131 frames in 10 s, 13.1 fps,
20.3 KB each, 259 KB/s, zero duplicates. Every byte on the wire is a
picture the viewer has not seen, and the rate follows the camera instead
of a constant.

Also records why this is MJPEG rather than passing the camera's own
compressed video through, which would be smoother, cheaper and use no
CPU. Probed the office camera: main 2304x1296@15, sub 800x448@15 - and
BOTH are H.265, despite stream paths ending in ".264". Browsers play
H.264 everywhere and H.265 only on some platforms, so passthrough cannot
rely on it, and transcoding HEVC on the shop PC would put a video encoder
on the machine already doing the recognition.

So probe_source now reports `codec`. It decides what is possible, an
installer can usually change it, and otherwise the only way to learn it is
to read RTSP by hand - which is how this was found.

The RTSP libraries used to establish that are NOT kept: they were only
ever imported by a spike test, and two large dependencies in a shipped
binary to answer a question OpenCV already knows is a bad trade. Their
`go get` had also silently bumped the agent to go 1.25 and broken the
desktop build, which is its own argument.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 16:59:27 +05:30
18686cbceb Live view at head office, relayed through the agent's outbound connection
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.

The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.

Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.

Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.

LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.

Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.

Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 16:46:23 +05:30
2cd7a78ddc A headless PC can be claimed, and a refused broker says so
Both found by operating the stack rather than writing it: the local
processes were OOM-killed and bringing them back hit two gaps.

The headless agent had no way to be claimed at all. Bootstrap lived only
in desktop/internal/cloud, so the one configuration the agent binary
exists for - a back-office PC with no window - could only be onboarded by
hand-editing agent.json, which is the state the desktop's Setup screen was
built to end. `behavision-agent claim <code>` closes it; the CLI joins its
arguments because the code is printed in groups for reading aloud and an
operator pasting it will paste the spaces too.

Second: after the site's broker password was re-rolled, mosquitto logged
"not authorised" while the agent logged "timed out". Those need opposite
actions - re-link this PC, or go and look at the network - and paho's
SetConnectRetry collapses them, because it retries internally and the
connect token never completes. describeStall asks whether a TCP socket
opens at all, and says what is known rather than guessing at a reason the
broker never gives.

Verified end to end: minted a code from the platform as the owner,
claimed with the new command, broker connected, and the shop went to
online: true with 1/1 cameras on w600k_r50.

Also corrects this machine's memory in CLAUDE.md from 16 GB to 8 GB. It
feeds the model-fallback reasoning, and the local gallery already holds
17 embeddings tagged w600k_mbf beside 19 tagged w600k_r50 - the fallback
has silently fired before.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 13:32:38 +05:30
e0ceb14589 Camera pictures without an object-storage bucket
Head office shows a camera's latest frame rather than live video, for a
reason that has not changed: the engine serves MJPEG on 127.0.0.1 on a PC
behind a shop's router with no inbound route, and relaying it needs
WebRTC/TURN. Pointing a browser straight at the shop PC is not the escape
either - the engine's API is Basic-authenticated with a credential it
generates locally and never sends anywhere, and shipping that to the
cloud so a web page could use it would put the key to the biometric API
and the live face feed in the server's database.

But that picture only worked if you had an S3 bucket. Without one,
attachSnapshots reported "This system is not storing images" for every
camera forever - on the two screens whose whole job is to show the
camera. Making them picture-led turned a missing feature into a wall of
empty tiles, on every local install and any self-hosted customer who does
not want a bucket.

migrations/009 adds camera_snapshots and the agent falls back to
PUT /api/agent/cameras/{camera}/snapshot when the presigned route answers
images_disabled - chosen by sentinel, never by matching the message, since
it picks between two routes. One row per camera is what makes this safe in
the database when face images are not: the key IS the camera, so storage
is (cameras x ~100 KB) and does not grow with footfall.

The read is session-authenticated rather than a signed link, which an
<img> cannot use - hence Shot.jsx and useAuthedImage, keyed on the URL
string rather than the snapshot object so a poll does not re-fetch 90 KB
per camera every few seconds, and revoking the object URL on cleanup.

Verified against the real office camera with no bucket configured: 90,587
bytes stored in Postgres, served as image/jpeg to a signed-in user, 401
without a session, rendered on both the Cameras and Shops cards.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 12:53:18 +05:30
c7024b57ca The shipped config could not load on a machine without a .env
config/default.yaml reads its camera's host, username and password from
${ENV}. Unset placeholders parse as YAML null, and CameraConfig had no
_normalize_blanks - the guard ApiSection and EmailSection have had all
along - so loading it raised three pydantic errors and the engine would
not start at all. On a developer's checkout .env is right there, which is
why this survived: the failing machine is every machine the product is
actually installed on, and installer/build.ps1 runs this suite, so the
Windows build would have failed on a fresh clone.

The normalisation is field-by-field, never a blanket None -> "": `webcam`
is an Optional[int] whose None means "this is not a webcam", and `tuning`
is a nested model. Sweeping either trades one validation error for
another - which it did, on the first attempt.

Second bug behind the same line: that camera entry would then have been
SEEDED into a fresh install, giving a shop a camera called cam1 that
nobody added, retrying a connection to "" forever, with the first task on
a new PC being to work out what it was. CameraConfig.addressed() says
what a camera entry needs to be one, and seed() drops the rest.

Found by cloning the repository into a temp directory and running the
tests there. Nothing in a working tree can find this class of bug.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
v0.1.0
2026-09-04 12:15:53 +05:30
2b69a0be3d Anchor the gitignore patterns; a fresh clone did not compile
`spool/` is unanchored, so it matched `agent/pkg/spool/` - the durable
queue, source code - and the repository excluded it. Cloning and building
was what found it; nothing in the working tree ever would, because the
files are right there.

`data/` and `agent.json` have the same shape and are anchored too.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 12:08:01 +05:30
5453c26e4c The schema applies itself, and the setup script stops hiding failures
Migrations were run by hand and nothing recorded which had run, so
re-running the setup script against an existing database failed on the
first CREATE TABLE, and shipping a new migration gave an operator no way
to know whether an estate had it. A missed migration is not a startup
error - it is a query referencing a column that is not there, surfacing
later on whichever endpoint touches it first.

server/internal/migrate applies pending migrations at boot and refuses to
start against a schema it does not match. One transaction per file
holding both the DDL and the row that records it; an advisory lock so two
servers starting at once cannot both apply 008; checksums so an edited
migration is refused by name rather than silently skipped; numeric
ordering so 010 does not run before 009. `migrate -baseline N` adopts a
database built before any of this existed, because "the clients table
exists" does not say whether 007's index does.

Verified on the live database: adopted 001-007, applied 008.

008 adds two indexes on `purchases`, found by asking the database which
foreign keys had nothing behind them and then checking what queries the
table. The conversion report filters client_id + occurred_at, which is
exactly the estate-wide case with no site to narrow it.

run-local.sh had two bugs, both found by running it rather than reading
it: it reused a broker container whose bind mount pointed at a directory
that no longer existed, and it discarded stderr on the mosquitto_passwd
call, so under `set -e` it exited at step 5 with no output at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 12:06:52 +05:30
dad04e8cda Behavision: face recognition for retail, edge to head office
Five components that ship as one product:

- behavision/  the recognition engine. RTSP ingest, YuNet detection, IoU
               tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
               FastAPI dashboard. Identity is decided once per TRACK from an
               average of at least three embeddings, never per frame.
- agent/       the Go edge agent: supervises the engine, holds a durable
               spool, and drains it to MQTT. Nothing is acked before the
               broker confirms.
- desktop/     the shop PC application (Wails + React + tray).
- server/      the cloud API, MQTT consumer, reports and assistant.
- web/         platform.loyaly.ai, the head-office app, embedded in the
               server binary.

The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.

CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
2026-09-04 11:14:18 +05:30