The schema applies itself, and the setup script stops hiding failures

Migrations were run by hand and nothing recorded which had run, so
re-running the setup script against an existing database failed on the
first CREATE TABLE, and shipping a new migration gave an operator no way
to know whether an estate had it. A missed migration is not a startup
error - it is a query referencing a column that is not there, surfacing
later on whichever endpoint touches it first.

server/internal/migrate applies pending migrations at boot and refuses to
start against a schema it does not match. One transaction per file
holding both the DDL and the row that records it; an advisory lock so two
servers starting at once cannot both apply 008; checksums so an edited
migration is refused by name rather than silently skipped; numeric
ordering so 010 does not run before 009. `migrate -baseline N` adopts a
database built before any of this existed, because "the clients table
exists" does not say whether 007's index does.

Verified on the live database: adopted 001-007, applied 008.

008 adds two indexes on `purchases`, found by asking the database which
foreign keys had nothing behind them and then checking what queries the
table. The conversion report filters client_id + occurred_at, which is
exactly the estate-wide case with no site to narrow it.

run-local.sh had two bugs, both found by running it rather than reading
it: it reused a broker container whose bind mount pointed at a directory
that no longer existed, and it discarded stderr on the mosquitto_passwd
call, so under `set -e` it exited at step 5 with no output at all.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
This commit is contained in:
2026-09-04 12:06:52 +05:30
parent dad04e8cda
commit 5453c26e4c
11 changed files with 941 additions and 8 deletions

187
CLAUDE.md
View File

@@ -1984,3 +1984,190 @@ JPEG downloaded → erase → presigned GET **404**, 0 templates, 0 profiles, la
setInput/forward; it runs once per track, so the lock costs nothing) and
`Gallery`/`IdentityStore`. Locks around shared JPEG state; SQLite in WAL.
- Tests are dependency-light (no camera or models needed) — keep it so.
## A shop PC that runs on its own
Until now every install was blocked on an enrolment code: the app opened on the
setup screen, and there was no way past it except a credential issued by a
server. That is wrong for the product it claims to be. Recognition, the
cameras, the tracker and the local gallery all run on the shop PC and need no
network at all — so a single-till shop with no head office was being refused
the thing the software is *for* until a component it does not need had blessed
it.
`RunStandalone()` is the second answer on that screen. It is persisted
(`Config.Standalone`), because a choice that lives only in memory puts the
enrolment-code screen back in front of a shop that has already answered, which
reads as the app forgetting it was ever set up.
- **"Not linked yet" and "not going to be linked" are different states**, and
the difference is load-bearing rather than cosmetic. An unclaimed PC keeps
its bridge running and queues every visit, deliberately: *"footfall from the
day it was installed is on disk waiting for credentials rather than lost."*
A standalone PC must not. Nothing is ever going to drain that queue, so it
would write up to `SpoolMax` visits — **each carrying a face template, which
is biometric personal data** — to disk for no purpose. `startPipeline`
returns early; the camera reconciler still runs, and reconciles against
nothing.
- **Cloud-only screens are hidden, not shown broken.** The customer record
lives on the server, so Customers disappears; Live and Cameras stay, because
they read the engine on loopback and always could.
- **Live says "Running on this PC only", not "Not linked to head office"** with
an idle dot beside a count of zero. The second is what a fault looks like.
- **Claiming later clears the flag and restarts the pipeline**, so a shop that
grows into a second branch loses nothing it recorded on its own.
`Session()` and `PipelineStatus()` both report `standalone` as
`cfg.Standalone && !cfg.Configured()` — one expression, in two places that must
never disagree, for the same reason the tray is a client of `EngineStatus()`
rather than a second copy of it.
## The camera make picker is now on both forms, from one file
`shared/cameraMakes.js`, imported by the head-office web app **and** the shop
PC's app rather than copied into each. A make that is right in one and stale in
the other is worse than not offering the list at all, because an installer
trusts a filled-in field.
It exists because the RTSP **path** is the one field nobody can look up: the
address is on a label and the password is in the installer's notes, but the
path is model-specific and undiscoverable, and getting it wrong produces
*"could not open stream"*, which reads like a password problem and is not.
The desktop form also picked up the autofill defence the web form already had:
a text input next to a password input is a sign-in form as far as a webview is
concerned, so without `autoComplete="new-password"` on the secret and a
non-login `name` on the account, the browser offers the operator's own
Behavision email as the camera's username — which then fails with a message
about credentials that points at the camera.
## The schema applies itself (`server/internal/migrate`)
The migrations were run by hand — `psql < 001.sql`, in order, by whoever
remembered — and **nothing anywhere recorded which had run**. Three
consequences, all of which had already happened:
- Re-running the setup script against an existing database failed on the first
`CREATE TABLE`, so it only ever worked once. Discovered by running it.
- Shipping a migration 008 gave an operator no way to know whether an estate
had it. A missed migration is not a startup error; it is a query referencing
a column that is not there, surfacing later on whichever endpoint touches it
first.
- An interrupted file left a schema nothing could describe.
Now: `schema_migrations`, and the server applies pending migrations at boot.
Applying at boot rather than as a deploy step is deliberate — an upgrade of
this product is *"copy the new binary and restart it"*, and a migration
somebody has to remember to run is one that does not get run. Failure **stops
the server**: one running against a schema it does not match writes wrong data,
and wrong data outlives the outage that stopping causes.
Decisions worth keeping:
- **One transaction per file, holding both the DDL and the row that records
it.** A migration that ran but was not recorded runs again next start; one
recorded but not run leaves a missing column nothing will ever add.
- **An advisory lock held on one connection for the whole run.** Two servers
starting at once is the normal shape of a rolling restart, and both deciding
008 is pending is not a hypothetical.
- **The content is checksummed.** An already-applied file that has since been
edited means the database does not contain what the repository says it does,
and running the new text now would apply half of it twice. It refuses and
names the file: the fix is a new migration, never an edited one.
- **Ordering is numeric, not alphabetical.** At ten migrations `010` sorts
before `009` as text, and the failure arrives on the day the project reaches
double figures. Two files claiming one version — the ordinary result of two
branches both adding `008` — is refused outright, because whichever ran first
would then be decided by the filesystem, which is not an order.
- **`-baseline N` adopts a database built before any of this existed**, marking
1..N applied without running them. Guessing was not an option: *"the clients
table exists"* does not say whether 007's index does. An operator states it
once, and the row is marked `baselined` so an adopted database never looks
like one this code built.
- **The migrations are `go:embed`ed**, so the schema travels inside the binary
it belongs to. The consequence to know: a stale binary reports "schema up to
date" about migrations it has never heard of. Rebuild, then migrate.
`behavision-server migrate [-status|-baseline N]` is the operator's view.
Verified on the live database: adopted 001–007, then applied 008 (two indexes
on `purchases`, found by asking the database which foreign keys had nothing
behind them and then checking what actually queries the table — the conversion
report filters `client_id` + `occurred_at`, which is precisely the estate-wide
case with no site to narrow it).
## The Windows package (`installer/`)
`installer/build.ps1` builds it and `installer/behavision.iss` lays it out.
**The build script must run on Windows, and that is not a preference.** Every
other artefact here cross-compiles from a Mac — the Go binaries with
`GOOS=windows`, the front ends with npm, verified — but PyInstaller freezes the
interpreter and the native wheels (onnxruntime, OpenCV) of the machine it runs
on. There is no cross-target flag and there never has been. So the engine .exe
is built on Windows or it is not built.
Installed layout, and why it is not flat:
```
C:\Program Files\Behavision\
Behavision.exe the app: window, tray, engine supervisor
behavision-agent.exe the headless agent, for an install with no UI
engine\behavision.exe the engine, plus ~150 native DLLs beside it
C:\ProgramData\Behavision\ everything written: database, logs, models, cameras
```
The engine keeps its own folder because it is a one-**folder** PyInstaller
build that brings its DLLs with it — and because Windows filenames are
case-insensitive, so `Behavision.exe` and `behavision.exe` could not share a
directory even if it were tidy to. `config.Defaults().EngineExe` names
`engine\behavision.exe` and `tests/test_installer.py` asserts the two agree:
if they ever disagree the app starts, shows a healthy window, and recognises
nobody.
- **Admin at install time, never at run time.** Program Files needs elevation;
spawning a child process does not. This is the same reason the product is not
a Windows service: a service runs in session 0 and cannot draw a tray icon.
- **Nothing writable under Program Files.** That split is `behavision/paths.py`
and `agent/pkg/paths`, and the installer must not contradict it — a seeded
writable file under `{app}` works for the administrator who installed it and
fails for the shop assistant who uses it.
- **Models are not bundled.** ~200 MB, downloaded resumably on first run;
bundling them quadruples the installer and forces a re-sign for a model
change. The wizard offers the download and a Start-menu shortcut repeats it,
because a shop PC being set up often has no working internet yet.
- **The WebView2 bootstrapper IS bundled.** Without the runtime the app opens
as an empty white rectangle — not an error, just nothing — which is the worst
failure to hand a shop. Present on Windows 11 and recent Windows 10, absent
on plenty of older machines, and a shop counter is exactly where an older
machine lives.
- **`CloseApplications=yes`.** Replacing the engine's DLLs while it holds the
SQLite WAL and the camera produces a half-upgraded install that fails on the
*next* start, long after anyone would connect the two events.
`tests/test_installer.py` is the same guard `tests/test_paths.py` is for the
PyInstaller spec: it asserts the installer and the build script name the same
files, that nothing writable is placed under the install root, and that no
model is bundled. It cannot prove the package works on Windows — only a Windows
box can — but it catches the class of mistake that would otherwise get that
far.
**Not yet done, and it needs a Windows machine:** no `wails build` has ever
run, no installer has been compiled, nothing is code-signed, and the frozen
engine has never been started. Unsigned, SmartScreen will warn on first launch.
## `run-local.sh` was hiding its own failures
Two bugs, both found by running it after a reboot rather than by reading it.
- **A reused container keeps the bind mount it was created with.** `bv-mqtt`
had been created while the working directory was somewhere else, so it came
back up with an empty `/mosquitto/config`, died with *"Unable to open config
file"*, and every `docker exec` after that failed for a reason having nothing
to do with what it was asked. The script now compares the mount source and
recreates the container when it has moved.
- **`>/dev/null 2>&1 || true` on the `mosquitto_passwd` call.** Under `set -e`
the script then exited at step 5 with **no output at all** — the single
hardest failure to diagnose, and it took three runs to find. stderr is no
longer discarded, and the broker is waited for and reported on if it will not
stay up. A failure there means the server cannot authenticate to its own
broker, which is exactly what this script exists to surface early.