Found by walking the scenario against production rather than by a test.
Two records, each with a phone; the survivor kept its own, and the
source's simply stopped existing. Searching for it returned nothing.
The first version's rule was "fill the survivor's blanks, never overwrite
what it has", which is right about which value WINS and said nothing
about the one that loses. One person can have two numbers, two spellings
of a name, a work address and a personal one - and a merge that quietly
deletes one is exactly the data loss this file already refuses elsewhere:
"silently turning Alice back into Visitor 3 is data loss the operator
cannot see happen."
The profile is now reconciled field by field in Go rather than in one
clever upsert, because the interesting case was never the winner. Blanks
are still filled and the survivor still keeps its own values, but every
losing value is returned in `discarded` AND appended to the survivor's
notes - the response is read once and the record is read forever.
Notes themselves are additive rather than a winner: two people writing
about one customer wrote two different true things.
mergeProfiles is pure, so the rule is asserted directly - four cases
including the ordinary one, a typed record joining a camera record with
no profile at all, which must add no noise.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
POST /api/customers and POST /api/visitors/{id}/merge. They ship together
because the first creates the need for the second: a customer typed in at
a counter has no face template, so when a camera sees that person later
the matcher has nothing to compare against and enrols them as somebody
new. That is the design working, not failing - and it means every
hand-created customer is a duplicate waiting to happen. Shipping the
create alone would manufacture duplicates into the state CLAUDE.md
already flags: "there is no merge endpoint server-side, so its
duplicates would be unrecoverable."
The number comes from clients.visitor_seq, taken exactly as RecordVisit
takes it. Two sources of visitor numbers that could disagree would be
worse than none: V-42 has to mean one person whichever way they arrived.
The label is the typed name, or "Visitor N" when they gave none - the
same string the engine writes, so a record created by hand is
indistinguishable from an enrolled one afterwards.
The merge is one transaction over FIVE tables, and the count is the
point. visits, purchases, visitor_embeddings, consents and
visitor_profiles all reference visitors ON DELETE CASCADE, so a table
this forgets to re-point is not an error - those rows are destroyed with
the source and nobody finds out until a customer's history is short.
visitor_profiles is UNIQUE on visitor_id, so the two cannot simply both
move and something has to win. Blanks on the survivor are filled from the
source and nothing it already holds is overwritten, which is exactly
right for the case this exists for: a hand-typed name and phone joining
the face that was recognised a week later.
Policies carried over from the edge gallery's merge, which had to settle
all of this once already: a human-assigned name outranks an auto
"Visitor N" whichever direction the operator merged; visit_count is
recomputed with COUNT(*) and never summed, because the stored counter may
be stale and the row count cannot be; first_seen_at takes the earlier of
the two, since it is one person and always was.
Two things that are this side's own:
- The source is deleted for real, not soft-deleted. A tombstone would
leave its number resolving to a record holding nothing, which reads as
"this customer exists and has never been here" - a worse answer than
"no such customer".
- The response names the RETIRED reference. Staff write V-42 on cards and
read it aloud; a merge that does not say which one stopped working
leaves somebody to discover it at a counter.
Manager and above, not staff. Apart from erasure this is the only
irreversible operation on a customer: two people welded together cannot
be separated, because nothing records which visit came from whom. It logs
at WARNING and writes an audit row for the same reason.
Also fixed while here: two s.Log.Printf calls - one of them mine, from
the password endpoint - that would panic on a nil logger. The package has
a nil-guarded s.logf and those were the only two not using it. The
password one sat in an error path no test reaches, which is exactly where
that bug waits.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
POST /api/auth/password. The cost of its absence was measured today
rather than argued: rotating three production accounts took a shell on
the host, three round trips, and briefly left a PLATFORM ADMIN - the
account that reads every company on the estate - with the password
PASTE_IT_HERE, because a placeholder in a pasted command was taken
literally and there was no way to correct it from the product.
A manager could always reset somebody ELSE's password. A platform admin
could be reset by nobody: they have no client, so the team routes are
not theirs, and `provision user` on the host was the only route. For
software that puts accounts on shop-floor PCs and staff phones, this is
not a feature - it is what makes every other credential decision
recoverable.
Three decisions:
- **authed, not tenantOnly.** A session is not a company's data, and the
account with no company is precisely the one that had no route. Scoping
this by client would have reproduced the hole it exists to close, which
is also why SetUserPassword is not scoped by client the way
ResetMemberPassword beside it is. The user id comes from the verified
session, never the request, so there is nothing to point at anyone else.
- **The current password is required.** An access token lives twelve
hours and travels on devices that get lost and shared; without this a
stolen one owns the account permanently instead of until it expires.
- **Every OTHER session is revoked, and the caller's is kept.** Somebody
changing their password because they believe it is known must not have
to wonder whether the device that already had it is still signed in -
and must not be signed out of the one in their hand while dealing with
it. A failure there is logged, not returned: the password IS changed by
then, and reporting an error would send them to retry with a current
password that no longer exists.
The suite's login() helper fatals on anything but 200, which is right
everywhere else and useless here - half of what these tests assert is
that a password has STOPPED working. loginCode() returns the status.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
"When was this customer last in" and "did they buy" are one question
staff ask in one breath, and answering it meant two calls and a join in
the client. Each visit row carries purchases, spend and currency.
LATERAL, not a join onto purchases. A plain join returns the visit TWICE
when it holds two sales, which would make a customer look like they came
more often than they did - a wrong number of exactly the kind this
product is otherwise careful about, arrived at by adding a feature.
Mixed currencies on one visit report the count and NO figure. Adding
rupees to dollars produces something that looks like money and is not,
and the sales still happened, so the count is the honest part to keep.
A purchase with no visit_id is deliberately absent: it belongs to the
customer rather than to a moment, and GET /api/sales?customer=V-42 lists
it. The two surfaces together cover every sale exactly once.
Both properties are asserted in the LIVE store tests, because both live
in the SQL. An in-memory fake asserting that a LATERAL does not duplicate
a row would only be checking the fake.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
/api/visits, /api/cameras, /api/sites, /api/visitors and
/api/reports/footfall, all in production, all before today's work. A
platform admin is defined by having NO client, and every tenant query
scopes on client_id = $1::uuid - so the empty string reaches Postgres as
''::uuid, which is a cast ERROR rather than an empty result. Found by
calling them while verifying the new routes, which have the same shape
and were failing the same way.
tenantOnly is the guard, beside adminOnly and for the opposite audience.
Per-query casts would have been the wrong fix twice over: it is a fix the
next query forgets, and the next query would then 500 in production
exactly as these did.
403, not adminOnly's 404, because the two hide opposite things. A tenant
must not learn a platform surface exists. A platform admin already knows
the tenant surface does - they are reading its data through /api/admin -
so nothing is concealed by pretending otherwise, and the refusal names
the route to use instead. "Forbidden" alone sends somebody hunting a
permissions problem that does not exist.
/api/auth/* stays on plain authed: a session is not a company's data, and
signing out or revoking a lost device must keep working for an account
with no tenant.
The fake could not have caught this either - it compares client ids as
strings and is perfectly content with "". The test asserts the contract
(403 and a message naming /api/admin) and a third case that matters more
than either: an ordinary tenant user still reaches all of it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
/api/admin/clients/not-a-uuid/sites returned 500. `c.id = $1::uuid` makes
Postgres cast the path segment, and casting a malformed string - or the
empty one the shape check handed back in its place - is an ERROR, not a
miss. `c.id::text = $1` cannot fail: an id that is not a uuid matches
nothing, which is the 404 a wrong URL should get.
The two sibling resolvers were already written this way and correctly
404'd the same input. I applied the rule to two of three places, which is
the shape of a rule that holds until somebody adds the next write path.
The shape check is gone with it - it existed only to produce the empty
string that then broke the cast.
The in-memory fake could not have caught this and did not: it resolves a
merchant with a map lookup, so every handler test passed, including the
one named for the case. That test stays, because 404-not-500 is still the
contract, but the property belongs to Postgres - so
api_admin_monitor_live_test.go asserts it where it lives, over every
free-text identifier these queries take. It skips without
TEST_DATABASE_URL, like the rest of the live store tests.
Also in deploy.sh, found by reading its own output: step 3 reported the
WRONG backup. `ls | tail -1` sorts alphabetically, so pre-...-demo-12
sorts before pre-...-demo-6 and it printed a dump from four days earlier.
A deploy that names the wrong safety net is worse than one that names
none, because that is the file somebody reaches for at the worst possible
moment. It echoes the filename it just wrote, and refuses to continue on
an empty one - pipefail catches a failing pg_dump, but a zero-byte gzip
would still have satisfied it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Three routes over data the server already stores.
GET /api/sales and /api/sales/{id}. The purchases table has existed
since the conversion report did, and nothing could read a row of it - so
"revenue was 41,000 last week" was a number that could not be checked
against a till. The list carries the customer reference the product
actually shows people (V-42) beside the uuid, and a sale with NO
customer is listed rather than joined away: an unidentified walk-in is
still revenue, and an inner join would make this disagree with the
conversion report computed over the same rows.
No cursor, deliberately. A keyset cursor needs a monotonic
server-assigned column and purchases has none; ordering by
(occurred_at, id) with a random uuid tie-break is exactly the shape that
silently dropped four of six simultaneous visits from the arrivals feed
before visits.seq existed. Offering one here would imply a delivery
guarantee this table cannot make, so the list is bounded by the date
window and a limit - which is how a sales list is browsed anyway.
GET /api/dashboard/summary. Four calls a client had to make and then
combine, which is how the desktop Footfall screen once produced its
headline by adding the daily bars up: silently too high, because a
customer who came twice is one person and two bucket-visitors. The
combining happens here, against Footfall and SiteHealth rather than new
SQL - a second definition of "unique visitor" or of "online" drifts, and
a home screen that disagrees with the report it links to is the one
nobody trusts afterwards. fraction_below_gate travels with the count for
the same reason it does everywhere else: it is what says whether the
headcount is a number or a floor.
Today is cut in the shop's timezone. In the one market this ships to,
UTC is five and a half hours wrong.
An unknown shop filter is a 400, not an ignored parameter. This API has
already been bitten once by a silently ignored filter handing back the
whole estate, which is a wrong number nobody would question.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Six read-only routes: merchant detail, its shops, one shop, its cameras,
one camera, and the platform totals. The console drills down
merchant -> store -> camera and every level below the first showed
'Backend integration required'.
They cannot be the tenant routes, and the reason is structural rather
than incidental. Every tenant handler derives the client from the
SESSION - that is what makes cross-tenant access impossible rather than
merely disallowed - and a platform admin has no client at all. The three
workarounds each make it worse: passing a company id to a tenant route
puts a caller-chosen tenant back in the one place this system refuses to
take one, filtering the estate in the browser ships every merchant's
data to render one, and signing in as the owner audits the wrong person.
So the tenant STORE functions are reused with an explicit client id -
they already take one - and the scoping the tenant handlers get from the
session is done in the handler instead.
AdminCamera is a separate type from Camera, for the same reason
AgentCamera is. It cannot carry host, port, path, username or
has_password. A tenant seeing those for their own camera is correct; a
platform admin browsing another company's estate is a different
question, and an RTSP host with a username beside it is most of a live
path into a customer's camera. Blanking fields on a shared struct leaves
'remember to redact, on every path, forever' as the only thing
preventing a leak. The test asserts on the raw JSON, because decoding
into the struct would discard exactly what it is looking for.
An unowned site is 404, never an empty list. The tenant resolver returns
a uuid untouched and lets client_id = downstream scope it, which is
sound only because that id comes from a session; here the caller names
both halves, so an unowned uuid would reach a query that quietly returns
nothing - 'this shop has no cameras' when the truth is 'not your shop'.
Both resolvers check the whole chain in one statement.
Two things the in-memory fake could not have caught, so neither was left
to it. The fake ignored clientID in SiteHealth and Cameras, which would
have made every cross-merchant test pass while returning another
company's shops; it is client-aware now for these paths. And the SQL was
written to make the documented $2-deduced-as-two-types bug impossible
rather than to be caught by a database later: id::text = $2 in place
of id = $2::uuid, one type per parameter, which also turns a malformed
path segment into the 404 it should be instead of a cast error.
Every read below the merchant list writes an audit row naming the admin
and the merchant - an admin is the one account for which nothing else
here leaves a trace. The counts-only summary does not: a console
refreshes it on a timer, and logging that buries the reads worth
finding. A suspended merchant stays readable, because that is precisely
what an admin opens the console to look at.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Every other thing in this product a person refers to already had a
readable reference: a shop is chennai, a camera cam1, a customer V-42, a
person their email. An audit of every list response found exactly one
gap, and it was the row people look at most - the arrivals feed showed a
visit as 36 hex characters.
012 argued no route takes a visit id so none was needed. That is true of
routing and false of everything else: it is what the feed shows, what a
support conversation quotes, and what somebody reading an API response
judges the product by.
Migration 014 mirrors the visitor scheme exactly - per client, so it
discloses no platform-wide volume, and beside the uuid rather than
instead of it. A stored counter is affordable on the busiest table
because visits from one tenant are already serialised by the consumer's
SetOrderMatters(true), so it adds no contention that was not already
there. A derived reference was the alternative and does not work:
several people through one door share occurred_at to the microsecond,
which is the collision 004 exists to handle.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
GET /api/cameras read only site_id, while every other filtered endpoint
takes both spellings through siteParam. So ?site=chennai was not a
filter at all but an unknown query parameter, silently ignored, and the
caller got every camera in the tenant believing it had one shop's.
Found by using it: a setup script saw another shop's cameras, concluded
three shops already had theirs and created none; then a delete aimed at
a test shop removed the live Coimbatore entrance camera, which had to be
restored. This is exactly the hazard already recorded for site vs
site_id - the note existed, the handler was simply missed.
One line to fix, and a test that asserts the whole class rather than
this one route: both spellings must narrow, and only an absent filter
may return more than one shop.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The display name was always meant to be editable and the slug frozen;
until now neither had a way in. PATCH /api/sites/{site} takes a name
and a timezone (manager and above), DELETE removes an empty shop
(owner). The shop drawer in head office gets both, with the short name
shown read-only and the reason beside it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Suspend or reinstate a company (PATCH /api/admin/clients/{id}), reset
its owner's password (shown once), and delete it - and an owner can
remove a shop opened by mistake (DELETE /api/sites/{site}, empty only).
Suspension ends every session the company holds in the same
transaction: login and ingest already refused an inactive client, but a
live access token would have kept reading for up to twelve hours, so
'suspend' would have meant 'suspend some time tomorrow'. Deletion is
deliberately two steps - the company must already be suspended and the
request repeats the slug - because the data under it is biometric.
Face images go first (a storage failure aborts with nothing touched),
then the broker logins, then the rows by cascade.
Exercised against the local Postgres and broker: create, open a shop,
remove it (two plugin commands), refuse delete while active, suspend
(owner's token 401 immediately), reset, delete, zero rows left.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The last step of onboarding that needed a shell: provision site printed
a broker password and a person typed it into Mosquitto's passwd file on
the host - mounted read-only in the container, so the first attempt
failed silently and the password was re-rolled. No tenant could open a
second branch without us.
The server now drives Mosquitto's dynamic-security plugin over its own
broker login: POST /api/sites (owner) writes the row and the sealed
password, registers the login and a per-site role with literal topics
(the 2.0 plugin does not substitute %u - measured), and removes the row
again if the broker refuses, so a shop cannot exist in the database and
not on the broker. provision site goes through the same path. The
head-office Shops screen gets 'Open a new shop'.
broker-init converts the existing passwd file into the plugin's store
with every hash intact - PBKDF2-SHA512 both sides - so the cutover
re-claims no shop PC. Rehearsed locally: old logins keep working,
isolation holds, the health probe works, and a PC claiming a shop opened
through the API connects as that shop. run-local.sh now brings the
broker up the same way.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The flow this product is sold on is three tiers: the platform admin
registers a merchant, the merchant registers their sales staff, the
staff sign in on a phone. Tier 1 handed the new owner a password. Tier 2
could not - a manager could only mint an invitation code, which the
salesperson had to redeem themselves, on their own phone, choosing their
own password. Good practice, and no use to a manager setting somebody up
before their first shift with a card and a pen.
POST /api/team/members mirrors POST /api/admin/clients: generated
password unless one is given, returned exactly once, bcrypt-hashed on
the way in and not recoverable after. Same permission shape as an
invitation - manager and above, only an owner mints an owner, admin
refused - so a manager cannot do through one door what they are refused
at the other. The invitation path stays; it is the better one whenever
the salesperson has their phone.
POST /api/team/{id}/password is the everyday case on a shop floor:
they forgot it. It sets a new one AND revokes every session they hold,
in one transaction, because the other reason a manager resets a
password is a lost phone, and a reset that left that phone signed in
would look complete while fixing nothing. Tenant-scoped in the UPDATE
itself; another company's user id is 404, never 403. No self-service
and no reset-by-email, deliberately: a floor account often has no
mailbox anyone checks, and the person who can vouch for the salesperson
standing in front of them is their manager.
RandomPassword moves from a private helper in the store to auth, so the
admin path, the merchant path and the reset all mint the same 80-bit
credential - rather than someone later writing a shorter one for the
"less important" account.
Verified: eight handler tests, and two against a real Postgres for the
things a fake cannot see - the RETURNING list scans on a row with no
last_login_at, the tenant scope holds, and the sessions row is actually
revoked. The tenant cleanup from yesterday held throughout.
API.md now documents the chain with both paths, and the note saying a
merchant could not create a login directly is gone because it is no
longer true.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Asked of the row the feed actually returns.
site_id had a reference all along and the feed was not sending it. A
client could read the shop's NAME off an arrival and still had no way to
ask for that shop except by uuid - the exact gap the reference scheme
exists to close. site_slug now travels with it.
visit_id stays a uuid and needs no reference: no route takes it, it is a
key a client de-duplicates on because delivery is at-least-once, and
nobody says a visit id out loud.
The uuid in a face URL must STAY random. visit_faces.id is
gen_random_uuid() and a derived or sequential one would let somebody
walk a shop's customers by date - the same reason bucket keys are random
rather than derived from the event id. A readable identifier is right
for a customer and wrong for the thing that points at their photograph.
And seq is now json:"-". visits.seq is a plain bigserial, so it counts
every visit on the PLATFORM, and shipping it put the total footfall of
every customer we have on every row of every tenant's feed - the same
German-tank estimate that decided visitors.number had to be per client.
It was a convenience for "have I fallen behind", nothing ever read it,
and the cursor answers that without disclosing a number. The SSE event
id was never the raw value; it has always been the opaque cursor.
The one test that broke was reading seq back off the wire to assert the
cursor pointed at the last row of a burst. It asserts against the seeded
position now: the property is unchanged, and the test can no longer see
what a client cannot.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Every id in the schema is a uuid and stays one. What was wrong was
putting one in front of a person: RecordVisit named every new customer
'Visitor ' || left(id::text, 8), so the arrivals feed, the shop PC and
the mobile app all read "Visitor 3446ec35" - the string a shop assistant
reads to a colleague and types into a search box. label is a stored
column staff can overwrite and SearchVisitors matches on, so formatting
around it in a front end would have left the data wrong on three
surfaces.
Migration 012 adds a per-client visitors.number, taken from a counter on
clients with UPDATE ... RETURNING inside the visit transaction. Per
client rather than global: a global sequence would tell any customer who
signs up how many people the whole platform has ever seen, from their
own first visitor number. The backfill numbers existing rows by
first_seen_at and relabels only the eight-hex pattern the old statement
produced, so a human-typed name is never overwritten.
Three of the four things anyone addresses by URL already had a human
name and the API simply refused it - a site has a slug, a camera has the
id the engine knows it by. refs.go accepts either form anywhere an id is
taken; a uuid resolves with no lookup, so every URL a client already
stored keeps working.
- An ambiguous camera name resolves to nothing, never to a guess: two
shops may each have an "Office1" and acting on the first row would
edit the wrong shop's camera.
- 404 on a path, 400 on a query filter. /api/visits answered fine and it
was the filter that was wrong.
- site and site_id are both accepted everywhere now. They differed per
endpoint, and an unknown query parameter is silently ignored, so
getting it the wrong way round returned the whole estate.
- The search matches V-13, which is what the product now shows.
Two bugs found by running it rather than testing it:
- 'Visitor ' || $2::text beside number = $2 makes Postgres deduce two
types for one parameter and refuse the insert. It compiled and passed
every in-memory test; the first real database rejected it, along with
the existing face tests that share the path.
- The fallback avatar said "V1" for Visitor 13, Visitor 10 and Visitor
15 alike, and read as the V-1 reference for a fourth person. It shows
the number now. The prop is customerRef, not ref - React reserves
that name and it would never have arrived.
Verified on the live database and through the running API: 13 hex labels
became Visitor 1-13 in first-seen order, two typed names left alone, and
the same customer reachable by uuid, V-13 and 13.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
A tenant had exactly the users somebody had created with a command on the
server. That is not a missing screen: a shop with an owner and four staff
either shared one password or raised a ticket per person, and a phone app
for the shop floor could not exist while there was one account to sign in
as.
Registration is by invitation, never open signup - the same line already
drawn around creating a company. The code carries the address and the role
and the request carries only a password, so a code that gets forwarded
cannot become somebody else's account, and a staff invitation cannot be
redeemed as an owner. Single use lives in the UPDATE and the account is
created in the same transaction.
Deactivating a member revokes their sessions in that transaction too. An
access token lives twelve hours, so without it "remove their access"
removed it sometime tomorrow. The session list and revoke that go with it
are the benefit of opaque tokens the product had been paying for and never
collecting: nothing could say what was signed in, let alone stop one.
Face images now work on a deployment with no object storage, which was
every local install and every self-hosted site - the arrivals feed said
"not storing customer photos" for every customer forever, on the screen
whose whole job is to show a face. Bounded to one row per visitor, so it
grows with the customer base and not with footfall; the bucket stays
primary wherever one exists.
Image.auth says whether a URL needs the session, because a browser img
cannot load one that does, a mobile image view can, and a webview can do
neither - the desktop client resolves those to a data URI in Go.
Found by running it, not by tests:
* UPDATE ... RETURNING gives the value AFTER the update, so the prune
read back empty keys, deleted nothing, and the table grew with
footfall exactly as if it were not there. The fake agreed with either
version; only the live Postgres test caught it.
* Trusting only the auth flag broke every shop card, because Sites.jsx
rebuilt a partial snapshot object and dropped it. A relative URL is
now sufficient on its own.
* ago() renders a future time as "just now", so a code valid for a week
read "expires just now".
Verified live against real Postgres: invite, preview, escalation refused,
register into a session, replay 404, staff forbidden, device revoked and
401 at once, last owner refused, and a 92,405-byte camera JPEG stored,
served to its owner, 401 with no session, 404 to another tenant, and
rendered in a browser.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
I got this wrong first time. "Head office cannot show live video cheaply"
conflated TRUE VIDEO with SEEING THE CAMERA NOW, and only the first needs
WebRTC and a TURN server.
The shop PC is behind a router with no inbound route, so head office
cannot pull the engine's MJPEG. It can answer the agent's outbound
requests, which is the shape of everything else here: the server holds a
poll open, the agent asks "is anyone watching?", and pushes JPEGs up for
exactly as long as somebody is.
Measured on the office camera: 98 KB full frame, 20.8 KB re-encoded at
640/q60, so one watcher costs ~83 KB/s. 47 frames arrived in 12 seconds -
4 fps, as configured. The UI says "about 4 frames a second" rather than
letting anyone conclude the camera stutters.
Nothing is uploaded when nobody is looking, which is the whole cost
argument: Publish returns false once the last viewer goes, interest lapses
on a timer each viewer refreshes as it reads (so a closed tab stops the
upload within seconds), one push is capped at five minutes, and the UI
streams one camera at a time.
LiveHub is deliberately the opposite of the arrivals Hub. There a doorbell
pushes nothing because nothing may be lost; here a dropped frame is the
correct outcome, so each viewer has a one-slot buffer that is overwritten -
the only frame worth having is the newest, and a queue would show an
ever-growing delay behind the shop instead of dropping back to live.
Ownership is proved once, before anything streams: the relay is keyed on a
camera id, a hub does not know whose camera it holds, and a camera id is
not a secret. Verified: another tenant gets 404, no session gets 401, and
an agent cannot push into another site's camera.
Also fixes a bug I introduced with it - the Live button was gated on
`connected`, which is head office's last report and up to two minutes
stale, so it hid itself during every reconnect. "Is that camera really
down?" is exactly when somebody wants to look, and a hidden control says
"you cannot" where the honest answer is "here is why".
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Head office shows a camera's latest frame rather than live video, for a
reason that has not changed: the engine serves MJPEG on 127.0.0.1 on a PC
behind a shop's router with no inbound route, and relaying it needs
WebRTC/TURN. Pointing a browser straight at the shop PC is not the escape
either - the engine's API is Basic-authenticated with a credential it
generates locally and never sends anywhere, and shipping that to the
cloud so a web page could use it would put the key to the biometric API
and the live face feed in the server's database.
But that picture only worked if you had an S3 bucket. Without one,
attachSnapshots reported "This system is not storing images" for every
camera forever - on the two screens whose whole job is to show the
camera. Making them picture-led turned a missing feature into a wall of
empty tiles, on every local install and any self-hosted customer who does
not want a bucket.
migrations/009 adds camera_snapshots and the agent falls back to
PUT /api/agent/cameras/{camera}/snapshot when the presigned route answers
images_disabled - chosen by sentinel, never by matching the message, since
it picks between two routes. One row per camera is what makes this safe in
the database when face images are not: the key IS the camera, so storage
is (cameras x ~100 KB) and does not grow with footfall.
The read is session-authenticated rather than a signed link, which an
<img> cannot use - hence Shot.jsx and useAuthedImage, keyed on the URL
string rather than the snapshot object so a poll does not re-fetch 90 KB
per camera every few seconds, and revoking the object URL on cleanup.
Verified against the real office camera with no bucket configured: 90,587
bytes stored in Postgres, served as image/jpeg to a signed-in user, 401
without a session, rendered on both the Cameras and Shops cards.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn
Five components that ship as one product:
- behavision/ the recognition engine. RTSP ingest, YuNet detection, IoU
tracking, ArcFace embeddings, a FAISS/SQLite gallery, and a
FastAPI dashboard. Identity is decided once per TRACK from an
average of at least three embeddings, never per frame.
- agent/ the Go edge agent: supervises the engine, holds a durable
spool, and drains it to MQTT. Nothing is acked before the
broker confirms.
- desktop/ the shop PC application (Wails + React + tray).
- server/ the cloud API, MQTT consumer, reports and assistant.
- web/ platform.loyaly.ai, the head-office app, embedded in the
server binary.
The gallery stores 512-float embeddings and timestamps - no images unless
`app.store_faces` is switched on. Those embeddings are biometric personal
data under GDPR and India's DPDP: template inversion reconstructs a
recognisable face from an ArcFace vector, so data/behavision.db is treated
as a biometric database and DELETE /api/visitors/{id} is a real erasure.
CLAUDE.md carries the reasoning behind every non-obvious decision here,
including the ones that were measured and the ones that were wrong first.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HViLj9gYNRtSr7YVZmW5sn