Stores: each shop's health in one line (offline, losing visits, camera
trouble, faces too poor - in severity order, one verdict), its cameras
as pictures with connection and 'proven to recognise a face' as two
different claims, add/edit/remove camera with the make picker, test
connection and placement checks claimed by the shop PC, and the
enrolment code a new shop PC types. Team: create an account (password
shown once), invite (code shown once), change role, remove access, reset
password, revoke pending invitations - with the rank rules the server
enforces mirrored in what the form offers. Account: who you are and
every device signed in, with 'sign out' per device and everywhere else.
Gone: StoreManagement, SecurityManager and ProfileForm that rendered
hard-coded arrays, the Business form with no backend, the /api/stores
route nothing served, and the cascading account menu's dead links.
Two things found by using the camera form, not by tests: Chrome filled
the operator's own email into 'camera username', and the first camera
saved with a password and no username because autofill wrote to the
input without React seeing it. The dialog sets autocomplete on the real
inputs and reads the credential fields from the DOM at submit.
Every route was exercised against production before this commit.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
The console defaulted to a backend on localhost, and features were built
against a locally modified server that production never had: Floor,
Commerce and their sales/customers routes answered 404 the day they were
deployed. The platform API is now https://mcp.loyaly.ai in every
environment; LOYALY_API_BASE remains only as an explicit override.
Removed what had no server behind it - Floor, Commerce, Lyts,
Leaderboard, and the Roles, Notifications, Billing, Integrations, API
keys and Preferences settings pages, all of which rendered hard-coded
arrays as if they were the merchant's data. Navigation is what the API
can honestly back. The removed code is in history if a real backend for
any of it is ever built.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KGcjxF1cNLcuwc3DAPcnfj
Two surfaces a merchant currently has no way to reach, and the BFF routes
behind them.
/floor is the only screen in this console aimed at somebody standing behind
a counter: who walked in, who is being served, and by whom. A visit is
claimed with attend, handed back with release, and closed with complete.
Claiming is single-winner — the platform decides, and a second person
tapping the same customer gets 409 CUSTOMER_ALREADY_TAKEN rather than a
silent overwrite. That was verified against a live platform: eight
concurrent claims, exactly one winner.
Naming a walk-in posts to /api/customers, which is the same permission as
PUT /api/visitors/{id}/profile upstream — attaching a name and a phone
number to a face is the floor's job, not a manager's.
Sale entry sits on the floor because that is where a sale happens. Lines
carry an intent, purchased or enquired, so a shop can record what somebody
asked about and did not buy; only purchased lines are billed. Money is
integer paise end to end and formatPaise is the only place it becomes
rupees — a float round-trip through the BFF was rejected once already and
must not come back.
Every write carries an idempotency key minted per dialog, so a double tap,
a timeout retry and a resubmit collapse into one sale rather than three.
Proven live: a replayed key returns 200 already_processed with the original
sale id.
/commerce gains the real list of those sales, replacing nothing invented —
it reads GET /api/sales and opens a detail view per row.
── What this does NOT do ────────────────────────────────────────────────
No mock, demo or placeholder data anywhere in it. Every figure comes off a
payload; an empty floor renders an empty state and says so.
── Known: the deployed backend does not serve these yet ─────────────────
/api/floor/visits, /api/customers, /api/sales and the three visit actions
all answer 404 on mcp.loyaly.ai today, which runs a build older than this
repository's first commit. Until that backend ships, /floor and the sales
panel will show error states, and the Floor nav entry points at a page that
cannot load its data. Committed deliberately so the two halves can be
deployed together rather than drifting further apart.
staff -> /floor has been the committed destination since 697b0d9; this is
the page it was always pointing at.
Verified: tsc clean, production build clean, lint unchanged at the existing
baseline. Exercised against a live local platform for all three roles —
attend/release/steal-prevention, customer creation, sale creation and
idempotent replay, and two-tenant isolation.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0161AMotQ8FxGPZ9gFGb5wiK
Production answered 502 on every url, /favicon.ico included, because the
container had no AUTH_SECRET and the boot check called process.exit(1): the
container died, so Traefik had no upstream and the one line explaining it was
trapped inside a restart-looping container. The exit is already gone (0dc865b).
This makes the configuration contract itself hard to get wrong.
One required variable, one resolver, two probe endpoints.
shared/config/authSecret.ts is now the only place the signing secret is
resolved. sessionToken.ts (HMAC of the identity cookie) and tokenStore.ts
(AES-256-GCM key for the platform token bundle) each read process.env
independently before, under rules that disagreed — one accepted a
whitespace-only value the other rejected. It accepts AUTH_SECRET, or
AUTH_SECRET_FILE for the Docker/Swarm secret convention when a dashboard field
mangles a value, trims both, and is a pure function of the environment.
That purity is load-bearing. Next 16 compiles proxy.ts for the NODE runtime
(its own docs: "Proxy defaults to using the Node.js runtime"), confirmed in the
build output — the proxy is in .next/server/chunks, not .next/server/edge. But
the proxy entry and the route entries are still SEPARATE BUNDLES with their own
copy of this module, so a secret invented in module scope would differ between
them, the proxy would reject every cookie the login route signed, and /login
would redirect forever. There is no generated fallback and there must not be.
LOYALY_API_BASE is no longer required in production. It accepted exactly one
origin, so an unset value could never have meant another, and requiring it added
a failure mode without adding a choice. Verified against the installed @next/env:
a real variable set to the EMPTY STRING is left empty and .env is NOT consulted,
so one blank dashboard field defeated the value shipped in the image and took
production down with "required in production". Any other host set explicitly is
still rejected by name, platform.loyaly.ai included.
/api/health and /api/ready are split. Health was returning 503 on a
configuration fault — readiness semantics on the name every orchestrator probes
by default. A Dockerfile HEALTHCHECK pointed there for one commit, and because
Dokploy runs applications as Swarm services, Swarm removed the task from the
load balancer and rescheduled it: the container was up, serving a 503 that named
the fault, and nothing could reach it to read that 503. Health is now liveness
and always 200 while the process answers; ready is readiness and 503 while a
variable is missing, for a DEPLOY gate (Order start-first + FailureAction
rollback) where failing keeps the previous good task serving. No HEALTHCHECK is
reintroduced.
Diagnostics answer the question that could not be answered from outside the
container: whether the variable never arrived or arrived empty, the secret's
source and length (never its value), and any environment variable whose NAME is
a near-miss for AUTH_SECRET — wrapped (NEXT_PUBLIC_AUTH_SECRET) or mistyped
(AUTH_SECERT, via bounded edit distance). Dokploy's Build Arguments and Build
Secrets are build-time only and absent at runtime, which from inside the
container is indistinguishable from never setting it; the boot log now tells
those apart.
Dockerfile, .env and .env.example changes are comments only — every directive
and every variable value is byte-identical to before.
Verified on the standalone payload the image ships: absent / empty / whitespace
/ typo'd name / wrong API host all keep the container ALIVE and answering 503
with x-loyaly-config: misconfigured; a valid secret gives / 307, /login 200,
/favicon.ico 200, health 200, ready 200. Cross-bundle auth, for both sources: a
cookie signed with the live secret is accepted by the proxy bundle (200) and
independently re-verified by the app/layout.tsx render bundle, while one signed
with a different secret is rejected by both (307). tsc --noEmit clean, eslint
clean on changed files, production build exit 0.
This does not by itself end the outage: AUTH_SECRET still has to be set on the
container, in Dokploy's runtime Environment Variables panel.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The HEALTHCHECK added an hour ago recreated the exact symptom. Dokploy runs
applications as Docker Swarm services, and Swarm does not merely report an
unhealthy task — it removes it from the service load balancer and reschedules
it. /api/health answers 503 while a required variable is missing, so:
AUTH_SECRET unset -> /api/health 503 -> task unhealthy -> pulled out of the
load balancer -> Traefik has no backend -> 502 Bad Gateway on every url.
The container was up and serving a 503 that names the fault, and nothing could
reach it to read that 503. "A broken deploy must not look healthy" is a real
concern, but enforcing it in the orchestrator destroys the diagnostics, and an
outage whose reason cannot be seen is the more expensive failure. The container
now stays in rotation whenever it can serve HTTP at all.
Also, two things that make a secret that was SET look like one that was not:
- Recommend `openssl rand -hex 32` everywhere instead of `openssl rand -base64
48`. A base64 value ends in '=' and may contain '+' and '/'; pasted into a
dashboard field or a KEY=VALUE editor that splits on the first '=', it can be
stored truncated or empty, which is indistinguishable from never setting it.
Hex is [0-9a-f] only, so there is nothing for a parser to mangle.
- Treat a whitespace-only AUTH_SECRET as missing, and print the secret's LENGTH
(never its value) in the boot log. `openssl rand -hex 32` is 64 characters, so
a much shorter number there is a value that arrived truncated — which
otherwise presents as sessions that do not verify, with nothing to explain it.
Verified on the rebuilt standalone payload: absent -> container alive, 503
`x-loyaly-config: misconfigured` on /, /login and /api/sites; whitespace -> the
same; a real hex secret -> / 307, /login 200, /favicon.ico 200, /api/health 200
and `auth secret set (64 chars)` in the log. tsc --noEmit and eslint clean,
production build exits 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A container started without AUTH_SECRET called process.exit(1) from the boot
check. The container died, Dokploy's Traefik had no upstream to proxy to, and
every url answered 502 Bad Gateway — /favicon.ico first, which is the line that
shows up in a browser console. The one message that explained it was on stderr
inside a restart-looping container, so the fastest fault in this app to fix
became the slowest to identify.
Reproduced against the real standalone payload: without AUTH_SECRET the process
exited 1; with it, /favicon.ico 200, /login 200, / 307.
The container now boots and stays up. While a required variable is missing the
proxy answers 503 with `x-loyaly-config: misconfigured` on every gated request,
and the boot log names the variable. A Docker HEALTHCHECK against the new
/api/health keeps the property the exit was protecting — a broken deploy still
reports unhealthy rather than presenting itself as a working one.
- configCheck.ts: one runtime-neutral check shared by the boot log, the proxy
and the health route, memoised so a healthy server pays an array-length read
per request rather than re-reading the environment.
- proxy.ts: the config gate runs before the session gate. verifySessionToken
reads AUTH_SECRET and throws ConfigError without one, which Next turns into a
500 per request — a status that says "this server has a bug" for a server that
is merely unconfigured.
- Neither the 503 body nor /api/health names the missing variable. Those
messages are operator information (configError.ts states the rule, the login
route already follows it); the names go to the container log.
- HEALTHCHECK probes with node, already the entrypoint, so it adds no package
and cannot break because a base image dropped a busybox applet.
- nginx.conf: marked dead. Nothing has installed nginx since 28258b5 and
.dockerignore keeps it out of the build context, but it is the first place
anyone looks at a 502 and the wrong one.
Verified: tsc --noEmit clean, eslint clean, and the Docker builder stage's
`NODE_ENV=production CI_BUILD=1 next build` exits 0. Misconfigured -> 503 on
/login, /dashboard, /api/sites with healthcheck exit 1; configured -> 200/307
with healthcheck exit 0.
This does not by itself end the outage: AUTH_SECRET still has to be set on the
container (Dokploy -> Environment). It makes the next occurrence legible.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three misconfigurations in a row were each diagnosed the same slow way: deploy,
try to sign in, read a status code, guess, read a response body, guess again.
That is a property of the design, not of bad luck. Every required variable is
read lazily, on the first request that needs it, so a container missing its
environment starts clean, serves the sign-in page, and passes a healthcheck
while being completely unable to authenticate anybody.
The lazy reads have to stay — reading at module scope fails `next build`, which
collects page data with NODE_ENV=production and none of these variables set
(that is the "Failed to collect page data for /api/assistant" failure already
documented in apiClient). So the check goes in an instrumentation hook instead,
which Next runs once per server start and NEVER during a build: it returns early
when NEXT_PHASE is 'phase-production-build', in
server/lib/router-utils/instrumentation-globals.external.js.
Boot now either prints what it resolved:
[loyaly] config ok — platform https://mcp.loyaly.ai, auth secret set, NODE_ENV=production
or refuses to start, naming every problem at once rather than one per deploy:
refusing to start — 2 configuration problem(s):
1. LOYALY_API_BASE is invalid: https://api.example.com is not a supported production API host...
2. AUTH_SECRET is required in production — it signs the session cookie...
Verified against a real standalone build, booted four ways: unconfigured, fully
configured, LOYALY_API_BASE pointed at the console, and both wrong at once.
The platform origin and its validation move to shared/config/platformApi. This
is load-bearing, not tidying: apiClient is `server-only`, and that package
resolves to a module which THROWS ON IMPORT outside a react-server condition —
which the instrumentation bundle is not. Importing apiClient from the boot check
would have crashed every start, correctly configured or not. The extracted
module imports nothing but ConfigError, so the boot check and the request path
run the same function against the same allowlist.
Also corrects a claim I made in the Dockerfile two commits ago. `next build`
DOES copy .env into .next/standalone — writeStandaloneDirectory takes exactly
.env and .env.production and nothing else — so the explicit COPY is redundant
rather than required. It stays, with an accurate reason: it only happens when
.env is in the build context, and .dockerignore excluded it until recently.
Keeping the COPY makes that dependency fail the Docker build loudly instead of
producing an unconfigured image.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`storeTokens()` and `createSessionToken()` ran outside any catch. Both read
AUTH_SECRET — one derives the AES key that encrypts the platform bundle, the
other signs the identity cookie — and in production both refuse to fall back to
the development key. So an unset AUTH_SECRET threw after the credentials had
already been accepted upstream, and the browser got a bare 500 on a sign-in
that was entirely valid.
The status was the smaller half. The upstream session minted moments earlier by
`authApi.login` was ORPHANED: a live refresh token, issued to somebody who did
not end up logged in, left to expire on its own. The platform-admin branch a few
lines above already revokes for precisely this reason — declining because the
server is broken is no different from declining because the account is wrong —
so this now revokes too, best-effort, on the same terms.
It fails closed. No cookie is set on this path, so a half-configured server
cannot hand out a session it will be unable to verify on the next request.
ConfigError moves to src/shared/errors/configError.ts because its throwers now
span two runtimes: apiClient and tokenStore are server-only, while sessionToken
is reached from src/proxy.ts, which Next compiles for Edge. Declaring it in
apiClient would have dragged the whole platform client, `server-only` guard and
all, into the proxy bundle to name one class. The new module imports nothing.
Verified by exercising both functions directly: with AUTH_SECRET unset under
NODE_ENV=production, sealTokens and createSessionToken each raise ConfigError
rather than a bare Error; with it set, both succeed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`POST /api/auth/login` answered 502 platform_unreachable — "Could not reach
Loyaly. Check your connection and try again." — for a fault that is entirely
ours and that no connection check can fix.
`buildUrl()` is where LOYALY_API_BASE is read, and it was called INSIDE the
try block whose catch turns a failed fetch into UpstreamError(0, 'network').
So the config guard threw, the catch swallowed it, and "nobody set
LOYALY_API_BASE" arrived at the route indistinguishable from "the platform is
down". Measured on the pre-fix code, all of these produced the identical
UpstreamError(status=0, code=network):
LOYALY_API_BASE unset
LOYALY_API_BASE=https://platform.loyaly.ai (the known-wrong host)
LOYALY_API_BASE=not-a-url
nothing listening on the far end (the only real network failure)
The message survived, so the truth was reachable, but only by reading the
prose of an error the code had already classified as a network fault — and
the login route had by then replaced it with advice about the user's wifi.
ConfigError now exists for this, buildUrl is resolved before the try in both
upstreamRequest and upstreamRaw, and callers branch on it: 500 misconfigured,
not 502 unreachable. 500 is the honest status — a bad gateway says the thing
upstream is unwell, and this server has not got as far as having an upstream.
Verified after the change: the four config faults raise ConfigError, and a
dead port still raises UpstreamError(0, 'network').
The detail is logged, never returned. It names an environment variable and the
hosts this console accepts, which belongs in the Dokploy log pane rather than
in an anonymous sign-in form's response body. `[loyaly] configuration error:`
is now the line to grep for.
Also fixes failJson labelling a 5xx as `unauthorized` in the envelope: the
code is derived from the status now, so a misconfigured server can no longer
tell a browser the password was wrong. That one costs somebody a password
reset for a fault they cannot see.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`GET /api/sites 401 (Unauthorized)` fired on the sign-in page for every
visitor who had not signed in yet.
WorkspaceProvider sits above the whole route tree, /login included, and it
already carried a comment saying the estate was "gated on the session ...
asking for the estate before anyone has signed in would put a guaranteed 401
in the console on every visit to the sign-in page". The gate was never
applied. `useSites()` was called unconditionally and `isAuthenticated` only
guarded the derived `stores` value below it — and a gate on a derived value
cannot hold a fetch that has already gone out.
useResource has taken `Endpoint<T> | null` for exactly this all along; the
effect returns early on a null key, so nothing is sent. The gate goes in
useSites rather than in one consumer because the rule belongs to the endpoint
— no session, no estate — and /stores calls it too.
Costs a signed-in user nothing: SessionProvider resolves `status` synchronously
from the server-rendered `initialSession`, so there is no 'loading' pass to
wait through before the request goes out. /stores is unaffected in the other
direction too — AuthGuard renders a spinner instead of children once status is
'unauthenticated', so no consumer sits on a permanently held resource.
Verified in the browser against a clean network buffer: with the gate reverted,
/login issues GET /api/sites → 401; with it in place, /login issues 31 requests
and none of them are /api/*.
Worth being explicit about what this does NOT change: platform.loyaly.ai/api/*
is the correct address for these calls. It is this app's own BFF, same-origin
by design, and the hop to mcp.loyaly.ai happens server-side where the token
lives. The 401 was a request that should never have been made, not a request
made to the wrong host.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>