The HEALTHCHECK added an hour ago recreated the exact symptom. Dokploy runs
applications as Docker Swarm services, and Swarm does not merely report an
unhealthy task — it removes it from the service load balancer and reschedules
it. /api/health answers 503 while a required variable is missing, so:
AUTH_SECRET unset -> /api/health 503 -> task unhealthy -> pulled out of the
load balancer -> Traefik has no backend -> 502 Bad Gateway on every url.
The container was up and serving a 503 that names the fault, and nothing could
reach it to read that 503. "A broken deploy must not look healthy" is a real
concern, but enforcing it in the orchestrator destroys the diagnostics, and an
outage whose reason cannot be seen is the more expensive failure. The container
now stays in rotation whenever it can serve HTTP at all.
Also, two things that make a secret that was SET look like one that was not:
- Recommend `openssl rand -hex 32` everywhere instead of `openssl rand -base64
48`. A base64 value ends in '=' and may contain '+' and '/'; pasted into a
dashboard field or a KEY=VALUE editor that splits on the first '=', it can be
stored truncated or empty, which is indistinguishable from never setting it.
Hex is [0-9a-f] only, so there is nothing for a parser to mangle.
- Treat a whitespace-only AUTH_SECRET as missing, and print the secret's LENGTH
(never its value) in the boot log. `openssl rand -hex 32` is 64 characters, so
a much shorter number there is a value that arrived truncated — which
otherwise presents as sessions that do not verify, with nothing to explain it.
Verified on the rebuilt standalone payload: absent -> container alive, 503
`x-loyaly-config: misconfigured` on /, /login and /api/sites; whitespace -> the
same; a real hex secret -> / 307, /login 200, /favicon.ico 200, /api/health 200
and `auth secret set (64 chars)` in the log. tsc --noEmit and eslint clean,
production build exits 0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The deployed console called its own origin instead of the platform. The BFF
route would throw "LOYALY_API_BASE is required in production — refusing to
guess the Loyaly platform host", and from the browser that reads as a broken
login form rather than as a missing variable.
The guard was right; nothing ever set the variable. `.gitignore` had a blanket
`.env*` and `.dockerignore` excluded `.env` and `.env.*`, so the image carried
no environment at all and the only copy of the production host was a comment
in `.env.example`. Injecting it by hand at the orchestrator was the single
point of failure, and it failed.
The platform host is not a secret, so it is now committed in `.env` and copied
into the runner stage. `next build` does not fold `.env` into
`.next/standalone`, which is why the COPY is explicit; server.js chdirs to
/app and Next runs loadEnvConfig there, so the file sits beside it at the
WORKDIR root. `npm run bundle` stages it the same way for a non-Docker deploy.
This pins nothing. @next/env never overwrites a variable already present in
process.env, so anything set in Dokploy still wins — verified against
@next/env directly: a bare image resolves https://mcp.loyaly.ai, an injected
LOYALY_API_BASE overrides it, and a leaked .env.local beats both.
That last case is why `.dockerignore` still excludes `.env.*`. A developer's
.env.local points at http://127.0.0.1:8088 and loads AHEAD of .env, so one
leaking into the build context would make the deployed console call localhost
with no error to read. Confirmed the context now carries `.env` and nothing
else.
AUTH_SECRET stays out of every committed file and out of the image. It signs
the session cookie and encrypts the token bundle, so a committed value is a
session-forging key in git — the thing 8b3fbab removed from the Dockerfile.
It remains a Dokploy secret, and production still refuses to sign without it.
`.env.example` is now the template for `.env.local` rather than a second copy
of the production values, so the two files cannot drift.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Three configuration defects and one screen of invented people.
API base. The client defaulted to https://platform.loyaly.ai when
LOYALY_API_BASE was unset, and that host serves THIS console, not the
Behavision API - verified live: it answers /api/auth/me with the console's
own 404 HTML and a login POST with the console's own BFF envelope. So an
unset variable in production made the BFF call its own origin, which fails
looking like a broken login form rather than a misconfiguration. There is
now no remote fallback: development defaults to 127.0.0.1:8088 and
production throws, naming the variable, the way tokenStore already refuses
to run without AUTH_SECRET. A wrong host that appears to work is worse than
a startup failure that says what is missing.
The template. .env.example documented that same wrong host, and .gitignore's
`.env*` matched the template itself, so it was never committed - a fresh
clone got no template at all, for an app that cannot start in production
without AUTH_SECRET. Added `!.env.example` after the ignore rule; .env.local
and every other .env* stay ignored. The template carries placeholders only,
no values.
Team. /settings/team listed five invented people - aravind@nearle.in,
Vikram Seth, Priya Sharma - with store names no endpoint supplies and roles
that do not exist upstream, behind four controls that mutated local state
and were lost on refresh. A merchant could not tell any of it from the real
thing. It now reads GET /api/team, which the platform already serves and
scopes by session, and renders what actually comes back: name, email, role,
whether the account is still active, and last sign-in (or "Never", which is
a fact worth seeing).
The route used to map each row through toAuthUser, which reads client_name -
a field GET /api/team does not send - so organisation was undefined on every
row while active, last_login_at and created_at were discarded. ApiTeamMember
now describes that payload properly and ApiUser is left to authentication.
The screen is READ-ONLY on purpose. Accounts are born from invitations, and
that flow already exists in the platform's own web app: a manager mints a
code, the holder redeems it and chooses their own password. A second way to
create a login does not belong here, least of all on the screen that lists
them. Role changes and deactivation are supported upstream by
PATCH /api/team/{id} and are deliberately not wired: deactivating revokes
every session that person holds immediately, so it wants a confirmation step
and 409 last_owner handling, neither of which belongs in a change whose
purpose is removing invented data.
types.ts also gains the Sales/Floor/Customer interfaces. They are inert here
- nothing imports them yet - and land with this commit so the screens that
consume them arrive as one reviewable change.
Verified against the live local platform: two tenants, correct member lists
for each, and no cross-tenant leakage.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0161AMotQ8FxGPZ9gFGb5wiK