Production answered 502 on every url, /favicon.ico included, because the
container had no AUTH_SECRET and the boot check called process.exit(1): the
container died, so Traefik had no upstream and the one line explaining it was
trapped inside a restart-looping container. The exit is already gone (0dc865b).
This makes the configuration contract itself hard to get wrong.
One required variable, one resolver, two probe endpoints.
shared/config/authSecret.ts is now the only place the signing secret is
resolved. sessionToken.ts (HMAC of the identity cookie) and tokenStore.ts
(AES-256-GCM key for the platform token bundle) each read process.env
independently before, under rules that disagreed — one accepted a
whitespace-only value the other rejected. It accepts AUTH_SECRET, or
AUTH_SECRET_FILE for the Docker/Swarm secret convention when a dashboard field
mangles a value, trims both, and is a pure function of the environment.
That purity is load-bearing. Next 16 compiles proxy.ts for the NODE runtime
(its own docs: "Proxy defaults to using the Node.js runtime"), confirmed in the
build output — the proxy is in .next/server/chunks, not .next/server/edge. But
the proxy entry and the route entries are still SEPARATE BUNDLES with their own
copy of this module, so a secret invented in module scope would differ between
them, the proxy would reject every cookie the login route signed, and /login
would redirect forever. There is no generated fallback and there must not be.
LOYALY_API_BASE is no longer required in production. It accepted exactly one
origin, so an unset value could never have meant another, and requiring it added
a failure mode without adding a choice. Verified against the installed @next/env:
a real variable set to the EMPTY STRING is left empty and .env is NOT consulted,
so one blank dashboard field defeated the value shipped in the image and took
production down with "required in production". Any other host set explicitly is
still rejected by name, platform.loyaly.ai included.
/api/health and /api/ready are split. Health was returning 503 on a
configuration fault — readiness semantics on the name every orchestrator probes
by default. A Dockerfile HEALTHCHECK pointed there for one commit, and because
Dokploy runs applications as Swarm services, Swarm removed the task from the
load balancer and rescheduled it: the container was up, serving a 503 that named
the fault, and nothing could reach it to read that 503. Health is now liveness
and always 200 while the process answers; ready is readiness and 503 while a
variable is missing, for a DEPLOY gate (Order start-first + FailureAction
rollback) where failing keeps the previous good task serving. No HEALTHCHECK is
reintroduced.
Diagnostics answer the question that could not be answered from outside the
container: whether the variable never arrived or arrived empty, the secret's
source and length (never its value), and any environment variable whose NAME is
a near-miss for AUTH_SECRET — wrapped (NEXT_PUBLIC_AUTH_SECRET) or mistyped
(AUTH_SECERT, via bounded edit distance). Dokploy's Build Arguments and Build
Secrets are build-time only and absent at runtime, which from inside the
container is indistinguishable from never setting it; the boot log now tells
those apart.
Dockerfile, .env and .env.example changes are comments only — every directive
and every variable value is byte-identical to before.
Verified on the standalone payload the image ships: absent / empty / whitespace
/ typo'd name / wrong API host all keep the container ALIVE and answering 503
with x-loyaly-config: misconfigured; a valid secret gives / 307, /login 200,
/favicon.ico 200, health 200, ready 200. Cross-bundle auth, for both sources: a
cookie signed with the live secret is accepted by the proxy bundle (200) and
independently re-verified by the app/layout.tsx render bundle, while one signed
with a different secret is rejected by both (307). tsc --noEmit clean, eslint
clean on changed files, production build exit 0.
This does not by itself end the outage: AUTH_SECRET still has to be set on the
container, in Dokploy's runtime Environment Variables panel.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
80 lines
4.5 KiB
Bash
80 lines
4.5 KiB
Bash
# ---------------------------------------------------------------------------
|
|
# Production runtime configuration. COMMITTED ON PURPOSE — carries no secret.
|
|
# ---------------------------------------------------------------------------
|
|
#
|
|
# This file is the production environment. It is read by `next build` and, more
|
|
# importantly, by the standalone `server.js` at boot (Next calls loadEnvConfig
|
|
# on the server's working directory), so the deployed container knows the
|
|
# platform host without anyone remembering to type it into a dashboard.
|
|
#
|
|
# ── Precedence, exactly as @next/env resolves it ────────────────────────────
|
|
#
|
|
# 1. real process.env (Dokploy / docker -e / systemd) ← always wins
|
|
# 2. .env.production.local
|
|
# 3. .env.local ← LOCAL DEV ONLY. Never enters the image.
|
|
# 4. .env.production
|
|
# 5. .env ← this file, the floor everything falls back to
|
|
#
|
|
# A value already present in process.env is never overwritten by a file, so
|
|
# setting LOYALY_API_BASE in Dokploy still overrides this — nothing here locks
|
|
# the deployment in. It only removes "unset" as a possible state.
|
|
#
|
|
# ── Working on this locally? ────────────────────────────────────────────────
|
|
# Put your overrides in `.env.local` (gitignored, loaded ahead of this file).
|
|
# Without one, `npm run dev` will talk to the PRODUCTION platform, because that
|
|
# is what this file says. `.env.example` has the local values to copy.
|
|
|
|
# The one shared Loyaly platform API (Behavision). Server-side only and
|
|
# deliberately NOT NEXT_PUBLIC: publishing the host would let a browser bypass
|
|
# the BFF, which is what keeps the access token out of JavaScript.
|
|
#
|
|
# NOT platform.loyaly.ai — that host serves THIS console, not the API. Pointing
|
|
# the variable there makes the BFF call its own origin, which fails in a way
|
|
# that looks like a broken login form rather than a misconfiguration.
|
|
# apiClient.ts rejects that hostname by name for exactly this reason.
|
|
#
|
|
# NOT REQUIRED in production any more. Production accepts exactly one origin, so
|
|
# an unset variable could never have meant another one, and platformApi resolves
|
|
# it to that origin on its own. It stays here so `docker run` is self-describing
|
|
# and so development has something to read.
|
|
#
|
|
# Why that change was needed: @next/env only fills a variable that is ABSENT.
|
|
# Verified against the installed copy — a real environment variable set to the
|
|
# EMPTY STRING stays empty and this file is NOT consulted. So one blank field in
|
|
# a dashboard silently defeated the value below and took production down with
|
|
# "LOYALY_API_BASE is required in production".
|
|
LOYALY_API_BASE=https://mcp.loyaly.ai
|
|
|
|
# Browser → this app's own BFF routes, which are same-origin. Empty is correct
|
|
# and is what makes the console work on any hostname it is served from:
|
|
# requests go to /api/... on whatever origin loaded the page (localhost:3100 in
|
|
# dev, platform.loyaly.ai in production) and the server hop above reaches the
|
|
# platform. Setting this to the platform host would send the browser straight
|
|
# at the API with no session cookie and no token — do not.
|
|
#
|
|
# It is NEXT_PUBLIC, so it is inlined at BUILD time, not read at runtime.
|
|
# Changing it in Dokploy's environment panel would do nothing without a rebuild.
|
|
NEXT_PUBLIC_API_BASE=
|
|
|
|
# AUTH_SECRET is deliberately NOT in this file. It is the ONLY variable this
|
|
# deployment requires, and the only one that cannot ship.
|
|
#
|
|
# It signs the session cookie and encrypts the platform token bundle, so a
|
|
# value committed here is a session-forging key in git — anyone who can read
|
|
# the repo could mint a cookie for any user. It was already removed from the
|
|
# Dockerfile once for that reason; do not reintroduce it here.
|
|
#
|
|
# Set it as a Dokploy environment variable in the RUNTIME panel — a value set as
|
|
# a BUILD argument is not present when the server runs, which looks exactly like
|
|
# never having set it. Alternatively mount the value and set AUTH_SECRET_FILE to
|
|
# its path (the Docker/Swarm secret convention); AUTH_SECRET wins if both exist.
|
|
#
|
|
# Production refuses to sign sessions without it. Generate with:
|
|
#
|
|
# openssl rand -hex 32
|
|
#
|
|
# Hex, not base64: a base64 value ends in '=' and can contain '+' and '/', and
|
|
# an environment editor that splits a line on the first '=' can store that
|
|
# truncated or empty. A silently-empty AUTH_SECRET looks exactly like an unset
|
|
# one, which is a slow afternoon. Hex has nothing a parser can mangle.
|