fix(deploy): harden runtime configuration
Production answered 502 on every url, /favicon.ico included, because the
container had no AUTH_SECRET and the boot check called process.exit(1): the
container died, so Traefik had no upstream and the one line explaining it was
trapped inside a restart-looping container. The exit is already gone (0dc865b).
This makes the configuration contract itself hard to get wrong.
One required variable, one resolver, two probe endpoints.
shared/config/authSecret.ts is now the only place the signing secret is
resolved. sessionToken.ts (HMAC of the identity cookie) and tokenStore.ts
(AES-256-GCM key for the platform token bundle) each read process.env
independently before, under rules that disagreed — one accepted a
whitespace-only value the other rejected. It accepts AUTH_SECRET, or
AUTH_SECRET_FILE for the Docker/Swarm secret convention when a dashboard field
mangles a value, trims both, and is a pure function of the environment.
That purity is load-bearing. Next 16 compiles proxy.ts for the NODE runtime
(its own docs: "Proxy defaults to using the Node.js runtime"), confirmed in the
build output — the proxy is in .next/server/chunks, not .next/server/edge. But
the proxy entry and the route entries are still SEPARATE BUNDLES with their own
copy of this module, so a secret invented in module scope would differ between
them, the proxy would reject every cookie the login route signed, and /login
would redirect forever. There is no generated fallback and there must not be.
LOYALY_API_BASE is no longer required in production. It accepted exactly one
origin, so an unset value could never have meant another, and requiring it added
a failure mode without adding a choice. Verified against the installed @next/env:
a real variable set to the EMPTY STRING is left empty and .env is NOT consulted,
so one blank dashboard field defeated the value shipped in the image and took
production down with "required in production". Any other host set explicitly is
still rejected by name, platform.loyaly.ai included.
/api/health and /api/ready are split. Health was returning 503 on a
configuration fault — readiness semantics on the name every orchestrator probes
by default. A Dockerfile HEALTHCHECK pointed there for one commit, and because
Dokploy runs applications as Swarm services, Swarm removed the task from the
load balancer and rescheduled it: the container was up, serving a 503 that named
the fault, and nothing could reach it to read that 503. Health is now liveness
and always 200 while the process answers; ready is readiness and 503 while a
variable is missing, for a DEPLOY gate (Order start-first + FailureAction
rollback) where failing keeps the previous good task serving. No HEALTHCHECK is
reintroduced.
Diagnostics answer the question that could not be answered from outside the
container: whether the variable never arrived or arrived empty, the secret's
source and length (never its value), and any environment variable whose NAME is
a near-miss for AUTH_SECRET — wrapped (NEXT_PUBLIC_AUTH_SECRET) or mistyped
(AUTH_SECERT, via bounded edit distance). Dokploy's Build Arguments and Build
Secrets are build-time only and absent at runtime, which from inside the
container is indistinguishable from never setting it; the boot log now tells
those apart.
Dockerfile, .env and .env.example changes are comments only — every directive
and every variable value is byte-identical to before.
Verified on the standalone payload the image ships: absent / empty / whitespace
/ typo'd name / wrong API host all keep the container ALIVE and answering 503
with x-loyaly-config: misconfigured; a valid secret gives / 307, /login 200,
/favicon.ico 200, health 200, ready 200. Cross-bundle auth, for both sources: a
cookie signed with the live secret is accepted by the proxy bundle (200) and
independently re-verified by the app/layout.tsx render bundle, while one signed
with a different secret is rejected by both (307). tsc --noEmit clean, eslint
clean on changed files, production build exit 0.
This does not by itself end the outage: AUTH_SECRET still has to be set on the
container, in Dokploy's runtime Environment Variables panel.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -1,50 +1,61 @@
|
||||
import {configStatus} from '@/shared/config/configCheck';
|
||||
|
||||
/**
|
||||
* GET /api/health — can this container serve?
|
||||
* GET /api/health — LIVENESS. "Is this process answering HTTP?"
|
||||
*
|
||||
* ── Why a container needs this ───────────────────────────────────────────
|
||||
* The boot check used to answer the same question by killing the process, on
|
||||
* the reasoning that a dead container is the only signal a platform cannot
|
||||
* misread. It is also the only signal a BROWSER cannot read: the reverse proxy
|
||||
* in front of it had nothing to connect to and returned 502 for every url,
|
||||
* which is what a missing AUTH_SECRET looked like from the outside.
|
||||
* Always 200 when the server can respond at all. It does NOT fail on a
|
||||
* configuration problem, and that is the entire point of separating it from
|
||||
* /api/ready.
|
||||
*
|
||||
* This is the half of that trade worth keeping. The container stays up and
|
||||
* explains itself, while the Dockerfile's HEALTHCHECK polls this route and
|
||||
* drives the container `unhealthy` when it answers 503 — so a misconfigured
|
||||
* deploy still cannot present itself as a working one.
|
||||
* ── Why the split exists ─────────────────────────────────────────────────
|
||||
* This route used to answer 503 while a required variable was missing, which
|
||||
* is READINESS semantics living on the name every orchestrator probes by
|
||||
* default. A Docker HEALTHCHECK was pointed at it for one commit, and because
|
||||
* Dokploy runs applications as Docker Swarm services, Swarm did not merely
|
||||
* report the task unhealthy — it removed it from the service load balancer and
|
||||
* rescheduled it. Traefik then had no backend and answered 502 Bad Gateway on
|
||||
* every url: the container was up, serving a 503 that named the fault, and
|
||||
* nothing could reach it to read that 503.
|
||||
*
|
||||
* ── What it deliberately does not say ────────────────────────────────────
|
||||
* Anonymous and public, so it publishes a state and a count, never the problem
|
||||
* messages: those name environment variables, which is operator information
|
||||
* (configError.ts states the rule; the login route follows it too). The names
|
||||
* are printed once at boot, in the container log, where only an operator sees
|
||||
* them.
|
||||
* Splitting the two makes that choice explicit instead of accidental. A probe
|
||||
* wired here can never remove a serving container from rotation. A probe wired
|
||||
* to /api/ready can, deliberately, and is the right thing for a DEPLOY gate
|
||||
* (Swarm Order=start-first + FailureAction=rollback) where failing keeps the
|
||||
* PREVIOUS healthy task serving.
|
||||
*
|
||||
* Exempt from the proxy's session gate — see HEALTH_PATH in src/proxy.ts — or
|
||||
* an unauthenticated healthcheck would read 401 as "unhealthy" on a perfectly
|
||||
* good container.
|
||||
* The body still reports configuration, so this one endpoint answers both
|
||||
* "is it alive?" and "why is it unhappy?" — it just never lies about the first
|
||||
* to signal the second.
|
||||
*/
|
||||
|
||||
export const dynamic = 'force-dynamic';
|
||||
|
||||
export function GET(): Response {
|
||||
const {problems} = configStatus();
|
||||
const healthy = problems.length === 0;
|
||||
const {problems, authSecretSource, authSecretLength} = configStatus();
|
||||
|
||||
return Response.json(
|
||||
{
|
||||
status: healthy ? 'ok' : 'misconfigured',
|
||||
// A count, not the messages. Enough to tell "one variable missing" from
|
||||
// "this container has nothing set at all" without publishing which.
|
||||
status: 'alive',
|
||||
/**
|
||||
* Reported, never enforced here. `ok` vs `misconfigured` tells a human
|
||||
* what is wrong without letting a probe tear the container down for it.
|
||||
*/
|
||||
configuration: problems.length === 0 ? 'ok' : 'misconfigured',
|
||||
problems: problems.length,
|
||||
/**
|
||||
* Source and length only — never the value. `env` / `file` /
|
||||
* `development` / `missing` plus a length distinguishes "not set" from
|
||||
* "set but truncated", which is the distinction that costs the most time
|
||||
* to make from outside a container. A length is not a meaningful
|
||||
* disclosure about a 256-bit random value.
|
||||
*/
|
||||
authSecret: {source: authSecretSource, length: authSecretLength},
|
||||
},
|
||||
{
|
||||
status: healthy ? 200 : 503,
|
||||
status: 200,
|
||||
headers: {
|
||||
'cache-control': 'no-store',
|
||||
'x-loyaly-config': healthy ? 'ok' : 'misconfigured',
|
||||
'x-loyaly-config': problems.length === 0 ? 'ok' : 'misconfigured',
|
||||
},
|
||||
},
|
||||
);
|
||||
|
||||
54
src/app/api/ready/route.ts
Normal file
54
src/app/api/ready/route.ts
Normal file
@@ -0,0 +1,54 @@
|
||||
import {configStatus} from '@/shared/config/configCheck';
|
||||
|
||||
/**
|
||||
* GET /api/ready — READINESS. "Can this container serve real traffic?"
|
||||
*
|
||||
* 200 only when every required variable is present and acceptable; 503
|
||||
* otherwise. Unlike /api/health (liveness), this one is MEANT to fail.
|
||||
*
|
||||
* ── What to point at it, and what not to ─────────────────────────────────
|
||||
* Point a DEPLOY gate here: Dokploy → Advanced → Swarm Settings, a Health
|
||||
* Check whose Test hits this path, with Update Config `Order: start-first` and
|
||||
* `FailureAction: rollback`. A newly deployed task that is missing AUTH_SECRET
|
||||
* then never becomes healthy, never replaces the running one, and is rolled
|
||||
* back — the previous good version keeps serving, and the broken deploy is
|
||||
* rejected before any production traffic reaches it. That is the contract this
|
||||
* endpoint exists for.
|
||||
*
|
||||
* Understand the one case it cannot save: if NO healthy task exists to fall
|
||||
* back to — a first deploy, or a service that is already broken — then a gate
|
||||
* here means nothing is in rotation and Traefik answers 502. A readiness gate
|
||||
* cannot invent a working version. Get production healthy FIRST, then turn the
|
||||
* gate on; it protects every deploy after that.
|
||||
*
|
||||
* Do NOT put this path in a Dockerfile HEALTHCHECK. That applies to the
|
||||
* container unconditionally, including when there is no predecessor, which is
|
||||
* exactly how a 502 was recreated once already. The deploy gate belongs in
|
||||
* Dokploy's Swarm settings, where start-first and rollback give it somewhere
|
||||
* safe to fail to.
|
||||
*/
|
||||
|
||||
export const dynamic = 'force-dynamic';
|
||||
|
||||
export function GET(): Response {
|
||||
const {problems, authSecretSource, authSecretLength} = configStatus();
|
||||
const ready = problems.length === 0;
|
||||
|
||||
return Response.json(
|
||||
{
|
||||
status: ready ? 'ready' : 'not_ready',
|
||||
// A count, not the messages: this is public and anonymous, and the
|
||||
// messages name environment variables, which is operator information.
|
||||
// The names are printed once at boot, in the container log.
|
||||
problems: problems.length,
|
||||
authSecret: {source: authSecretSource, length: authSecretLength},
|
||||
},
|
||||
{
|
||||
status: ready ? 200 : 503,
|
||||
headers: {
|
||||
'cache-control': 'no-store',
|
||||
'x-loyaly-config': ready ? 'ok' : 'misconfigured',
|
||||
},
|
||||
},
|
||||
);
|
||||
}
|
||||
Reference in New Issue
Block a user