fix(deploy): serve a 503 that names the fault instead of dying into a 502

A container started without AUTH_SECRET called process.exit(1) from the boot
check. The container died, Dokploy's Traefik had no upstream to proxy to, and
every url answered 502 Bad Gateway — /favicon.ico first, which is the line that
shows up in a browser console. The one message that explained it was on stderr
inside a restart-looping container, so the fastest fault in this app to fix
became the slowest to identify.

Reproduced against the real standalone payload: without AUTH_SECRET the process
exited 1; with it, /favicon.ico 200, /login 200, / 307.

The container now boots and stays up. While a required variable is missing the
proxy answers 503 with `x-loyaly-config: misconfigured` on every gated request,
and the boot log names the variable. A Docker HEALTHCHECK against the new
/api/health keeps the property the exit was protecting — a broken deploy still
reports unhealthy rather than presenting itself as a working one.

- configCheck.ts: one runtime-neutral check shared by the boot log, the proxy
  and the health route, memoised so a healthy server pays an array-length read
  per request rather than re-reading the environment.
- proxy.ts: the config gate runs before the session gate. verifySessionToken
  reads AUTH_SECRET and throws ConfigError without one, which Next turns into a
  500 per request — a status that says "this server has a bug" for a server that
  is merely unconfigured.
- Neither the 503 body nor /api/health names the missing variable. Those
  messages are operator information (configError.ts states the rule, the login
  route already follows it); the names go to the container log.
- HEALTHCHECK probes with node, already the entrypoint, so it adds no package
  and cannot break because a base image dropped a busybox applet.
- nginx.conf: marked dead. Nothing has installed nginx since 28258b5 and
  .dockerignore keeps it out of the build context, but it is the first place
  anyone looks at a 502 and the wrong one.

Verified: tsc --noEmit clean, eslint clean, and the Docker builder stage's
`NODE_ENV=production CI_BUILD=1 next build` exits 0. Misconfigured -> 503 on
/login, /dashboard, /api/sites with healthcheck exit 1; configured -> 200/307
with healthcheck exit 0.

This does not by itself end the outage: AUTH_SECRET still has to be set on the
container (Dokploy -> Environment). It makes the next occurrence legible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-17 22:54:41 +05:30
parent 958d352abe
commit aed8598eb9
7 changed files with 283 additions and 69 deletions

View File

@@ -0,0 +1,51 @@
import {configStatus} from '@/shared/config/configCheck';
/**
* GET /api/health — can this container serve?
*
* ── Why a container needs this ───────────────────────────────────────────
* The boot check used to answer the same question by killing the process, on
* the reasoning that a dead container is the only signal a platform cannot
* misread. It is also the only signal a BROWSER cannot read: the reverse proxy
* in front of it had nothing to connect to and returned 502 for every url,
* which is what a missing AUTH_SECRET looked like from the outside.
*
* This is the half of that trade worth keeping. The container stays up and
* explains itself, while the Dockerfile's HEALTHCHECK polls this route and
* drives the container `unhealthy` when it answers 503 — so a misconfigured
* deploy still cannot present itself as a working one.
*
* ── What it deliberately does not say ────────────────────────────────────
* Anonymous and public, so it publishes a state and a count, never the problem
* messages: those name environment variables, which is operator information
* (configError.ts states the rule; the login route follows it too). The names
* are printed once at boot, in the container log, where only an operator sees
* them.
*
* Exempt from the proxy's session gate — see HEALTH_PATH in src/proxy.ts — or
* an unauthenticated healthcheck would read 401 as "unhealthy" on a perfectly
* good container.
*/
export const dynamic = 'force-dynamic';
export function GET(): Response {
const {problems} = configStatus();
const healthy = problems.length === 0;
return Response.json(
{
status: healthy ? 'ok' : 'misconfigured',
// A count, not the messages. Enough to tell "one variable missing" from
// "this container has nothing set at all" without publishing which.
problems: problems.length,
},
{
status: healthy ? 200 : 503,
headers: {
'cache-control': 'no-store',
'x-loyaly-config': healthy ? 'ok' : 'misconfigured',
},
},
);
}