fix(deploy): serve a 503 that names the fault instead of dying into a 502
A container started without AUTH_SECRET called process.exit(1) from the boot
check. The container died, Dokploy's Traefik had no upstream to proxy to, and
every url answered 502 Bad Gateway — /favicon.ico first, which is the line that
shows up in a browser console. The one message that explained it was on stderr
inside a restart-looping container, so the fastest fault in this app to fix
became the slowest to identify.
Reproduced against the real standalone payload: without AUTH_SECRET the process
exited 1; with it, /favicon.ico 200, /login 200, / 307.
The container now boots and stays up. While a required variable is missing the
proxy answers 503 with `x-loyaly-config: misconfigured` on every gated request,
and the boot log names the variable. A Docker HEALTHCHECK against the new
/api/health keeps the property the exit was protecting — a broken deploy still
reports unhealthy rather than presenting itself as a working one.
- configCheck.ts: one runtime-neutral check shared by the boot log, the proxy
and the health route, memoised so a healthy server pays an array-length read
per request rather than re-reading the environment.
- proxy.ts: the config gate runs before the session gate. verifySessionToken
reads AUTH_SECRET and throws ConfigError without one, which Next turns into a
500 per request — a status that says "this server has a bug" for a server that
is merely unconfigured.
- Neither the 503 body nor /api/health names the missing variable. Those
messages are operator information (configError.ts states the rule, the login
route already follows it); the names go to the container log.
- HEALTHCHECK probes with node, already the entrypoint, so it adds no package
and cannot break because a base image dropped a busybox applet.
- nginx.conf: marked dead. Nothing has installed nginx since 28258b5 and
.dockerignore keeps it out of the build context, but it is the first place
anyone looks at a 502 and the wrong one.
Verified: tsc --noEmit clean, eslint clean, and the Docker builder stage's
`NODE_ENV=production CI_BUILD=1 next build` exits 0. Misconfigured -> 503 on
/login, /dashboard, /api/sites with healthcheck exit 1; configured -> 200/307
with healthcheck exit 0.
This does not by itself end the outage: AUTH_SECRET still has to be set on the
container (Dokploy -> Environment). It makes the next occurrence legible.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
79
src/shared/config/configCheck.ts
Normal file
79
src/shared/config/configCheck.ts
Normal file
@@ -0,0 +1,79 @@
|
||||
/**
|
||||
* The one list of "what must be true before this server can serve a request",
|
||||
* shared by everything that needs to ask.
|
||||
*
|
||||
* ── Why it is its own module ─────────────────────────────────────────────
|
||||
* Its three callers live in three different runtimes:
|
||||
*
|
||||
* src/instrumentation-node.ts Node, once at boot
|
||||
* src/proxy.ts Edge, once per request
|
||||
* src/app/api/health/route.ts Node, on demand
|
||||
*
|
||||
* So this file may import nothing runtime-specific. `platformApi` and
|
||||
* `configError` are both deliberately free of `server-only` and of node:
|
||||
* builtins for the same reason — see the headers on those two.
|
||||
*
|
||||
* Reimplementing the rules in any of the three would create a second source of
|
||||
* truth, and the first thing it would do is drift: a boot check that passes
|
||||
* while the request path throws is worse than no boot check at all.
|
||||
*/
|
||||
|
||||
import {ConfigError} from '@/shared/errors/configError';
|
||||
import {resolvePlatformOrigin} from '@/shared/config/platformApi';
|
||||
|
||||
export interface ConfigStatus {
|
||||
/** Operator-facing messages. Empty means the server can serve. */
|
||||
problems: string[];
|
||||
/** The resolved upstream origin, or null when it could not be resolved. */
|
||||
platform: string | null;
|
||||
}
|
||||
|
||||
let cached: ConfigStatus | null = null;
|
||||
|
||||
/**
|
||||
* Both required variables are read from the real environment, which cannot
|
||||
* change while the process lives, so this is memoised. A misconfigured server
|
||||
* therefore pays the check once and answers from a boolean afterwards — which
|
||||
* is what makes it affordable in the proxy, on every request.
|
||||
*/
|
||||
export function configStatus(): ConfigStatus {
|
||||
if (cached) return cached;
|
||||
|
||||
const problems: string[] = [];
|
||||
let platform: string | null = null;
|
||||
|
||||
/**
|
||||
* Resolved through the SAME function a request uses, against the same
|
||||
* allowlist, so a check that passes here cannot be contradicted by the first
|
||||
* real call. A non-ConfigError is a bug in the resolver rather than a
|
||||
* deployment fault, so it propagates.
|
||||
*/
|
||||
try {
|
||||
platform = resolvePlatformOrigin();
|
||||
} catch (err) {
|
||||
if (!(err instanceof ConfigError)) throw err;
|
||||
problems.push(err.message);
|
||||
}
|
||||
|
||||
/**
|
||||
* AUTH_SECRET is checked by presence, because presence is the whole rule:
|
||||
* sessionToken and tokenStore refuse the development key in production and
|
||||
* accept anything else. Calling the signing functions to probe would mean
|
||||
* inventing a throwaway payload to sign.
|
||||
*/
|
||||
if (process.env.NODE_ENV === 'production' && !process.env.AUTH_SECRET) {
|
||||
problems.push(
|
||||
'AUTH_SECRET is not set — it signs the session cookie and encrypts the ' +
|
||||
'platform token bundle, and production refuses to fall back to the ' +
|
||||
'development key. Set it on the container (Dokploy → Environment). ' +
|
||||
'Generate one with: openssl rand -base64 48',
|
||||
);
|
||||
}
|
||||
|
||||
return (cached = {problems, platform});
|
||||
}
|
||||
|
||||
/** True when every required variable is present and acceptable. */
|
||||
export function isConfigured(): boolean {
|
||||
return configStatus().problems.length === 0;
|
||||
}
|
||||
Reference in New Issue
Block a user