feat(config): refuse to start a misconfigured container

Three misconfigurations in a row were each diagnosed the same slow way: deploy,
try to sign in, read a status code, guess, read a response body, guess again.
That is a property of the design, not of bad luck. Every required variable is
read lazily, on the first request that needs it, so a container missing its
environment starts clean, serves the sign-in page, and passes a healthcheck
while being completely unable to authenticate anybody.

The lazy reads have to stay — reading at module scope fails `next build`, which
collects page data with NODE_ENV=production and none of these variables set
(that is the "Failed to collect page data for /api/assistant" failure already
documented in apiClient). So the check goes in an instrumentation hook instead,
which Next runs once per server start and NEVER during a build: it returns early
when NEXT_PHASE is 'phase-production-build', in
server/lib/router-utils/instrumentation-globals.external.js.

Boot now either prints what it resolved:

  [loyaly] config ok — platform https://mcp.loyaly.ai, auth secret set, NODE_ENV=production

or refuses to start, naming every problem at once rather than one per deploy:

  refusing to start — 2 configuration problem(s):
    1. LOYALY_API_BASE is invalid: https://api.example.com is not a supported production API host...
    2. AUTH_SECRET is required in production — it signs the session cookie...

Verified against a real standalone build, booted four ways: unconfigured, fully
configured, LOYALY_API_BASE pointed at the console, and both wrong at once.

The platform origin and its validation move to shared/config/platformApi. This
is load-bearing, not tidying: apiClient is `server-only`, and that package
resolves to a module which THROWS ON IMPORT outside a react-server condition —
which the instrumentation bundle is not. Importing apiClient from the boot check
would have crashed every start, correctly configured or not. The extracted
module imports nothing but ConfigError, so the boot check and the request path
run the same function against the same allowlist.

Also corrects a claim I made in the Dockerfile two commits ago. `next build`
DOES copy .env into .next/standalone — writeStandaloneDirectory takes exactly
.env and .env.production and nothing else — so the explicit COPY is redundant
rather than required. It stays, with an accurate reason: it only happens when
.env is in the build context, and .dockerignore excluded it until recently.
Keeping the COPY makes that dependency fail the Docker build loudly instead of
producing an unconfigured image.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-17 19:53:15 +05:30
parent c01c750436
commit 1ee4b5d24c
4 changed files with 285 additions and 150 deletions

100
src/instrumentation.ts Normal file
View File

@@ -0,0 +1,100 @@
/**
* Boot-time configuration check.
*
* ── The problem this exists to end ───────────────────────────────────────
* Every required variable in this app is read LAZILY, on the first request
* that needs it. That is deliberate and must stay that way: `next build`
* collects page data with NODE_ENV=production and none of these variables
* present, so reading them at module scope fails the BUILD instead of the
* deployment (it did, once — "Failed to collect page data for /api/assistant",
* see apiClient.ts).
*
* The cost of lazy reads is that a misconfigured container LOOKS healthy. It
* starts, it serves the sign-in page, and the fault only appears when somebody
* tries to use it — as a 500 on a form, with the real reason buried in a
* response body. Three separate misconfigurations were diagnosed that way, one
* round trip at a time.
*
* This closes the gap without touching the lazy reads. `register()` runs once
* when a server instance starts and NEVER during a build — Next itself returns
* early when NEXT_PHASE is 'phase-production-build' (see
* server/lib/router-utils/instrumentation-globals.external.js). So the checks
* below run in exactly the situation they are about: a real server, booting,
* with a real environment.
*
* ── Why it throws ────────────────────────────────────────────────────────
* A container missing either variable cannot serve a single authenticated
* request. Refusing to start turns that into a failed deploy with a named
* cause in the log pane, which Dokploy surfaces immediately, instead of a
* green healthcheck in front of a console nobody can sign into. It also stops
* a broken image from replacing a working one.
*/
export async function register() {
/**
* Node only. `register()` is invoked once per runtime, and src/proxy.ts makes
* this app compile an Edge one too — without this guard the same check would
* run and log twice per boot. The node server is the process that serves
* every route handler, so validating there is what matters.
*/
if (process.env.NEXT_RUNTIME !== 'nodejs') return;
const isProduction = process.env.NODE_ENV === 'production';
/**
* shared/config/platformApi, NOT apiClient. apiClient is `server-only`, and
* that package resolves to a module which throws on import outside a
* react-server condition — which this bundle is not. Importing it here would
* crash every boot, correctly configured or not.
*/
const {resolvePlatformOrigin} = await import('@/shared/config/platformApi');
const {ConfigError} = await import('@/shared/errors/configError');
const problems: string[] = [];
/**
* Resolve the platform origin exactly the way a request would — same
* function, same validation, same allowlist. A check that reimplemented the
* rules would be a second source of truth and would drift.
*/
let platform: string | null = null;
try {
platform = resolvePlatformOrigin();
} catch (err) {
if (!(err instanceof ConfigError)) throw err;
problems.push(err.message);
}
/**
* AUTH_SECRET is checked by presence rather than by calling the signing
* functions, because those are pure and would have to be handed a throwaway
* payload to probe. Presence is the whole rule — sessionToken and tokenStore
* both refuse the development key in production and accept anything else.
*/
if (isProduction && !process.env.AUTH_SECRET) {
problems.push(
'AUTH_SECRET is required in production — it signs the session cookie and ' +
'encrypts the platform token bundle. Generate one with: openssl rand -base64 48',
);
}
if (problems.length > 0) {
// Numbered, because a container missing its environment is usually missing
// more than one variable, and fixing them one deploy at a time is the slow
// way to find that out.
const detail = problems.map((p, i) => ` ${i + 1}. ${p}`).join('\n');
throw new ConfigError(
`refusing to start — ${problems.length} configuration problem(s):\n${detail}\n` +
'\nSet these as environment variables on the container (Dokploy → Environment). ' +
'LOYALY_API_BASE also ships in the committed .env; AUTH_SECRET never does.',
);
}
// The healthy path says what it resolved, so a log reader can confirm the
// host WITHOUT having to trigger a request. Never prints AUTH_SECRET, only
// whether one was supplied.
console.log(
`[loyaly] config ok — platform ${platform}, ` +
`auth secret ${process.env.AUTH_SECRET ? 'set' : 'using development key'}, ` +
`NODE_ENV=${process.env.NODE_ENV}`,
);
}