feat(config): refuse to start a misconfigured container

Three misconfigurations in a row were each diagnosed the same slow way: deploy,
try to sign in, read a status code, guess, read a response body, guess again.
That is a property of the design, not of bad luck. Every required variable is
read lazily, on the first request that needs it, so a container missing its
environment starts clean, serves the sign-in page, and passes a healthcheck
while being completely unable to authenticate anybody.

The lazy reads have to stay — reading at module scope fails `next build`, which
collects page data with NODE_ENV=production and none of these variables set
(that is the "Failed to collect page data for /api/assistant" failure already
documented in apiClient). So the check goes in an instrumentation hook instead,
which Next runs once per server start and NEVER during a build: it returns early
when NEXT_PHASE is 'phase-production-build', in
server/lib/router-utils/instrumentation-globals.external.js.

Boot now either prints what it resolved:

  [loyaly] config ok — platform https://mcp.loyaly.ai, auth secret set, NODE_ENV=production

or refuses to start, naming every problem at once rather than one per deploy:

  refusing to start — 2 configuration problem(s):
    1. LOYALY_API_BASE is invalid: https://api.example.com is not a supported production API host...
    2. AUTH_SECRET is required in production — it signs the session cookie...

Verified against a real standalone build, booted four ways: unconfigured, fully
configured, LOYALY_API_BASE pointed at the console, and both wrong at once.

The platform origin and its validation move to shared/config/platformApi. This
is load-bearing, not tidying: apiClient is `server-only`, and that package
resolves to a module which THROWS ON IMPORT outside a react-server condition —
which the instrumentation bundle is not. Importing apiClient from the boot check
would have crashed every start, correctly configured or not. The extracted
module imports nothing but ConfigError, so the boot check and the request path
run the same function against the same allowlist.

Also corrects a claim I made in the Dockerfile two commits ago. `next build`
DOES copy .env into .next/standalone — writeStandaloneDirectory takes exactly
.env and .env.production and nothing else — so the explicit COPY is redundant
rather than required. It stays, with an accurate reason: it only happens when
.env is in the build context, and .dockerignore excluded it until recently.
Keeping the COPY makes that dependency fail the Docker build loudly instead of
producing an unconfigured image.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-17 19:53:15 +05:30
parent c01c750436
commit 1ee4b5d24c
4 changed files with 285 additions and 150 deletions

View File

@@ -0,0 +1,167 @@
/**
* Where the Loyaly platform API lives, and the rules about what may be called
* one. Pure configuration resolution: no I/O, no crypto, no `server-only`.
*
* ── Why it is not in apiClient ───────────────────────────────────────────
* apiClient is `server-only`, and that package resolves to a module which
* THROWS ON IMPORT outside a react-server condition. src/instrumentation.ts
* runs this same validation at boot and is NOT compiled in that condition, so
* importing apiClient from it would crash the server on start — for every
* deployment, correctly configured or not.
*
* Splitting it also keeps one source of truth: the boot check and the request
* path call the SAME function against the SAME allowlist, so a check that
* passes at startup cannot be contradicted by the first request.
*/
import {ConfigError} from '@/shared/errors/configError';
/**
* Server-side only — deliberately NOT NEXT_PUBLIC. Publishing the platform
* host would let a browser bypass the BFF, which is the whole point of it.
*
* ── Why there is no remote fallback ──────────────────────────────────────
* This used to default to `https://platform.loyaly.ai`, which is NOT the
* Behavision API — that host serves this very console. Measured: it answers
* `GET /api/auth/me` with the console's own 404 HTML page, and a login POST
* with the console's own `{error:{code,message}}` envelope rather than the
* platform's flat `{error,message}`. So an unset variable did not fail; it
* quietly pointed the BFF at its own origin, and every upstream call became a
* request the console made to itself.
*
* A wrong host that *works* is worse than a startup failure, so production
* refuses to SERVE without the variable — the same stance `tokenStore.ts`
* takes on AUTH_SECRET, and for the same reason. Development falls back to the
* local backend, which is the only host a dev machine can usefully mean.
*
* local http://127.0.0.1:8088
* production https://mcp.loyaly.ai
*
* ── Why this is resolved lazily and not at module scope ──────────────────
* It used to be `const BASE = resolveBase()`, evaluated the moment any module
* imported this one. That broke `next build`: the "Collecting page data" step
* imports every route module, the Docker builder stage sets NODE_ENV=production,
* and LOYALY_API_BASE is a RUNTIME value that is not present while building an
* image. So the guard fired against the build instead of against a
* misconfigured server, and the deploy failed with "Failed to collect page data
* for /api/assistant".
*
* Deferring to first use draws the line where it belongs: building an image
* needs no platform host, serving a request does. `tokenStore.key()` is a
* function for exactly this reason — this now matches it rather than only
* claiming to. The result is memoised, so the environment is read once per
* process and a healthy server pays nothing per request.
*/
const DEV_API_BASE = 'http://127.0.0.1:8088';
/**
* The only origin that serves the Loyaly platform API in production.
*
* Production is an allowlist of exactly one entry rather than a shape check,
* because "looks like a URL" is what let the wrong host through before. A new
* environment — staging, a regional deployment — is a deliberate line added
* here, not something a typo in a dashboard can invent.
*/
const PRODUCTION_API_ORIGIN = 'https://mcp.loyaly.ai';
/**
* Hosts that are definitely NOT the API, and why.
*
* Rejected in EVERY environment, development included: this is not a
* production-hardening rule, it is a statement of fact about what the host
* serves. Naming the reason matters — "rejected" alone sends somebody looking
* for a firewall or a DNS problem, when the actual fix is one word in a
* variable.
*/
const KNOWN_WRONG_HOSTS: Record<string, string> = {
'platform.loyaly.ai':
'serves this console, not the Loyaly API — pointing the BFF there makes ' +
'it call its own origin',
};
/**
* `buildUrl` — which resolves this variable — is called OUTSIDE the try block
* that turns a failed fetch into `UpstreamError(0, 'network')`. It used to be
* inside it, which made "nobody set LOYALY_API_BASE" indistinguishable from
* "the platform is down" at every call site, and had the login route report
* both as 502 platform_unreachable.
*/
function configError(detail: string): ConfigError {
return new ConfigError(
`LOYALY_API_BASE is invalid: ${detail}. ` +
`Set it to ${PRODUCTION_API_ORIGIN} in production, or ${DEV_API_BASE} locally.`,
);
}
/**
* Validate a configured value and reduce it to an origin.
*
* The path is dropped on purpose rather than preserved: `new URL(path, base)`
* has always discarded a base path, so a value like `https://host/v1` never
* did what whoever wrote it expected. Returning the origin makes that visible
* instead of silently ignored.
*/
function validateBase(raw: string, isProduction: boolean): string {
let url: URL;
try {
url = new URL(raw);
} catch {
throw configError(`"${raw}" is not an absolute URL`);
}
if (url.protocol !== 'https:' && url.protocol !== 'http:') {
throw configError(`"${url.protocol}" is not an http(s) URL`);
}
const wrong = KNOWN_WRONG_HOSTS[url.hostname];
if (wrong) throw configError(`${url.hostname} ${wrong}`);
if (isProduction) {
if (url.origin !== PRODUCTION_API_ORIGIN) {
throw configError(
`${url.origin} is not a supported production API host`,
);
}
return url.origin;
}
/**
* Development stays permissive by design. A dev legitimately points this at
* a LAN address, a tunnel or a container host, and breaking that to enforce
* a production rule would cost more than it protects — nothing a dev machine
* reaches is production. The known-wrong list above still applies.
*/
return url.origin;
}
let cachedBase: string | null = null;
function resolveBase(): string {
if (cachedBase !== null) return cachedBase;
const isProduction = process.env.NODE_ENV === 'production';
const configured = process.env.LOYALY_API_BASE?.trim();
if (configured) {
// NOT cached before validating: an invalid value must throw on every
// request, the same way a missing one does.
return (cachedBase = validateBase(configured, isProduction));
}
if (isProduction) {
throw new ConfigError(
'LOYALY_API_BASE is required in production — refusing to guess the ' +
'Loyaly platform host. Set it to https://mcp.loyaly.ai.',
);
}
return (cachedBase = DEV_API_BASE);
}
/**
* The upstream origin, resolved on first use and memoised. Call it; do not
* hoist it — see the note on lazy resolution above.
*/
export {resolveBase as resolvePlatformOrigin};
/** The one production origin, for messages that need to name it. */
export {PRODUCTION_API_ORIGIN};