Files
loyaly-merchant/Dockerfile
Aravind aed8598eb9 fix(deploy): serve a 503 that names the fault instead of dying into a 502
A container started without AUTH_SECRET called process.exit(1) from the boot
check. The container died, Dokploy's Traefik had no upstream to proxy to, and
every url answered 502 Bad Gateway — /favicon.ico first, which is the line that
shows up in a browser console. The one message that explained it was on stderr
inside a restart-looping container, so the fastest fault in this app to fix
became the slowest to identify.

Reproduced against the real standalone payload: without AUTH_SECRET the process
exited 1; with it, /favicon.ico 200, /login 200, / 307.

The container now boots and stays up. While a required variable is missing the
proxy answers 503 with `x-loyaly-config: misconfigured` on every gated request,
and the boot log names the variable. A Docker HEALTHCHECK against the new
/api/health keeps the property the exit was protecting — a broken deploy still
reports unhealthy rather than presenting itself as a working one.

- configCheck.ts: one runtime-neutral check shared by the boot log, the proxy
  and the health route, memoised so a healthy server pays an array-length read
  per request rather than re-reading the environment.
- proxy.ts: the config gate runs before the session gate. verifySessionToken
  reads AUTH_SECRET and throws ConfigError without one, which Next turns into a
  500 per request — a status that says "this server has a bug" for a server that
  is merely unconfigured.
- Neither the 503 body nor /api/health names the missing variable. Those
  messages are operator information (configError.ts states the rule, the login
  route already follows it); the names go to the container log.
- HEALTHCHECK probes with node, already the entrypoint, so it adds no package
  and cannot break because a base image dropped a busybox applet.
- nginx.conf: marked dead. Nothing has installed nginx since 28258b5 and
  .dockerignore keeps it out of the build context, but it is the first place
  anyone looks at a 502 and the wrong one.

Verified: tsc --noEmit clean, eslint clean, and the Docker builder stage's
`NODE_ENV=production CI_BUILD=1 next build` exits 0. Misconfigured -> 503 on
/login, /dashboard, /api/sites with healthcheck exit 1; configured -> 200/307
with healthcheck exit 0.

This does not by itself end the outage: AUTH_SECRET still has to be set on the
container (Dokploy -> Environment). It makes the next occurrence legible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 22:54:41 +05:30

121 lines
5.5 KiB
Docker

# syntax=docker/dockerfile:1
# Stage 1: Install dependencies
FROM node:22-alpine AS deps
RUN apk add --no-cache libc6-compat
WORKDIR /app
COPY package.json package-lock.json ./
# devDeps are required to build (typescript, tailwind, eslint-config-next).
# This whole stage is discarded — none of it reaches the runner.
RUN npm ci --no-audit --no-fund
# Stage 2: Build the Next.js application
FROM node:22-alpine AS builder
WORKDIR /app
COPY --from=deps /app/node_modules ./node_modules
COPY . .
ENV NEXT_TELEMETRY_DISABLED=1
ENV NODE_ENV=production
# Each Docker build starts from a clean layer, so Turbopack's .next/cache is
# written but never restored. Skipping it cuts ~20% of build CPU (the metric
# that matters on a 1-vCPU host) and 116 MB off this layer.
ENV CI_BUILD=1
RUN npm run build
# Stage 3: Production runner with Next.js Standalone
FROM node:22-alpine AS runner
RUN apk add --no-cache libc6-compat
WORKDIR /app
ENV NODE_ENV=production
ENV NEXT_TELEMETRY_DISABLED=1
ENV PORT=3000
ENV HOSTNAME="0.0.0.0"
# ── Runtime configuration ────────────────────────────────────────────────
#
# Two variables are required to SERVE a request. Neither is required to BUILD:
# LOYALY_API_BASE is resolved on first use rather than at module load (see
# apiClient.ts) and sessionToken/tokenStore derive their key per call, so page
# data collection reads neither.
#
# LOYALY_API_BASE the Behavision API origin — https://mcp.loyaly.ai
# (NOT platform.loyaly.ai, which serves this console)
# → SHIPPED, in the .env copied below. Not a secret.
#
# AUTH_SECRET signs the session cookie and encrypts the platform token
# bundle. Generate with: openssl rand -base64 48
# → NOT shipped. Set it as a Dokploy secret.
#
# A container started WITHOUT AUTH_SECRET no longer dies. It boots, prints the
# missing variable on stderr, answers 503 with `x-loyaly-config: misconfigured`
# on every request, and fails the HEALTHCHECK below. Exiting instead is what
# made this fault present as a bare 502 Bad Gateway from the reverse proxy on
# every url — /favicon.ico first, in the browser console — with the one line
# that explained it trapped inside a restart-looping container. See
# src/instrumentation-node.ts.
#
# AUTH_SECRET used to be an ENV line here with a literal value, which put a
# session-forging key in git: anyone who could read the repo could mint a
# cookie for any user, and every built image carried it in a layer that
# `docker history` prints. Docker's own linter flags the pattern
# (SecretsUsedInArgOrEnv). It is gone; rotate the old value. That is why the
# split above exists — "inject everything" also meant injecting the one value
# that is public knowledge, and forgetting it took the console down.
# Run as a non-root user; nextjs owns nothing it does not need to write.
RUN addgroup -g 1001 -S nodejs && adduser -u 1001 -S nextjs -G nodejs
# Copy public static assets and standalone build output.
# These three paths are the ENTIRE runtime payload (~57 MB). Never copy the
# whole .next/ directory here — .next/dev and .next/cache are build-host-only
# and account for ~1.96 GB.
COPY --from=builder --chown=nextjs:nodejs /app/public ./public
COPY --from=builder --chown=nextjs:nodejs /app/.next/standalone ./
COPY --from=builder --chown=nextjs:nodejs /app/.next/static ./.next/static
# The production environment, as a file the server reads at boot.
#
# Deliberately redundant, and worth keeping. `next build` already copies .env
# (and .env.production, and nothing else — see writeStandaloneDirectory in
# next/dist/build/index.js) into .next/standalone, so the line above lands one
# at /app/.env on its own. But it only does that when .env was in the BUILD
# CONTEXT, and .dockerignore excluded it until recently — which is precisely
# how images shipped with no LOYALY_API_BASE at all.
#
# This line turns that silent outcome into a loud one: exclude .env again and
# the Docker build FAILS here with "file not found" instead of producing an
# unconfigured image that starts and then rejects every sign-in.
#
# It does not pin the deployment either way: @next/env never overwrites a
# variable already present in process.env, so anything set in Dokploy wins.
COPY --chown=nextjs:nodejs .env ./.env
USER nextjs
EXPOSE 3000
# Is this container able to serve, or merely running?
#
# /api/health answers 200 only when every required variable is present and
# acceptable, and 503 otherwise, so a container missing AUTH_SECRET reports
# `unhealthy` rather than presenting itself as a working deploy. That is the
# property the old process.exit(1) was protecting; this keeps it without
# taking the site down to do it.
#
# Probed with node rather than curl or busybox wget: node is already the
# container's entrypoint, so this adds no package and cannot break because a
# base image dropped an applet. `r.ok` is the check — a 503 from /api/health
# exits 1, which is what marks the container unhealthy — and the catch covers
# a server that is not listening at all.
#
# start-period covers first boot: Next reports ready in well under 20s here,
# and an unready server must not be reported as a broken one.
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s --retries=3 \
CMD node -e "fetch('http://127.0.0.1:'+(process.env.PORT||3000)+'/api/health').then(r=>process.exit(r.ok?0:1)).catch(()=>process.exit(1))"
CMD ["node", "server.js"]