The HEALTHCHECK added an hour ago recreated the exact symptom. Dokploy runs applications as Docker Swarm services, and Swarm does not merely report an unhealthy task — it removes it from the service load balancer and reschedules it. /api/health answers 503 while a required variable is missing, so: AUTH_SECRET unset -> /api/health 503 -> task unhealthy -> pulled out of the load balancer -> Traefik has no backend -> 502 Bad Gateway on every url. The container was up and serving a 503 that names the fault, and nothing could reach it to read that 503. "A broken deploy must not look healthy" is a real concern, but enforcing it in the orchestrator destroys the diagnostics, and an outage whose reason cannot be seen is the more expensive failure. The container now stays in rotation whenever it can serve HTTP at all. Also, two things that make a secret that was SET look like one that was not: - Recommend `openssl rand -hex 32` everywhere instead of `openssl rand -base64 48`. A base64 value ends in '=' and may contain '+' and '/'; pasted into a dashboard field or a KEY=VALUE editor that splits on the first '=', it can be stored truncated or empty, which is indistinguishable from never setting it. Hex is [0-9a-f] only, so there is nothing for a parser to mangle. - Treat a whitespace-only AUTH_SECRET as missing, and print the secret's LENGTH (never its value) in the boot log. `openssl rand -hex 32` is 64 characters, so a much shorter number there is a value that arrived truncated — which otherwise presents as sessions that do not verify, with nothing to explain it. Verified on the rebuilt standalone payload: absent -> container alive, 503 `x-loyaly-config: misconfigured` on /, /login and /api/sites; whitespace -> the same; a real hex secret -> / 307, /login 200, /favicon.ico 200, /api/health 200 and `auth secret set (64 chars)` in the log. tsc --noEmit and eslint clean, production build exits 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
134 lines
6.2 KiB
Docker
134 lines
6.2 KiB
Docker
# syntax=docker/dockerfile:1
|
|
|
|
# Stage 1: Install dependencies
|
|
FROM node:22-alpine AS deps
|
|
RUN apk add --no-cache libc6-compat
|
|
WORKDIR /app
|
|
|
|
COPY package.json package-lock.json ./
|
|
# devDeps are required to build (typescript, tailwind, eslint-config-next).
|
|
# This whole stage is discarded — none of it reaches the runner.
|
|
RUN npm ci --no-audit --no-fund
|
|
|
|
# Stage 2: Build the Next.js application
|
|
FROM node:22-alpine AS builder
|
|
WORKDIR /app
|
|
COPY --from=deps /app/node_modules ./node_modules
|
|
COPY . .
|
|
|
|
ENV NEXT_TELEMETRY_DISABLED=1
|
|
ENV NODE_ENV=production
|
|
# Each Docker build starts from a clean layer, so Turbopack's .next/cache is
|
|
# written but never restored. Skipping it cuts ~20% of build CPU (the metric
|
|
# that matters on a 1-vCPU host) and 116 MB off this layer.
|
|
ENV CI_BUILD=1
|
|
|
|
RUN npm run build
|
|
|
|
# Stage 3: Production runner with Next.js Standalone
|
|
FROM node:22-alpine AS runner
|
|
RUN apk add --no-cache libc6-compat
|
|
WORKDIR /app
|
|
|
|
ENV NODE_ENV=production
|
|
ENV NEXT_TELEMETRY_DISABLED=1
|
|
ENV PORT=3000
|
|
ENV HOSTNAME="0.0.0.0"
|
|
|
|
# ── Runtime configuration ────────────────────────────────────────────────
|
|
#
|
|
# Two variables are required to SERVE a request. Neither is required to BUILD:
|
|
# LOYALY_API_BASE is resolved on first use rather than at module load (see
|
|
# apiClient.ts) and sessionToken/tokenStore derive their key per call, so page
|
|
# data collection reads neither.
|
|
#
|
|
# LOYALY_API_BASE the Behavision API origin — https://mcp.loyaly.ai
|
|
# (NOT platform.loyaly.ai, which serves this console)
|
|
# → SHIPPED, in the .env copied below. Not a secret.
|
|
#
|
|
# AUTH_SECRET signs the session cookie and encrypts the platform token
|
|
# bundle. Generate with: openssl rand -hex 32
|
|
# → NOT shipped. Set it as a Dokploy secret.
|
|
#
|
|
# Hex rather than base64 on purpose. `openssl rand -base64 48` ends in '=' and
|
|
# may contain '+' and '/'. Pasted into a dashboard field or a KEY=VALUE env
|
|
# editor that splits on the first '=', that value can be stored truncated — or
|
|
# not at all — and the result is indistinguishable from never having set it.
|
|
# Hex is [0-9a-f] only, so there is nothing for a parser to mangle. 32 bytes is
|
|
# 256 bits, more than the HMAC and the AES-256 key derived from it need.
|
|
#
|
|
# A container started WITHOUT AUTH_SECRET no longer dies. It boots, prints the
|
|
# missing variable on stderr, answers 503 with `x-loyaly-config: misconfigured`
|
|
# on every request, and fails the HEALTHCHECK below. Exiting instead is what
|
|
# made this fault present as a bare 502 Bad Gateway from the reverse proxy on
|
|
# every url — /favicon.ico first, in the browser console — with the one line
|
|
# that explained it trapped inside a restart-looping container. See
|
|
# src/instrumentation-node.ts.
|
|
#
|
|
# AUTH_SECRET used to be an ENV line here with a literal value, which put a
|
|
# session-forging key in git: anyone who could read the repo could mint a
|
|
# cookie for any user, and every built image carried it in a layer that
|
|
# `docker history` prints. Docker's own linter flags the pattern
|
|
# (SecretsUsedInArgOrEnv). It is gone; rotate the old value. That is why the
|
|
# split above exists — "inject everything" also meant injecting the one value
|
|
# that is public knowledge, and forgetting it took the console down.
|
|
|
|
# Run as a non-root user; nextjs owns nothing it does not need to write.
|
|
RUN addgroup -g 1001 -S nodejs && adduser -u 1001 -S nextjs -G nodejs
|
|
|
|
# Copy public static assets and standalone build output.
|
|
# These three paths are the ENTIRE runtime payload (~57 MB). Never copy the
|
|
# whole .next/ directory here — .next/dev and .next/cache are build-host-only
|
|
# and account for ~1.96 GB.
|
|
COPY --from=builder --chown=nextjs:nodejs /app/public ./public
|
|
COPY --from=builder --chown=nextjs:nodejs /app/.next/standalone ./
|
|
COPY --from=builder --chown=nextjs:nodejs /app/.next/static ./.next/static
|
|
|
|
# The production environment, as a file the server reads at boot.
|
|
#
|
|
# Deliberately redundant, and worth keeping. `next build` already copies .env
|
|
# (and .env.production, and nothing else — see writeStandaloneDirectory in
|
|
# next/dist/build/index.js) into .next/standalone, so the line above lands one
|
|
# at /app/.env on its own. But it only does that when .env was in the BUILD
|
|
# CONTEXT, and .dockerignore excluded it until recently — which is precisely
|
|
# how images shipped with no LOYALY_API_BASE at all.
|
|
#
|
|
# This line turns that silent outcome into a loud one: exclude .env again and
|
|
# the Docker build FAILS here with "file not found" instead of producing an
|
|
# unconfigured image that starts and then rejects every sign-in.
|
|
#
|
|
# It does not pin the deployment either way: @next/env never overwrites a
|
|
# variable already present in process.env, so anything set in Dokploy wins.
|
|
COPY --chown=nextjs:nodejs .env ./.env
|
|
|
|
USER nextjs
|
|
|
|
EXPOSE 3000
|
|
|
|
# NO HEALTHCHECK, on purpose.
|
|
#
|
|
# One was added here and removed within the hour, because it recreated the
|
|
# exact 502 it was meant to replace. Dokploy runs applications as Docker Swarm
|
|
# services, and Swarm does not merely REPORT an unhealthy task — it pulls it
|
|
# out of the service load balancer and reschedules it. So a healthcheck wired
|
|
# to /api/health, which answers 503 while a required variable is missing, meant:
|
|
#
|
|
# AUTH_SECRET unset -> /api/health 503 -> task unhealthy -> removed from the
|
|
# load balancer and restarted -> Traefik has no backend -> 502 Bad Gateway on
|
|
# every url, which is precisely the symptom this whole change exists to end.
|
|
#
|
|
# The container would have been up, serving a 503 that names the fault, and
|
|
# nobody could have reached it. "A broken deploy must not look healthy" is a
|
|
# real concern, but enforcing it in the orchestrator destroys the diagnostics —
|
|
# and an outage you cannot see the reason for is the more expensive failure.
|
|
#
|
|
# So: the container stays in rotation whenever it can serve HTTP at all, and
|
|
# the configuration state is reported where it can actually be read — 503 with
|
|
# `x-loyaly-config: misconfigured` on every gated request, /api/health for a
|
|
# direct answer, and the named variable in the boot log.
|
|
#
|
|
# If a healthcheck is ever added back, it must probe LIVENESS (is the server
|
|
# answering?) and never configuration, or this comment is being relearned.
|
|
|
|
CMD ["node", "server.js"]
|