fix(deploy): drop the healthcheck that reproduced the 502 it was meant to end

The HEALTHCHECK added an hour ago recreated the exact symptom. Dokploy runs
applications as Docker Swarm services, and Swarm does not merely report an
unhealthy task — it removes it from the service load balancer and reschedules
it. /api/health answers 503 while a required variable is missing, so:

  AUTH_SECRET unset -> /api/health 503 -> task unhealthy -> pulled out of the
  load balancer -> Traefik has no backend -> 502 Bad Gateway on every url.

The container was up and serving a 503 that names the fault, and nothing could
reach it to read that 503. "A broken deploy must not look healthy" is a real
concern, but enforcing it in the orchestrator destroys the diagnostics, and an
outage whose reason cannot be seen is the more expensive failure. The container
now stays in rotation whenever it can serve HTTP at all.

Also, two things that make a secret that was SET look like one that was not:

- Recommend `openssl rand -hex 32` everywhere instead of `openssl rand -base64
  48`. A base64 value ends in '=' and may contain '+' and '/'; pasted into a
  dashboard field or a KEY=VALUE editor that splits on the first '=', it can be
  stored truncated or empty, which is indistinguishable from never setting it.
  Hex is [0-9a-f] only, so there is nothing for a parser to mangle.
- Treat a whitespace-only AUTH_SECRET as missing, and print the secret's LENGTH
  (never its value) in the boot log. `openssl rand -hex 32` is 64 characters, so
  a much shorter number there is a value that arrived truncated — which
  otherwise presents as sessions that do not verify, with nothing to explain it.

Verified on the rebuilt standalone payload: absent -> container alive, 503
`x-loyaly-config: misconfigured` on /, /login and /api/sites; whitespace -> the
same; a real hex secret -> / 307, /login 200, /favicon.ico 200, /api/health 200
and `auth secret set (64 chars)` in the log. tsc --noEmit and eslint clean,
production build exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-17 23:08:03 +05:30
parent aed8598eb9
commit 0dc865ba66
5 changed files with 54 additions and 21 deletions

View File

@@ -47,9 +47,16 @@ ENV HOSTNAME="0.0.0.0"
# → SHIPPED, in the .env copied below. Not a secret.
#
# AUTH_SECRET signs the session cookie and encrypts the platform token
# bundle. Generate with: openssl rand -base64 48
# bundle. Generate with: openssl rand -hex 32
# → NOT shipped. Set it as a Dokploy secret.
#
# Hex rather than base64 on purpose. `openssl rand -base64 48` ends in '=' and
# may contain '+' and '/'. Pasted into a dashboard field or a KEY=VALUE env
# editor that splits on the first '=', that value can be stored truncated — or
# not at all — and the result is indistinguishable from never having set it.
# Hex is [0-9a-f] only, so there is nothing for a parser to mangle. 32 bytes is
# 256 bits, more than the HMAC and the AES-256 key derived from it need.
#
# A container started WITHOUT AUTH_SECRET no longer dies. It boots, prints the
# missing variable on stderr, answers 503 with `x-loyaly-config: misconfigured`
# on every request, and fails the HEALTHCHECK below. Exiting instead is what
@@ -98,23 +105,29 @@ USER nextjs
EXPOSE 3000
# Is this container able to serve, or merely running?
# NO HEALTHCHECK, on purpose.
#
# /api/health answers 200 only when every required variable is present and
# acceptable, and 503 otherwise, so a container missing AUTH_SECRET reports
# `unhealthy` rather than presenting itself as a working deploy. That is the
# property the old process.exit(1) was protecting; this keeps it without
# taking the site down to do it.
# One was added here and removed within the hour, because it recreated the
# exact 502 it was meant to replace. Dokploy runs applications as Docker Swarm
# services, and Swarm does not merely REPORT an unhealthy task — it pulls it
# out of the service load balancer and reschedules it. So a healthcheck wired
# to /api/health, which answers 503 while a required variable is missing, meant:
#
# Probed with node rather than curl or busybox wget: node is already the
# container's entrypoint, so this adds no package and cannot break because a
# base image dropped an applet. `r.ok` is the check — a 503 from /api/health
# exits 1, which is what marks the container unhealthy — and the catch covers
# a server that is not listening at all.
# AUTH_SECRET unset -> /api/health 503 -> task unhealthy -> removed from the
# load balancer and restarted -> Traefik has no backend -> 502 Bad Gateway on
# every url, which is precisely the symptom this whole change exists to end.
#
# start-period covers first boot: Next reports ready in well under 20s here,
# and an unready server must not be reported as a broken one.
HEALTHCHECK --interval=30s --timeout=5s --start-period=20s --retries=3 \
CMD node -e "fetch('http://127.0.0.1:'+(process.env.PORT||3000)+'/api/health').then(r=>process.exit(r.ok?0:1)).catch(()=>process.exit(1))"
# The container would have been up, serving a 503 that names the fault, and
# nobody could have reached it. "A broken deploy must not look healthy" is a
# real concern, but enforcing it in the orchestrator destroys the diagnostics —
# and an outage you cannot see the reason for is the more expensive failure.
#
# So: the container stays in rotation whenever it can serve HTTP at all, and
# the configuration state is reported where it can actually be read — 503 with
# `x-loyaly-config: misconfigured` on every gated request, /api/health for a
# direct answer, and the named variable in the boot log.
#
# If a healthcheck is ever added back, it must probe LIVENESS (is the server
# answering?) and never configuration, or this comment is being relearned.
CMD ["node", "server.js"]