Two faults compounded into one symptom: with the database unreachable, every
route on the service returned 502 - including /docs, which never touches it.
_connect() passed no connect_timeout. A host that DROPS packets rather than
refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked
until the OS gave up - roughly 130 seconds on Linux. Every caller inherited
that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS,
defaulting to 5. Measured against an unroutable host: /api/health went from
hanging past 25s to answering 200 in 5.07s.
The container healthcheck then probed /api/health, so that hang timed out the
check, the container was marked unhealthy, and the platform stopped routing to
it. That is the part that turned a degraded dependency into a total outage, and
it was introduced with the healthcheck itself.
A healthcheck is a LIVENESS question, because the platform's answer to "no" is
to take the container out of service. It may only ask whether the process is
still serving HTTP. /api/health is a READINESS report - it dials Postgres and
Ollama to say whether they are reachable, and coupling the container's
existence to its dependencies is what made a running API unreachable. It now
probes "/", which is served from memory and does no I/O, so it can fail only if
the app really is gone.
Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200,
and the healthcheck exits 0. With the app stopped it still exits 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dokploy routes the domain to port 3000, but the container only bound 8000, so
the proxy had nothing to talk to and the domain returned 502 with a perfectly
healthy process behind it.
The frontend image already solved this by answering on both 80 and 3000
(`listen 80; listen 3000;` in nginx.conf). Do the same here rather than swap one
guess for another: 3000 is what the platform routes to, and 8000 is what the
README, the vite dev proxy and docker-compose all target, so binding both means
the container works whichever one it is pointed at.
uvicorn's CLI takes a single --port, but Server.run() accepts pre-bound
sockets, so serve.py binds each port and hands the list to one uvicorn - no
extra worker or second process to supervise. PORT still pins a single port for
anyone who wants one; PORTS changes the pair.
A port that cannot be bound is logged and skipped rather than being fatal,
since losing one of the two should not take down a service the platform only
routes to on the other. It exits non-zero only when nothing is listening at
all, so a genuinely dead container is still reported as failed.
The healthcheck moves into the same file and passes if either port answers,
which keeps it from drifting out of sync with what is actually bound.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>