Stop an unreachable database from presenting as Bad Gateway
Two faults compounded into one symptom: with the database unreachable, every route on the service returned 502 - including /docs, which never touches it. _connect() passed no connect_timeout. A host that DROPS packets rather than refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked until the OS gave up - roughly 130 seconds on Linux. Every caller inherited that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS, defaulting to 5. Measured against an unroutable host: /api/health went from hanging past 25s to answering 200 in 5.07s. The container healthcheck then probed /api/health, so that hang timed out the check, the container was marked unhealthy, and the platform stopped routing to it. That is the part that turned a degraded dependency into a total outage, and it was introduced with the healthcheck itself. A healthcheck is a LIVENESS question, because the platform's answer to "no" is to take the container out of service. It may only ask whether the process is still serving HTTP. /api/health is a READINESS report - it dials Postgres and Ollama to say whether they are reachable, and coupling the container's existence to its dependencies is what made a running API unreachable. It now probes "/", which is served from memory and does no I/O, so it can fail only if the app really is gone. Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200, and the healthcheck exits 0. With the app stopped it still exits 1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
16
Dockerfile
16
Dockerfile
@@ -81,14 +81,16 @@ VOLUME ["/app/data", "/app/app/intelligence/artifacts"]
|
||||
ENV PORTS=3000,8000
|
||||
EXPOSE 3000 8000
|
||||
|
||||
# Liveness only, and passes if EITHER port answers. /api/health always returns
|
||||
# 200 - it reports Postgres and Ollama in the body as "degraded" rather than
|
||||
# failing - which is deliberate: a check that went red whenever Postgres blinked
|
||||
# would have Dokploy restart a perfectly healthy API in a loop.
|
||||
# Liveness only, and passes if EITHER port answers.
|
||||
#
|
||||
# The 10s timeout is not padding: the handler probes Ollama over HTTP with a 3s
|
||||
# timeout of its own, so an unreachable Ollama makes every check take ~3s.
|
||||
# start-period covers first boot, where the venv is still cold.
|
||||
# It probes "/", which is served from memory, NOT /api/health, which dials
|
||||
# Postgres and Ollama. That is the whole point: the platform's response to a
|
||||
# failed healthcheck is to stop routing traffic, so this may only ask "is the
|
||||
# process still serving HTTP". Tying it to the database meant an unreachable
|
||||
# Postgres blocked the handler for the OS TCP timeout, the check timed out, the
|
||||
# container was marked unhealthy, and a perfectly healthy API returned Bad
|
||||
# Gateway on every route. Use /api/health to ask whether dependencies are up;
|
||||
# it reports them in the body and always answers 200.
|
||||
HEALTHCHECK --interval=30s --timeout=10s --start-period=40s --retries=3 \
|
||||
CMD ["python", "serve.py", "--healthcheck"]
|
||||
|
||||
|
||||
Reference in New Issue
Block a user