Stop an unreachable database from presenting as Bad Gateway
Two faults compounded into one symptom: with the database unreachable, every route on the service returned 502 - including /docs, which never touches it. _connect() passed no connect_timeout. A host that DROPS packets rather than refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked until the OS gave up - roughly 130 seconds on Linux. Every caller inherited that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS, defaulting to 5. Measured against an unroutable host: /api/health went from hanging past 25s to answering 200 in 5.07s. The container healthcheck then probed /api/health, so that hang timed out the check, the container was marked unhealthy, and the platform stopped routing to it. That is the part that turned a degraded dependency into a total outage, and it was introduced with the healthcheck itself. A healthcheck is a LIVENESS question, because the platform's answer to "no" is to take the container out of service. It may only ask whether the process is still serving HTTP. /api/health is a READINESS report - it dials Postgres and Ollama to say whether they are reachable, and coupling the container's existence to its dependencies is what made a running API unreachable. It now probes "/", which is served from memory and does no I/O, so it can fail only if the app really is gone. Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200, and the healthcheck exits 0. With the app stopped it still exits 1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
24
serve.py
24
serve.py
@@ -129,15 +129,31 @@ def _healthcheck() -> int:
|
||||
Shares configured_ports() with the server so the check cannot drift from
|
||||
what is actually bound - the reason this lives here rather than being a
|
||||
python -c one-liner in the Dockerfile.
|
||||
|
||||
Probes "/" rather than /api/health, and the distinction matters. This is a
|
||||
LIVENESS check: the only question it may ask is "is this process still
|
||||
serving HTTP", because the platform's answer to "no" is to stop routing
|
||||
traffic to the container.
|
||||
|
||||
/api/health is a READINESS report - it dials Postgres and Ollama to say
|
||||
whether they are reachable. Using it here couples the container's existence
|
||||
to its dependencies: an unreachable database made /api/health block for the
|
||||
OS TCP timeout, the check timed out, the container was marked unhealthy, and
|
||||
a service that was running perfectly well returned Bad Gateway on every
|
||||
route - including the ones that never touch the database. "/" is served from
|
||||
memory and does no I/O at all, so it can only fail if the app really is gone.
|
||||
"""
|
||||
ports = configured_ports()
|
||||
for port in ports:
|
||||
try:
|
||||
with urllib.request.urlopen(
|
||||
f"http://127.0.0.1:{port}/api/health", timeout=8
|
||||
) as resp:
|
||||
if 200 <= resp.status < 400:
|
||||
with urllib.request.urlopen(f"http://127.0.0.1:{port}/", timeout=8) as resp:
|
||||
if 200 <= resp.status < 500:
|
||||
return 0
|
||||
except urllib.error.HTTPError as exc:
|
||||
# An HTTP status - even 404 - means something is listening and
|
||||
# routing. That is exactly what liveness asks.
|
||||
if exc.code < 500:
|
||||
return 0
|
||||
except (urllib.error.URLError, OSError, ValueError):
|
||||
continue
|
||||
print(
|
||||
|
||||
Reference in New Issue
Block a user