Stop an unreachable database from presenting as Bad Gateway
Two faults compounded into one symptom: with the database unreachable, every route on the service returned 502 - including /docs, which never touches it. _connect() passed no connect_timeout. A host that DROPS packets rather than refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked until the OS gave up - roughly 130 seconds on Linux. Every caller inherited that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS, defaulting to 5. Measured against an unroutable host: /api/health went from hanging past 25s to answering 200 in 5.07s. The container healthcheck then probed /api/health, so that hang timed out the check, the container was marked unhealthy, and the platform stopped routing to it. That is the part that turned a degraded dependency into a total outage, and it was introduced with the healthcheck itself. A healthcheck is a LIVENESS question, because the platform's answer to "no" is to take the container out of service. It may only ask whether the process is still serving HTTP. /api/health is a READINESS report - it dials Postgres and Ollama to say whether they are reachable, and coupling the container's existence to its dependencies is what made a running API unreachable. It now probes "/", which is served from memory and does no I/O, so it can fail only if the app really is gone. Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200, and the healthcheck exits 0. With the app stopped it still exits 1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -127,6 +127,14 @@ DB_PORT = os.getenv("DB_PORT", "5432")
|
||||
DB_NAME = os.getenv("DB_NAME", "pgvector")
|
||||
DB_USER = os.getenv("DB_USER", "postgres")
|
||||
DB_PASSWORD = _require("DB_PASSWORD", feature_flag="USE_PGVECTOR") if USE_PGVECTOR else os.getenv("DB_PASSWORD", "")
|
||||
# How long to wait for the TCP connect before giving up. Matters more than it
|
||||
# looks: a host that DROPS packets (a firewall, a typo'd DB_HOST) otherwise
|
||||
# blocks until the OS timeout - about 130 seconds on Linux - and every request
|
||||
# that touches the database inherits that wait, including /api/health. A short
|
||||
# ceiling turns "the database is unreachable" into a fast, honest error instead
|
||||
# of a hung worker and a container the platform decides is unhealthy.
|
||||
DB_CONNECT_TIMEOUT_SECONDS = int(os.getenv("DB_CONNECT_TIMEOUT_SECONDS", "5"))
|
||||
|
||||
DATABASE_URL = os.getenv(
|
||||
"DATABASE_URL",
|
||||
f"postgresql://{DB_USER}:{DB_PASSWORD}@{DB_HOST}:{DB_PORT}/{DB_NAME}",
|
||||
|
||||
@@ -7,7 +7,10 @@ import logging
|
||||
import re
|
||||
import psycopg
|
||||
|
||||
from app.infrastructure.settings import DATABASE_URL, USE_PGVECTOR, DB_HOST, DB_PORT, DB_NAME, DB_USER, DB_PASSWORD
|
||||
from app.infrastructure.settings import (
|
||||
DATABASE_URL, USE_PGVECTOR, DB_HOST, DB_PORT, DB_NAME, DB_USER, DB_PASSWORD,
|
||||
DB_CONNECT_TIMEOUT_SECONDS,
|
||||
)
|
||||
from app.services.brand_registry import BRAND_ALIASES, resolve_parent_brand
|
||||
|
||||
|
||||
@@ -91,7 +94,15 @@ def _connect() -> Optional[psycopg.Connection]:
|
||||
dbname=DB_NAME,
|
||||
user=DB_USER,
|
||||
password=DB_PASSWORD,
|
||||
autocommit=True
|
||||
autocommit=True,
|
||||
# Without this, a host that DROPS packets rather than refusing them
|
||||
# - a firewall, a wrong DB_HOST - blocks here until the OS gives up,
|
||||
# which is around 130 seconds on Linux. Every caller of _connect()
|
||||
# inherits that: /api/health stops answering, the container's
|
||||
# healthcheck times out, and the platform pulls the service out of
|
||||
# its load balancer. "Database unreachable" then presents as a Bad
|
||||
# Gateway on every route, including ones that never touch the DB.
|
||||
connect_timeout=DB_CONNECT_TIMEOUT_SECONDS,
|
||||
)
|
||||
except Exception as e:
|
||||
logger.error(f"Vector DB connection failed: {e}")
|
||||
|
||||
Reference in New Issue
Block a user