Two faults compounded into one symptom: with the database unreachable, every route on the service returned 502 - including /docs, which never touches it. _connect() passed no connect_timeout. A host that DROPS packets rather than refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked until the OS gave up - roughly 130 seconds on Linux. Every caller inherited that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS, defaulting to 5. Measured against an unroutable host: /api/health went from hanging past 25s to answering 200 in 5.07s. The container healthcheck then probed /api/health, so that hang timed out the check, the container was marked unhealthy, and the platform stopped routing to it. That is the part that turned a degraded dependency into a total outage, and it was introduced with the healthcheck itself. A healthcheck is a LIVENESS question, because the platform's answer to "no" is to take the container out of service. It may only ask whether the process is still serving HTTP. /api/health is a READINESS report - it dials Postgres and Ollama to say whether they are reachable, and coupling the container's existence to its dependencies is what made a running API unreachable. It now probes "/", which is served from memory and does no I/O, so it can fail only if the app really is gone. Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200, and the healthcheck exits 0. With the app stopped it still exits 1. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
100 lines
4.7 KiB
Docker
100 lines
4.7 KiB
Docker
# Multi-stage, mirroring catalogue_frontend/Dockerfile: a build stage that
|
|
# resolves dependencies, then a clean runtime stage that copies in only the
|
|
# result. There it is `npm ci` -> dist/; here it is pip -> a virtualenv.
|
|
|
|
# ---- Build stage ----
|
|
FROM python:3.11-slim AS build
|
|
WORKDIR /app
|
|
|
|
# Dependencies land in a self-contained venv so the runtime stage can take that
|
|
# one directory and leave pip, its HTTP cache and the downloaded wheels behind.
|
|
#
|
|
# requirements.txt is copied on its own, ahead of the source, for the same
|
|
# reason the frontend stage copies package.json before the rest of the app:
|
|
# this layer is cached on the file's checksum, so editing a router does not
|
|
# reinstall torch.
|
|
#
|
|
# psycopg[binary] avoids needing libpq-dev; sentence-transformers/scikit-learn/
|
|
# scipy all ship prebuilt wheels for this image, so no compiler is needed and
|
|
# this stage installs no build toolchain. Playwright's Python package installs,
|
|
# but its browser binary is NOT installed here - it's only a last-resort
|
|
# image-search fallback (see requirements.txt); run `playwright install
|
|
# chromium` in the container if you need that specific fallback tier.
|
|
COPY requirements.txt .
|
|
RUN python -m venv /opt/venv \
|
|
&& /opt/venv/bin/pip install --no-cache-dir -r requirements.txt
|
|
|
|
|
|
# ---- Runtime stage ----
|
|
FROM python:3.11-slim AS runtime
|
|
WORKDIR /app
|
|
|
|
# PATH: putting the venv first is what makes a bare `python`/`uvicorn` resolve
|
|
# to it - there is no "activate" step in a container.
|
|
# PYTHONUNBUFFERED: without it Dokploy's log view stays empty until a buffer
|
|
# happens to fill, so startup errors surface minutes after the container died.
|
|
# PYTHONDONTWRITEBYTECODE: no .pyc to write into a read-only-ish image layer.
|
|
ENV PATH="/opt/venv/bin:$PATH"
|
|
ENV PYTHONUNBUFFERED=1
|
|
ENV PYTHONDONTWRITEBYTECODE=1
|
|
|
|
COPY --from=build /opt/venv /opt/venv
|
|
|
|
COPY app ./app
|
|
COPY cli ./cli
|
|
COPY scripts ./scripts
|
|
COPY data ./data
|
|
COPY serve.py .
|
|
|
|
# Pristine copies of everything the app also WRITES to, kept at a path that is
|
|
# never mounted over.
|
|
#
|
|
# /app/data/seed_catalogs and /app/app/intelligence/artifacts both need volumes
|
|
# (products added through the UI are appended to the first, retrained models
|
|
# are written to the second - otherwise a redeploy throws both away). But
|
|
# mounting a volume there hides the copies shipped in this image: a *named*
|
|
# volume is seeded from the image on first use, a *bind* mount is not, and
|
|
# Dokploy offers both. A bind mount would leave the API with zero seed catalogs
|
|
# and zero trained models, with nothing in the logs saying why.
|
|
#
|
|
# So keep a second copy here. On startup restore_bundled_assets()
|
|
# (app/infrastructure/persistence.py) copies in whatever the mounted directory
|
|
# is missing, and never overwrites what is already there.
|
|
RUN mkdir -p /app/.bundled \
|
|
&& cp -a /app/data/seed_catalogs /app/.bundled/seed_catalogs \
|
|
&& cp -a /app/app/intelligence/artifacts /app/.bundled/artifacts
|
|
|
|
# Declared so `docker run` without an explicit -v still gets an anonymous
|
|
# volume rather than writing into the container layer. Dokploy (and the compose
|
|
# file) name them properly; this is the floor, not the recommended setup.
|
|
VOLUME ["/app/data", "/app/app/intelligence/artifacts"]
|
|
|
|
# Answers on BOTH ports, the same way the frontend image does (nginx.conf has
|
|
# `listen 80; listen 3000;`). 3000 is what Dokploy routes a domain to; 8000 is
|
|
# what this project's README, the vite dev proxy and docker-compose all use.
|
|
# Serving both means the container works whichever one the platform is pointed
|
|
# at, instead of returning 502 from a perfectly healthy process.
|
|
#
|
|
# serve.py binds both sockets and hands them to one uvicorn - see the note
|
|
# there. To pin a single port, set PORT (PORT=8080 serves only 8080); to change
|
|
# the pair, set PORTS.
|
|
ENV PORTS=3000,8000
|
|
EXPOSE 3000 8000
|
|
|
|
# Liveness only, and passes if EITHER port answers.
|
|
#
|
|
# It probes "/", which is served from memory, NOT /api/health, which dials
|
|
# Postgres and Ollama. That is the whole point: the platform's response to a
|
|
# failed healthcheck is to stop routing traffic, so this may only ask "is the
|
|
# process still serving HTTP". Tying it to the database meant an unreachable
|
|
# Postgres blocked the handler for the OS TCP timeout, the check timed out, the
|
|
# container was marked unhealthy, and a perfectly healthy API returned Bad
|
|
# Gateway on every route. Use /api/health to ask whether dependencies are up;
|
|
# it reports them in the body and always answers 200.
|
|
HEALTHCHECK --interval=30s --timeout=10s --start-period=40s --retries=3 \
|
|
CMD ["python", "serve.py", "--healthcheck"]
|
|
|
|
# Exec form: python is PID 1, so Docker's SIGTERM reaches it directly and a
|
|
# redeploy shuts down cleanly instead of waiting out the 10s kill timeout.
|
|
CMD ["python", "serve.py"]
|