The image was 2.45GB, of which the venv is 1.73GB. The build was already
multi-stage and already installed CPU-only torch, so the remaining weight was
not build tooling - it was payload inside the installed packages that the
running service can never execute.
Removed in the build stage, before the runtime stage copies /opt/venv, so the
bytes never enter the final image:
- bundled test suites (~237MB; torch/test is 83MB, pandas/tests 40MB)
- torch/include (62MB), C++ headers for compiling against libtorch
- torch/bin (50MB), gtest binaries and protoc; torch_shm_manager is kept
pytest and httpx2 move to requirements-dev.txt: the image copies app/, cli/,
scripts/, data/ and serve.py, never tests/, so the test stack was unusable
there regardless.
sympy was checked and deliberately kept - 'import sentence_transformers' does
pull it in through torch.fx, so removing it would break embeddings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Measured inside the image: 2.7GB of nvidia/ CUDA libraries and 691MB of
triton/, on a single-CPU VPS with no GPU. sentence-transformers pulls torch in
transitively, and pip's default Linux wheel bundles the whole CUDA stack because
it cannot know the target has none. That is ~3.4GB of code that can never
execute, in a 9.2GB image on a 48GB disk shared with a dozen other services -
and a build here has already failed once on "no space left on device".
torch now installs first from PyTorch's CPU index, so the requirements.txt pass
finds it satisfied and leaves it alone.
playwright is commented out rather than deleted. It is the last-resort
image-search tier and the Dockerfile never installs its browser binary, so in a
container the tier is skipped at runtime regardless while the package still
costs 137MB. Its import is lazy, so absence takes the existing "not installed"
path rather than breaking anything.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The container was never getting its configuration, so it started with nothing
set and Traefik reported a Bad Gateway on every route.
Dokploy writes its own .env into the build context from the service's
Environment tab AFTER cloning the repository. That tab is empty, so it wrote a
zero-byte file over the committed one, and `COPY .env .` faithfully copied the
empty result into the image. The checkout showed it exactly: every file
timestamped 08:33, and .env alone at 08:34 with a size of 0. Inside the running
container, /app/.env was 0 bytes.
Nothing about this is visible from the outside. The build log shows the COPY
succeeding, the image is produced, and the platform reports only a 502.
Dokploy does not manage .env.production, so the config now travels under that
name and the Dockerfile copies it to /app/.env in the image. Anything set in the
Environment tab still wins at runtime, because settings.py calls load_dotenv()
without override=True.
Verified by reconstructing the build context the way Dokploy does - git archive
of HEAD, then an empty .env written over it - applying .dockerignore and the
COPY lines, and booting the result with an empty environment: /app/.env is 4938
bytes, the sign-in passwords are absent, and scripts/check_deploy.sh reports 6
passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The configuration was committed but excluded twice over: .dockerignore listed
.env, and the Dockerfile's explicit COPY lines never mentioned it. So the image
built cleanly, the container started with no configuration at all, and exited on
the first required setting - the same RuntimeError and the same Bad Gateway that
committing .env was meant to fix.
Silent in both directions. The build log shows every COPY succeeding, and the
platform reports only a 502, because the process is gone before it can say
anything. Nothing about "build completed" hints that the container has no
configuration.
Verified by reconstructing the image filesystem from the Dockerfile's COPY
lines with .dockerignore applied, then booting from it with an empty
environment: /app/.env is present, SIGNIN_PASSWORDS.txt is not, and /, /docs and
/api/health all answer 200. That reconstruction is the check the earlier "boots
from .env alone" testing was missing - it ran on the host, where .env sits next
to app/ whether the image would have contained it or not.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Two faults compounded into one symptom: with the database unreachable, every
route on the service returned 502 - including /docs, which never touches it.
_connect() passed no connect_timeout. A host that DROPS packets rather than
refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked
until the OS gave up - roughly 130 seconds on Linux. Every caller inherited
that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS,
defaulting to 5. Measured against an unroutable host: /api/health went from
hanging past 25s to answering 200 in 5.07s.
The container healthcheck then probed /api/health, so that hang timed out the
check, the container was marked unhealthy, and the platform stopped routing to
it. That is the part that turned a degraded dependency into a total outage, and
it was introduced with the healthcheck itself.
A healthcheck is a LIVENESS question, because the platform's answer to "no" is
to take the container out of service. It may only ask whether the process is
still serving HTTP. /api/health is a READINESS report - it dials Postgres and
Ollama to say whether they are reachable, and coupling the container's
existence to its dependencies is what made a running API unreachable. It now
probes "/", which is served from memory and does no I/O, so it can fail only if
the app really is gone.
Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200,
and the healthcheck exits 0. With the app stopped it still exits 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Dokploy routes the domain to port 3000, but the container only bound 8000, so
the proxy had nothing to talk to and the domain returned 502 with a perfectly
healthy process behind it.
The frontend image already solved this by answering on both 80 and 3000
(`listen 80; listen 3000;` in nginx.conf). Do the same here rather than swap one
guess for another: 3000 is what the platform routes to, and 8000 is what the
README, the vite dev proxy and docker-compose all target, so binding both means
the container works whichever one it is pointed at.
uvicorn's CLI takes a single --port, but Server.run() accepts pre-bound
sockets, so serve.py binds each port and hands the list to one uvicorn - no
extra worker or second process to supervise. PORT still pins a single port for
anyone who wants one; PORTS changes the pair.
A port that cannot be bound is logged and skipped rather than being fatal,
since losing one of the two should not take down a service the platform only
routes to on the other. It exits non-zero only when nothing is listening at
all, so a genuinely dead container is still reported as failed.
The healthcheck moves into the same file and passes if either port answers,
which keeps it from drifting out of sync with what is actually bound.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
/api/upload/nutrition called json.dumps() in a module that never imported
json, so every request to it raised NameError, was swallowed by the broad
except, and came back as "500 Database import failed". Import json.
Persist the three directories the app writes to at runtime. Products added
through the UI are appended to data/seed_catalogs/*.json and retrained models
are written to app/intelligence/artifacts/*.joblib; both live inside the image,
so a redeploy silently discarded them. The paths now come from settings
(DATA_DIR / SEED_CATALOG_DIR / MODEL_ARTIFACTS_DIR) so a volume can be mounted
on them, and catalog_engine.save_catalog resolves against DATA_DIR instead of
a working-directory-relative "data/", which landed somewhere different
depending on where the process was started from.
Mounting those volumes would otherwise have made things worse: Docker seeds a
named volume from the image on first use, but a bind mount starts empty and
just hides what the image shipped. A bind mount on /app/data would have left
the API with no seed catalogs, so the next product added would write a JSON
file containing only that product. The image now keeps pristine copies at
/app/.bundled, and restore_bundled_assets() tops up whatever a freshly mounted
directory is missing at startup without overwriting anything already there.
Configure CORS for the split-domain deployment: the React app is served from
catalogue.nearle.ai.in and calls the API on mcp.catalogue.nearle.ai.in, so the
frontend origin has to be in API_CORS_ORIGINS. A wrong list fails only in the
browser while the server logs a healthy 200, so the effective origins are now
logged at startup with a warning when they are localhost-only.
Fix FRONTEND_DIST, which looked for a sibling "frontend/" directory that is
actually named "catalogue_frontend/", so the single-port unified-serving branch
could never activate even with a build sitting next to it.
Rebuild the Dockerfile on the frontend's multi-stage pattern: dependencies
resolve into a venv in a build stage, the runtime stage copies only that.
Adds PYTHONUNBUFFERED so startup errors reach Dokploy's log pane, a liveness
HEALTHCHECK (/api/health answers 200 even when Postgres is down, so a database
blip cannot restart-loop the container), and an overridable PORT. The CMD execs
uvicorn so SIGTERM reaches it rather than the sh wrapper.
Add "from __future__ import annotations" to ollama_service and image_search,
which used PEP 604 unions in runtime-evaluated signatures and so could not be
imported below Python 3.10.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>