app/services/enrichment/barcode/retry.py imports tenacity at module scope,
and that module is on app/main.py's import path (main -> store_catalog router
-> store_catalog_pipeline -> barcode enrichment -> sources -> retry). tenacity
was in no requirements file, so the deployed image exited 1 during startup:
File "/app/app/main.py", line 33, in <module>
from app.api.routers import store_catalog
...
File "/app/app/services/enrichment/barcode/retry.py", line 16
from tenacity import (
ModuleNotFoundError: No module named 'tenacity'
This surfaced as "100% CPU", not as a crash, which is why it was mis-read.
uvicorn never bound a socket, Swarm restarted the task, and each restart
re-ran the ~21s of eager pandas/scipy/sklearn imports that the analytics,
recommendations and nutrition routers pull in at module scope. On this
1-vCPU host that loop pins the only core indefinitely.
retry.py's own docstring asserted tenacity was "already a project dependency
(see requirements.txt)". It never was - corrected to say the opposite, and to
record that it is a hard startup dependency rather than an optional extra.
Verified on the deployed image, not just locally: with tenacity present the
container reaches health=healthy with restarts=0, and /, /api/health,
/api/brands, /api/system/status and /docs all return 200. Idle cost is
0.16% CPU / 168MiB. An AST scan of every import in app/, cli/, scripts/ and
serve.py against the image finds no other missing module (playwright is the
one remaining absence and is deliberate - commented out in requirements.txt
and imported lazily inside a function).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7bVBpxH4AbJtK7Kp3MzDR
The image was 2.45GB, of which the venv is 1.73GB. The build was already
multi-stage and already installed CPU-only torch, so the remaining weight was
not build tooling - it was payload inside the installed packages that the
running service can never execute.
Removed in the build stage, before the runtime stage copies /opt/venv, so the
bytes never enter the final image:
- bundled test suites (~237MB; torch/test is 83MB, pandas/tests 40MB)
- torch/include (62MB), C++ headers for compiling against libtorch
- torch/bin (50MB), gtest binaries and protoc; torch_shm_manager is kept
pytest and httpx2 move to requirements-dev.txt: the image copies app/, cli/,
scripts/, data/ and serve.py, never tests/, so the test stack was unusable
there regardless.
sympy was checked and deliberately kept - 'import sentence_transformers' does
pull it in through torch.fx, so removing it would break embeddings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
_ensure_client() returns None - not False - when USE_OLLAMA is false, because
it returns before it ever probes. SystemStatusOut.ollama_connected is typed
bool, so pydantic rejected the None and the endpoint answered 500 on exactly
the configuration this deployment runs.
Found by smoke-testing the live host: every other read route answered 200 and
this one alone was a server error, which read like a database problem and was
not one. "Ollama is off" now reports as ollama_connected: false.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The frontend loaded but could not call the API at all: every preflight from the
browser came back with no Access-Control-Allow-Origin header, so each request
was blocked client-side while the server logged nothing wrong.
API_CORS_ORIGINS was set to catalogue.nearle.ai.in. The host Traefik actually
serves is catalouge.nearle.ai.in - the o and u transposed. The correctly spelled
domain does not resolve at all, which is why checking the "frontend" only ever
returned a connection error and looked like a network problem.
Both spellings are now listed, so this keeps working if the typo is corrected in
Dokploy later. Verified against the live API: a preflight from
https://catalouge.nearle.ai.in previously returned 400 with no allow-origin.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured inside the image: 2.7GB of nvidia/ CUDA libraries and 691MB of
triton/, on a single-CPU VPS with no GPU. sentence-transformers pulls torch in
transitively, and pip's default Linux wheel bundles the whole CUDA stack because
it cannot know the target has none. That is ~3.4GB of code that can never
execute, in a 9.2GB image on a 48GB disk shared with a dozen other services -
and a build here has already failed once on "no space left on device".
torch now installs first from PyTorch's CPU index, so the requirements.txt pass
finds it satisfied and leaves it alone.
playwright is commented out rather than deleted. It is the last-resort
image-search tier and the Dockerfile never installs its browser binary, so in a
container the tier is skipped at runtime regardless while the package still
costs 137MB. Its import is lazy, so absence takes the existing "not installed"
path rather than breaking anything.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The container was never getting its configuration, so it started with nothing
set and Traefik reported a Bad Gateway on every route.
Dokploy writes its own .env into the build context from the service's
Environment tab AFTER cloning the repository. That tab is empty, so it wrote a
zero-byte file over the committed one, and `COPY .env .` faithfully copied the
empty result into the image. The checkout showed it exactly: every file
timestamped 08:33, and .env alone at 08:34 with a size of 0. Inside the running
container, /app/.env was 0 bytes.
Nothing about this is visible from the outside. The build log shows the COPY
succeeding, the image is produced, and the platform reports only a 502.
Dokploy does not manage .env.production, so the config now travels under that
name and the Dockerfile copies it to /app/.env in the image. Anything set in the
Environment tab still wins at runtime, because settings.py calls load_dotenv()
without override=True.
Verified by reconstructing the build context the way Dokploy does - git archive
of HEAD, then an empty .env written over it - applying .dockerignore and the
COPY lines, and booting the result with an empty environment: /app/.env is 4938
bytes, the sign-in passwords are absent, and scripts/check_deploy.sh reports 6
passed, 0 failed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>