app/services/enrichment/barcode/retry.py imports tenacity at module scope,
and that module is on app/main.py's import path (main -> store_catalog router
-> store_catalog_pipeline -> barcode enrichment -> sources -> retry). tenacity
was in no requirements file, so the deployed image exited 1 during startup:
File "/app/app/main.py", line 33, in <module>
from app.api.routers import store_catalog
...
File "/app/app/services/enrichment/barcode/retry.py", line 16
from tenacity import (
ModuleNotFoundError: No module named 'tenacity'
This surfaced as "100% CPU", not as a crash, which is why it was mis-read.
uvicorn never bound a socket, Swarm restarted the task, and each restart
re-ran the ~21s of eager pandas/scipy/sklearn imports that the analytics,
recommendations and nutrition routers pull in at module scope. On this
1-vCPU host that loop pins the only core indefinitely.
retry.py's own docstring asserted tenacity was "already a project dependency
(see requirements.txt)". It never was - corrected to say the opposite, and to
record that it is a hard startup dependency rather than an optional extra.
Verified on the deployed image, not just locally: with tenacity present the
container reaches health=healthy with restarts=0, and /, /api/health,
/api/brands, /api/system/status and /docs all return 200. Idle cost is
0.16% CPU / 168MiB. An AST scan of every import in app/, cli/, scripts/ and
serve.py against the image finds no other missing module (playwright is the
one remaining absence and is deliberate - commented out in requirements.txt
and imported lazily inside a function).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7bVBpxH4AbJtK7Kp3MzDR
The image was 2.45GB, of which the venv is 1.73GB. The build was already
multi-stage and already installed CPU-only torch, so the remaining weight was
not build tooling - it was payload inside the installed packages that the
running service can never execute.
Removed in the build stage, before the runtime stage copies /opt/venv, so the
bytes never enter the final image:
- bundled test suites (~237MB; torch/test is 83MB, pandas/tests 40MB)
- torch/include (62MB), C++ headers for compiling against libtorch
- torch/bin (50MB), gtest binaries and protoc; torch_shm_manager is kept
pytest and httpx2 move to requirements-dev.txt: the image copies app/, cli/,
scripts/, data/ and serve.py, never tests/, so the test stack was unusable
there regardless.
sympy was checked and deliberately kept - 'import sentence_transformers' does
pull it in through torch.fx, so removing it would break embeddings.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Measured inside the image: 2.7GB of nvidia/ CUDA libraries and 691MB of
triton/, on a single-CPU VPS with no GPU. sentence-transformers pulls torch in
transitively, and pip's default Linux wheel bundles the whole CUDA stack because
it cannot know the target has none. That is ~3.4GB of code that can never
execute, in a 9.2GB image on a 48GB disk shared with a dozen other services -
and a build here has already failed once on "no space left on device".
torch now installs first from PyTorch's CPU index, so the requirements.txt pass
finds it satisfied and leaves it alone.
playwright is commented out rather than deleted. It is the last-resort
image-search tier and the Dockerfile never installs its browser binary, so in a
container the tier is skipped at runtime regardless while the package still
costs 137MB. Its import is lazy, so absence takes the existing "not installed"
path rather than breaking anything.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a FastMCP server mounted onto the existing FastAPI app, so it ships in the
same container and answers on the same host rather than needing a process of
its own.
Fifteen tools, all read-only: catalog search and browse, nutrition facts and
health scores, healthier alternatives, per-store inventory and pricing,
discounts, trending and sales analytics. Each wraps a service function the REST
API already reaches through a GET. None of the write or compute endpoints are
exposed, because a tool list is chosen from by a model rather than by a person,
and catalog generation or model retraining is not something to leave one tool
call away.
The tools are written by hand rather than generated from the OpenAPI schema.
Mirroring all 63 routes would work, but a model picks a tool by reading its
description, and 63 near-identical generated entries is a worse thing to choose
from than a dozen written to be told apart.
Authentication reuses the access token from POST /api/auth/login - no separate
MCP credential, the same Principal and expiry as the REST API. The check lives
in one middleware rather than at the top of each tool, so a tool added later
cannot be left unguarded by forgetting a line. Note that this requires passing
include={"authorization"} to get_http_headers(), which strips that header by
default to avoid forwarding it downstream; without it the header is invisible
and every request looks unauthenticated, valid ones included.
Mounting a sub-app does not run its lifespan - only the outermost app's is
executed - so the MCP app's lifespan is chained through the FastAPI one. Without
that the endpoint accepts a connection and then fails on the first message with
a session manager that was never started.
Adds GET /api/mcp/info and POST /api/mcp/tools/{name} for the admin UI. The MCP
endpoint speaks streamable HTTP with session handling, so rendering a tool list
in the browser would otherwise mean shipping a full MCP client in React.
fastmcp needs Python 3.10+. The image is 3.11; on anything older the import
fails and the REST API starts without the MCP endpoint instead of not starting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>