app/services/enrichment/barcode/retry.py imports tenacity at module scope,
and that module is on app/main.py's import path (main -> store_catalog router
-> store_catalog_pipeline -> barcode enrichment -> sources -> retry). tenacity
was in no requirements file, so the deployed image exited 1 during startup:
File "/app/app/main.py", line 33, in <module>
from app.api.routers import store_catalog
...
File "/app/app/services/enrichment/barcode/retry.py", line 16
from tenacity import (
ModuleNotFoundError: No module named 'tenacity'
This surfaced as "100% CPU", not as a crash, which is why it was mis-read.
uvicorn never bound a socket, Swarm restarted the task, and each restart
re-ran the ~21s of eager pandas/scipy/sklearn imports that the analytics,
recommendations and nutrition routers pull in at module scope. On this
1-vCPU host that loop pins the only core indefinitely.
retry.py's own docstring asserted tenacity was "already a project dependency
(see requirements.txt)". It never was - corrected to say the opposite, and to
record that it is a hard startup dependency rather than an optional extra.
Verified on the deployed image, not just locally: with tenacity present the
container reaches health=healthy with restarts=0, and /, /api/health,
/api/brands, /api/system/status and /docs all return 200. Idle cost is
0.16% CPU / 168MiB. An AST scan of every import in app/, cli/, scripts/ and
serve.py against the image finds no other missing module (playwright is the
one remaining absence and is deliberate - commented out in requirements.txt
and imported lazily inside a function).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7bVBpxH4AbJtK7Kp3MzDR
100 lines
4.2 KiB
Plaintext
100 lines
4.2 KiB
Plaintext
# --- Web API ---
|
|
fastapi>=0.115.0
|
|
uvicorn[standard]>=0.30.6
|
|
pydantic>=2.9.2
|
|
python-multipart>=0.0.6
|
|
|
|
# --- Config ---
|
|
python-dotenv>=1.0.1
|
|
|
|
# --- Authentication ---
|
|
# Signs/verifies the access tokens issued by /api/auth/login. Pure Python, no
|
|
# compiled extension - nothing extra to build on the slim image. Password
|
|
# hashing uses hashlib.pbkdf2_hmac from the standard library, so there is
|
|
# deliberately no bcrypt/argon2/passlib dependency here.
|
|
PyJWT>=2.9.0
|
|
|
|
# --- MCP (Model Context Protocol) ---
|
|
# Serves the catalog as tools for AI clients at /mcp (see app/mcp_server.py),
|
|
# mounted onto the FastAPI app so it needs no separate process or container.
|
|
# Requires Python >=3.10; the Dockerfile's python:3.11-slim satisfies that, and
|
|
# app/main.py degrades to serving the REST API alone if the import fails.
|
|
fastmcp>=3.4.7
|
|
|
|
# --- Database / pgvector ---
|
|
psycopg[binary]>=3.2.3
|
|
pgvector>=0.2.5
|
|
|
|
# --- Embeddings (RAG retrieval) ---
|
|
# sentence-transformers pulls in its own CPU-friendly torch wheel
|
|
# automatically - no need to pin torch separately. Imports are LAZY
|
|
# (see app/services/embeddings_service.py) so the API still boots fast
|
|
# even before this is installed/loaded.
|
|
sentence-transformers>=3.0.1
|
|
|
|
# --- HTTP clients ---
|
|
requests>=2.31.0
|
|
aiohttp>=3.9.5
|
|
httpx>=0.27.2
|
|
|
|
# --- Catalog ingestion pipeline (discovery / images / pricing) ---
|
|
# Only required if you plan to run cli/ingest_brand.py or
|
|
# POST /api/catalog/generate to pull in NEW brands. If you only ever use
|
|
# the bundled seed data (scripts/seed_sample_data.py) + RAG chat/search,
|
|
# you can skip everything below this line.
|
|
# Retry/backoff for the barcode enrichment sources
|
|
# (app/services/enrichment/barcode/retry.py). This is NOT an optional extra:
|
|
# retry.py imports it at module scope, and that module is reached from
|
|
# app/main.py's own import of the store_catalog router - so a missing tenacity
|
|
# is not a degraded feature, it is the container exiting 1 on boot with
|
|
# ModuleNotFoundError before uvicorn ever binds a socket. Swarm then restarts
|
|
# it, and each restart re-runs ~21s of pandas/scipy/sklearn imports, which on a
|
|
# 1-vCPU host reads as pinned-at-100% CPU rather than as a crash.
|
|
tenacity>=8.2.3
|
|
|
|
beautifulsoup4>=4.12.3
|
|
lxml>=4.9.3
|
|
python-slugify>=8.0.4
|
|
ddgs>=9.14.4
|
|
Pillow>=10.0.0
|
|
aiofiles>=23.2.1
|
|
|
|
# Playwright (Python) - last-resort image-search fallback only.
|
|
# Run `playwright install chromium` once after pip install to enable it;
|
|
# the pipeline works fine without it (just one fewer fallback source).
|
|
# playwright>=1.47.0
|
|
# Commented out, not deleted. It is the last-resort image-search tier, and the
|
|
# Dockerfile never installs its browser binary - so in a container the tier is
|
|
# skipped at runtime no matter what, while the package still costs 137MB. The
|
|
# import is lazy (app/services/playwright_image_fallback.py does it inside the
|
|
# function), so its absence is handled on the existing "not installed" path.
|
|
# Uncomment, and run `playwright install chromium`, if you ever want that tier.
|
|
|
|
# --- S3 / DigitalOcean Spaces (optional product image storage) ---
|
|
boto3>=1.34.162
|
|
|
|
# --- ML / analytics (app/intelligence/*, incl. nutrition similarity & ---
|
|
# --- clustering) - CPU-only estimators only (GradientBoosting/Random- ---
|
|
# --- Forest/KMeans/NearestNeighbors), no GPU dependency, per the 8GB ---
|
|
# --- RAM / CPU-only environment this project targets. Was already an ---
|
|
# --- implicit dependency (app/intelligence/*.py imports these) but was ---
|
|
# --- missing from this file - declared explicitly now. ---
|
|
scikit-learn>=1.5.2
|
|
pandas>=2.2.2
|
|
# pandas' Excel readers are optional extras it does not install itself, and the
|
|
# upload endpoints (/api/user/products/upload-file, /api/upload/*) accept .xlsx
|
|
# and .xls. Undeclared, they happened to be present in some environments and
|
|
# absent in others - so an Excel upload that worked locally failed in the
|
|
# container with "Missing optional dependency 'openpyxl'". openpyxl reads
|
|
# .xlsx/.xlsm; xlrd is only for the legacy .xls format.
|
|
openpyxl>=3.1.5
|
|
xlrd>=2.0.1
|
|
numpy>=1.26.4
|
|
scipy>=1.13.1
|
|
joblib>=1.4.2
|
|
|
|
# --- Dev/test tooling ---
|
|
# Moved to requirements-dev.txt so the Docker image does not carry the test
|
|
# stack it can never run. To work on this project, install both:
|
|
# pip install -r requirements.txt -r requirements-dev.txt
|