Files
catalogue_backend/requirements.txt
Suriyakumarvijayanayagam c101f2c8ba Install CPU-only torch and drop playwright - 9.2GB image on a GPU-less VPS
Measured inside the image: 2.7GB of nvidia/ CUDA libraries and 691MB of
triton/, on a single-CPU VPS with no GPU. sentence-transformers pulls torch in
transitively, and pip's default Linux wheel bundles the whole CUDA stack because
it cannot know the target has none. That is ~3.4GB of code that can never
execute, in a 9.2GB image on a 48GB disk shared with a dozen other services -
and a build here has already failed once on "no space left on device".

torch now installs first from PyTorch's CPU index, so the requirements.txt pass
finds it satisfied and leaves it alone.

playwright is commented out rather than deleted. It is the last-resort
image-search tier and the Dockerfile never installs its browser binary, so in a
container the tier is skipped at runtime regardless while the package still
costs 137MB. Its import is lazy, so absence takes the existing "not installed"
path rather than breaking anything.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 13:06:03 +05:30

85 lines
3.3 KiB
Plaintext

# --- Web API ---
fastapi>=0.115.0
uvicorn[standard]>=0.30.6
pydantic>=2.9.2
python-multipart>=0.0.6
# --- Config ---
python-dotenv>=1.0.1
# --- Authentication ---
# Signs/verifies the access tokens issued by /api/auth/login. Pure Python, no
# compiled extension - nothing extra to build on the slim image. Password
# hashing uses hashlib.pbkdf2_hmac from the standard library, so there is
# deliberately no bcrypt/argon2/passlib dependency here.
PyJWT>=2.9.0
# --- MCP (Model Context Protocol) ---
# Serves the catalog as tools for AI clients at /mcp (see app/mcp_server.py),
# mounted onto the FastAPI app so it needs no separate process or container.
# Requires Python >=3.10; the Dockerfile's python:3.11-slim satisfies that, and
# app/main.py degrades to serving the REST API alone if the import fails.
fastmcp>=3.4.7
# --- Database / pgvector ---
psycopg[binary]>=3.2.3
pgvector>=0.2.5
# --- Embeddings (RAG retrieval) ---
# sentence-transformers pulls in its own CPU-friendly torch wheel
# automatically - no need to pin torch separately. Imports are LAZY
# (see app/services/embeddings_service.py) so the API still boots fast
# even before this is installed/loaded.
sentence-transformers>=3.0.1
# --- HTTP clients ---
requests>=2.31.0
aiohttp>=3.9.5
httpx>=0.27.2
# --- Catalog ingestion pipeline (discovery / images / pricing) ---
# Only required if you plan to run cli/ingest_brand.py or
# POST /api/catalog/generate to pull in NEW brands. If you only ever use
# the bundled seed data (scripts/seed_sample_data.py) + RAG chat/search,
# you can skip everything below this line.
beautifulsoup4>=4.12.3
lxml>=4.9.3
python-slugify>=8.0.4
ddgs>=9.14.4
Pillow>=10.0.0
aiofiles>=23.2.1
# Playwright (Python) - last-resort image-search fallback only.
# Run `playwright install chromium` once after pip install to enable it;
# the pipeline works fine without it (just one fewer fallback source).
# playwright>=1.47.0
# Commented out, not deleted. It is the last-resort image-search tier, and the
# Dockerfile never installs its browser binary - so in a container the tier is
# skipped at runtime no matter what, while the package still costs 137MB. The
# import is lazy (app/services/playwright_image_fallback.py does it inside the
# function), so its absence is handled on the existing "not installed" path.
# Uncomment, and run `playwright install chromium`, if you ever want that tier.
# --- S3 / DigitalOcean Spaces (optional product image storage) ---
boto3>=1.34.162
# --- ML / analytics (app/intelligence/*, incl. nutrition similarity & ---
# --- clustering) - CPU-only estimators only (GradientBoosting/Random- ---
# --- Forest/KMeans/NearestNeighbors), no GPU dependency, per the 8GB ---
# --- RAM / CPU-only environment this project targets. Was already an ---
# --- implicit dependency (app/intelligence/*.py imports these) but was ---
# --- missing from this file - declared explicitly now. ---
scikit-learn>=1.5.2
pandas>=2.2.2
numpy>=1.26.4
scipy>=1.13.1
joblib>=1.4.2
# --- Dev/test tooling ---
pytest>=8.3.3
# Starlette's TestClient deprecates the httpx 0.x backend and emits a
# StarletteDeprecationWarning without this. pytest.ini turns warnings into
# errors, so it is a hard requirement of the suite, not a nicety. Test-only:
# the application itself uses the `httpx` pinned above.
httpx2>=2.10.0