Files
catalogue_backend/requirements.txt
Suriya 5e373ea20c Trim ~350MB of unreachable payload from the backend image
The image was 2.45GB, of which the venv is 1.73GB. The build was already
multi-stage and already installed CPU-only torch, so the remaining weight was
not build tooling - it was payload inside the installed packages that the
running service can never execute.

Removed in the build stage, before the runtime stage copies /opt/venv, so the
bytes never enter the final image:

  - bundled test suites (~237MB; torch/test is 83MB, pandas/tests 40MB)
  - torch/include (62MB), C++ headers for compiling against libtorch
  - torch/bin (50MB), gtest binaries and protoc; torch_shm_manager is kept

pytest and httpx2 move to requirements-dev.txt: the image copies app/, cli/,
scripts/, data/ and serve.py, never tests/, so the test stack was unusable
there regardless.

sympy was checked and deliberately kept - 'import sentence_transformers' does
pull it in through torch.fx, so removing it would break embeddings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 09:52:52 +05:30

82 lines
3.1 KiB
Plaintext

# --- Web API ---
fastapi>=0.115.0
uvicorn[standard]>=0.30.6
pydantic>=2.9.2
python-multipart>=0.0.6
# --- Config ---
python-dotenv>=1.0.1
# --- Authentication ---
# Signs/verifies the access tokens issued by /api/auth/login. Pure Python, no
# compiled extension - nothing extra to build on the slim image. Password
# hashing uses hashlib.pbkdf2_hmac from the standard library, so there is
# deliberately no bcrypt/argon2/passlib dependency here.
PyJWT>=2.9.0
# --- MCP (Model Context Protocol) ---
# Serves the catalog as tools for AI clients at /mcp (see app/mcp_server.py),
# mounted onto the FastAPI app so it needs no separate process or container.
# Requires Python >=3.10; the Dockerfile's python:3.11-slim satisfies that, and
# app/main.py degrades to serving the REST API alone if the import fails.
fastmcp>=3.4.7
# --- Database / pgvector ---
psycopg[binary]>=3.2.3
pgvector>=0.2.5
# --- Embeddings (RAG retrieval) ---
# sentence-transformers pulls in its own CPU-friendly torch wheel
# automatically - no need to pin torch separately. Imports are LAZY
# (see app/services/embeddings_service.py) so the API still boots fast
# even before this is installed/loaded.
sentence-transformers>=3.0.1
# --- HTTP clients ---
requests>=2.31.0
aiohttp>=3.9.5
httpx>=0.27.2
# --- Catalog ingestion pipeline (discovery / images / pricing) ---
# Only required if you plan to run cli/ingest_brand.py or
# POST /api/catalog/generate to pull in NEW brands. If you only ever use
# the bundled seed data (scripts/seed_sample_data.py) + RAG chat/search,
# you can skip everything below this line.
beautifulsoup4>=4.12.3
lxml>=4.9.3
python-slugify>=8.0.4
ddgs>=9.14.4
Pillow>=10.0.0
aiofiles>=23.2.1
# Playwright (Python) - last-resort image-search fallback only.
# Run `playwright install chromium` once after pip install to enable it;
# the pipeline works fine without it (just one fewer fallback source).
# playwright>=1.47.0
# Commented out, not deleted. It is the last-resort image-search tier, and the
# Dockerfile never installs its browser binary - so in a container the tier is
# skipped at runtime no matter what, while the package still costs 137MB. The
# import is lazy (app/services/playwright_image_fallback.py does it inside the
# function), so its absence is handled on the existing "not installed" path.
# Uncomment, and run `playwright install chromium`, if you ever want that tier.
# --- S3 / DigitalOcean Spaces (optional product image storage) ---
boto3>=1.34.162
# --- ML / analytics (app/intelligence/*, incl. nutrition similarity & ---
# --- clustering) - CPU-only estimators only (GradientBoosting/Random- ---
# --- Forest/KMeans/NearestNeighbors), no GPU dependency, per the 8GB ---
# --- RAM / CPU-only environment this project targets. Was already an ---
# --- implicit dependency (app/intelligence/*.py imports these) but was ---
# --- missing from this file - declared explicitly now. ---
scikit-learn>=1.5.2
pandas>=2.2.2
numpy>=1.26.4
scipy>=1.13.1
joblib>=1.4.2
# --- Dev/test tooling ---
# Moved to requirements-dev.txt so the Docker image does not carry the test
# stack it can never run. To work on this project, install both:
# pip install -r requirements.txt -r requirements-dev.txt