The image was 2.45GB, of which the venv is 1.73GB. The build was already multi-stage and already installed CPU-only torch, so the remaining weight was not build tooling - it was payload inside the installed packages that the running service can never execute. Removed in the build stage, before the runtime stage copies /opt/venv, so the bytes never enter the final image: - bundled test suites (~237MB; torch/test is 83MB, pandas/tests 40MB) - torch/include (62MB), C++ headers for compiling against libtorch - torch/bin (50MB), gtest binaries and protoc; torch_shm_manager is kept pytest and httpx2 move to requirements-dev.txt: the image copies app/, cli/, scripts/, data/ and serve.py, never tests/, so the test stack was unusable there regardless. sympy was checked and deliberately kept - 'import sentence_transformers' does pull it in through torch.fx, so removing it would break embeddings. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
82 lines
3.1 KiB
Plaintext
82 lines
3.1 KiB
Plaintext
# --- Web API ---
|
|
fastapi>=0.115.0
|
|
uvicorn[standard]>=0.30.6
|
|
pydantic>=2.9.2
|
|
python-multipart>=0.0.6
|
|
|
|
# --- Config ---
|
|
python-dotenv>=1.0.1
|
|
|
|
# --- Authentication ---
|
|
# Signs/verifies the access tokens issued by /api/auth/login. Pure Python, no
|
|
# compiled extension - nothing extra to build on the slim image. Password
|
|
# hashing uses hashlib.pbkdf2_hmac from the standard library, so there is
|
|
# deliberately no bcrypt/argon2/passlib dependency here.
|
|
PyJWT>=2.9.0
|
|
|
|
# --- MCP (Model Context Protocol) ---
|
|
# Serves the catalog as tools for AI clients at /mcp (see app/mcp_server.py),
|
|
# mounted onto the FastAPI app so it needs no separate process or container.
|
|
# Requires Python >=3.10; the Dockerfile's python:3.11-slim satisfies that, and
|
|
# app/main.py degrades to serving the REST API alone if the import fails.
|
|
fastmcp>=3.4.7
|
|
|
|
# --- Database / pgvector ---
|
|
psycopg[binary]>=3.2.3
|
|
pgvector>=0.2.5
|
|
|
|
# --- Embeddings (RAG retrieval) ---
|
|
# sentence-transformers pulls in its own CPU-friendly torch wheel
|
|
# automatically - no need to pin torch separately. Imports are LAZY
|
|
# (see app/services/embeddings_service.py) so the API still boots fast
|
|
# even before this is installed/loaded.
|
|
sentence-transformers>=3.0.1
|
|
|
|
# --- HTTP clients ---
|
|
requests>=2.31.0
|
|
aiohttp>=3.9.5
|
|
httpx>=0.27.2
|
|
|
|
# --- Catalog ingestion pipeline (discovery / images / pricing) ---
|
|
# Only required if you plan to run cli/ingest_brand.py or
|
|
# POST /api/catalog/generate to pull in NEW brands. If you only ever use
|
|
# the bundled seed data (scripts/seed_sample_data.py) + RAG chat/search,
|
|
# you can skip everything below this line.
|
|
beautifulsoup4>=4.12.3
|
|
lxml>=4.9.3
|
|
python-slugify>=8.0.4
|
|
ddgs>=9.14.4
|
|
Pillow>=10.0.0
|
|
aiofiles>=23.2.1
|
|
|
|
# Playwright (Python) - last-resort image-search fallback only.
|
|
# Run `playwright install chromium` once after pip install to enable it;
|
|
# the pipeline works fine without it (just one fewer fallback source).
|
|
# playwright>=1.47.0
|
|
# Commented out, not deleted. It is the last-resort image-search tier, and the
|
|
# Dockerfile never installs its browser binary - so in a container the tier is
|
|
# skipped at runtime no matter what, while the package still costs 137MB. The
|
|
# import is lazy (app/services/playwright_image_fallback.py does it inside the
|
|
# function), so its absence is handled on the existing "not installed" path.
|
|
# Uncomment, and run `playwright install chromium`, if you ever want that tier.
|
|
|
|
# --- S3 / DigitalOcean Spaces (optional product image storage) ---
|
|
boto3>=1.34.162
|
|
|
|
# --- ML / analytics (app/intelligence/*, incl. nutrition similarity & ---
|
|
# --- clustering) - CPU-only estimators only (GradientBoosting/Random- ---
|
|
# --- Forest/KMeans/NearestNeighbors), no GPU dependency, per the 8GB ---
|
|
# --- RAM / CPU-only environment this project targets. Was already an ---
|
|
# --- implicit dependency (app/intelligence/*.py imports these) but was ---
|
|
# --- missing from this file - declared explicitly now. ---
|
|
scikit-learn>=1.5.2
|
|
pandas>=2.2.2
|
|
numpy>=1.26.4
|
|
scipy>=1.13.1
|
|
joblib>=1.4.2
|
|
|
|
# --- Dev/test tooling ---
|
|
# Moved to requirements-dev.txt so the Docker image does not carry the test
|
|
# stack it can never run. To work on this project, install both:
|
|
# pip install -r requirements.txt -r requirements-dev.txt
|