137 lines
6.2 KiB
Plaintext
137 lines
6.2 KiB
Plaintext
# --- Web API ---
|
|
fastapi>=0.115.0
|
|
uvicorn[standard]>=0.30.6
|
|
pydantic>=2.9.2
|
|
python-multipart>=0.0.6
|
|
|
|
# --- Config ---
|
|
python-dotenv>=1.0.1
|
|
|
|
# --- Authentication ---
|
|
# Signs/verifies the access tokens issued by /api/auth/login. Pure Python, no
|
|
# compiled extension - nothing extra to build on the slim image. Password
|
|
# hashing uses hashlib.pbkdf2_hmac from the standard library, so there is
|
|
# deliberately no bcrypt/argon2/passlib dependency here.
|
|
PyJWT>=2.9.0
|
|
|
|
# --- MCP (Model Context Protocol) ---
|
|
# Serves the catalog as tools for AI clients at /mcp (see app/mcp_server.py),
|
|
# mounted onto the FastAPI app so it needs no separate process or container.
|
|
# Requires Python >=3.10; the Dockerfile's python:3.11-slim satisfies that, and
|
|
# app/main.py degrades to serving the REST API alone if the import fails.
|
|
fastmcp>=3.4.7
|
|
|
|
# --- Database / pgvector ---
|
|
psycopg[binary]>=3.2.3
|
|
pgvector>=0.2.5
|
|
|
|
# --- Embeddings (RAG retrieval) ---
|
|
# sentence-transformers pulls in its own CPU-friendly torch wheel
|
|
# automatically - no need to pin torch separately. Imports are LAZY
|
|
# (see app/services/embeddings_service.py) so the API still boots fast
|
|
# even before this is installed/loaded.
|
|
sentence-transformers>=3.0.1
|
|
|
|
# --- HTTP clients ---
|
|
requests>=2.31.0
|
|
aiohttp>=3.9.5
|
|
httpx>=0.27.2
|
|
|
|
# --- Catalog ingestion pipeline (discovery / images / pricing) ---
|
|
# Only required if you plan to run cli/ingest_brand.py or
|
|
# POST /api/catalog/generate to pull in NEW brands. If you only ever use
|
|
# the bundled seed data (scripts/seed_sample_data.py) + RAG chat/search,
|
|
# you can skip everything below this line.
|
|
# Retry/backoff for the barcode enrichment sources
|
|
# (app/services/enrichment/barcode/retry.py). This is NOT an optional extra:
|
|
# retry.py imports it at module scope, and that module is reached at BOOT, via
|
|
# app/main.py -> routers/{uploads,batch_catalog} -> api/batch_common
|
|
# -> core/store_catalog_pipeline -> enrichment/barcode/retry
|
|
# so a missing tenacity is not a degraded feature, it is the container exiting 1
|
|
# on boot with ModuleNotFoundError before uvicorn ever binds a socket. Swarm
|
|
# then restarts it, and each restart re-runs ~21s of pandas/scipy/sklearn
|
|
# imports, which on a 1-vCPU host reads as pinned-at-100% CPU rather than as a
|
|
# crash.
|
|
#
|
|
# This note used to name the store_catalog router as the importer. That router
|
|
# was deleted with the Store Catalog Ingestion tab; the chain above is the one
|
|
# that survives, and it runs through every remaining ingestion route. Re-check
|
|
# it here before ever concluding this dependency has become unused.
|
|
tenacity>=8.2.3
|
|
|
|
beautifulsoup4>=4.12.3
|
|
lxml>=4.9.3
|
|
python-slugify>=8.0.4
|
|
ddgs>=9.14.4
|
|
Pillow>=10.0.0
|
|
# TFLite runtime for the img_vector image embedder (app/services/image_embedder.py).
|
|
# Runtime only - no TensorFlow - about 20MB; wheels exist for the container's
|
|
# cp311 manylinux and for cp314 Windows. Imported lazily on first use, so a
|
|
# missing wheel means "no image vectors", never a failed boot.
|
|
ai-edge-litert>=2.2.0
|
|
# The Nearle Flutter app preprocesses with OpenCV (centre crop, INTER_AREA
|
|
# resize, 0..1) and its vectors must match ours to a few decimals, so the
|
|
# backend preprocesses with the same library. Headless: no GUI, no libGL.
|
|
# abi3 wheel, ~55MB. Pinned below 5 to stay on the app's major version.
|
|
opencv-python-headless>=4.8,<5
|
|
# Server-side OCR for POST /api/search/identify (app/services/ocr_service.py):
|
|
# rapidocr's PP-OCR models on onnxruntime, CPU. rapidocr ITSELF IS NOT LISTED
|
|
# HERE, on purpose - its metadata demands opencv-python (the GUI build), which
|
|
# a normal install unpacks over opencv-python-headless above and which then
|
|
# fails on libGL.so.1 in the slim image, breaking the image embedder too. It is
|
|
# installed with `--no-deps` from requirements-ocr.txt (see the Dockerfile),
|
|
# and the dependencies it really needs are declared below instead:
|
|
# - onnxruntime: the inference engine, which rapidocr does not declare (it is
|
|
# one of several it can use). Imported lazily; missing means "no server
|
|
# OCR", never a failed boot.
|
|
# - the rest are pure-Python helpers from its metadata (Shapely and pyclipper
|
|
# for the detection polygons, omegaconf/PyYAML for its config).
|
|
onnxruntime>=1.17,<2
|
|
omegaconf>=2.1,!=2.2.1
|
|
pyclipper>=1.2.0
|
|
Shapely>=1.7.1,!=2.0.4
|
|
PyYAML>=6.0
|
|
tqdm>=4.60
|
|
colorlog>=6.0
|
|
six>=1.15.0
|
|
aiofiles>=23.2.1
|
|
|
|
# Playwright (Python) - last-resort image-search fallback only.
|
|
# Run `playwright install chromium` once after pip install to enable it;
|
|
# the pipeline works fine without it (just one fewer fallback source).
|
|
# playwright>=1.47.0
|
|
# Commented out, not deleted. It is the last-resort image-search tier, and the
|
|
# Dockerfile never installs its browser binary - so in a container the tier is
|
|
# skipped at runtime no matter what, while the package still costs 137MB. The
|
|
# import is lazy (app/services/playwright_image_fallback.py does it inside the
|
|
# function), so its absence is handled on the existing "not installed" path.
|
|
# Uncomment, and run `playwright install chromium`, if you ever want that tier.
|
|
|
|
# --- S3 / DigitalOcean Spaces (optional product image storage) ---
|
|
boto3>=1.34.162
|
|
|
|
# --- ML / analytics (app/intelligence/*, incl. nutrition similarity & ---
|
|
# --- clustering) - CPU-only estimators only (GradientBoosting/Random- ---
|
|
# --- Forest/KMeans/NearestNeighbors), no GPU dependency, per the 8GB ---
|
|
# --- RAM / CPU-only environment this project targets. Was already an ---
|
|
# --- implicit dependency (app/intelligence/*.py imports these) but was ---
|
|
# --- missing from this file - declared explicitly now. ---
|
|
scikit-learn>=1.5.2
|
|
pandas>=2.2.2
|
|
# pandas' Excel readers are optional extras it does not install itself, and the
|
|
# upload endpoints (/api/user/products/upload-file, /api/upload/*) accept .xlsx
|
|
# and .xls. Undeclared, they happened to be present in some environments and
|
|
# absent in others - so an Excel upload that worked locally failed in the
|
|
# container with "Missing optional dependency 'openpyxl'". openpyxl reads
|
|
# .xlsx/.xlsm; xlrd is only for the legacy .xls format.
|
|
openpyxl>=3.1.5
|
|
xlrd>=2.0.1
|
|
numpy>=1.26.4
|
|
scipy>=1.13.1
|
|
joblib>=1.4.2
|
|
|
|
# --- Dev/test tooling ---
|
|
# Moved to requirements-dev.txt so the Docker image does not carry the test
|
|
# stack it can never run. To work on this project, install both:
|
|
# pip install -r requirements.txt -r requirements-dev.txt
|