Files
catalogue_backend/requirements.txt
Suriyakumarvijayanayagam ad7bb9250b Expose the catalog to AI clients over MCP at /mcp
Adds a FastMCP server mounted onto the existing FastAPI app, so it ships in the
same container and answers on the same host rather than needing a process of
its own.

Fifteen tools, all read-only: catalog search and browse, nutrition facts and
health scores, healthier alternatives, per-store inventory and pricing,
discounts, trending and sales analytics. Each wraps a service function the REST
API already reaches through a GET. None of the write or compute endpoints are
exposed, because a tool list is chosen from by a model rather than by a person,
and catalog generation or model retraining is not something to leave one tool
call away.

The tools are written by hand rather than generated from the OpenAPI schema.
Mirroring all 63 routes would work, but a model picks a tool by reading its
description, and 63 near-identical generated entries is a worse thing to choose
from than a dozen written to be told apart.

Authentication reuses the access token from POST /api/auth/login - no separate
MCP credential, the same Principal and expiry as the REST API. The check lives
in one middleware rather than at the top of each tool, so a tool added later
cannot be left unguarded by forgetting a line. Note that this requires passing
include={"authorization"} to get_http_headers(), which strips that header by
default to avoid forwarding it downstream; without it the header is invisible
and every request looks unauthenticated, valid ones included.

Mounting a sub-app does not run its lifespan - only the outermost app's is
executed - so the MCP app's lifespan is chained through the FastAPI one. Without
that the endpoint accepts a connection and then fails on the first message with
a session manager that was never started.

Adds GET /api/mcp/info and POST /api/mcp/tools/{name} for the admin UI. The MCP
endpoint speaks streamable HTTP with session handling, so rendering a tool list
in the browser would otherwise mean shipping a full MCP client in React.

fastmcp needs Python 3.10+. The image is 3.11; on anything older the import
fails and the REST API starts without the MCP endpoint instead of not starting.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:14:00 +05:30

79 lines
2.8 KiB
Plaintext

# --- Web API ---
fastapi>=0.115.0
uvicorn[standard]>=0.30.6
pydantic>=2.9.2
python-multipart>=0.0.6
# --- Config ---
python-dotenv>=1.0.1
# --- Authentication ---
# Signs/verifies the access tokens issued by /api/auth/login. Pure Python, no
# compiled extension - nothing extra to build on the slim image. Password
# hashing uses hashlib.pbkdf2_hmac from the standard library, so there is
# deliberately no bcrypt/argon2/passlib dependency here.
PyJWT>=2.9.0
# --- MCP (Model Context Protocol) ---
# Serves the catalog as tools for AI clients at /mcp (see app/mcp_server.py),
# mounted onto the FastAPI app so it needs no separate process or container.
# Requires Python >=3.10; the Dockerfile's python:3.11-slim satisfies that, and
# app/main.py degrades to serving the REST API alone if the import fails.
fastmcp>=3.4.7
# --- Database / pgvector ---
psycopg[binary]>=3.2.3
pgvector>=0.2.5
# --- Embeddings (RAG retrieval) ---
# sentence-transformers pulls in its own CPU-friendly torch wheel
# automatically - no need to pin torch separately. Imports are LAZY
# (see app/services/embeddings_service.py) so the API still boots fast
# even before this is installed/loaded.
sentence-transformers>=3.0.1
# --- HTTP clients ---
requests>=2.31.0
aiohttp>=3.9.5
httpx>=0.27.2
# --- Catalog ingestion pipeline (discovery / images / pricing) ---
# Only required if you plan to run cli/ingest_brand.py or
# POST /api/catalog/generate to pull in NEW brands. If you only ever use
# the bundled seed data (scripts/seed_sample_data.py) + RAG chat/search,
# you can skip everything below this line.
beautifulsoup4>=4.12.3
lxml>=4.9.3
python-slugify>=8.0.4
ddgs>=9.14.4
Pillow>=10.0.0
aiofiles>=23.2.1
# Playwright (Python) - last-resort image-search fallback only.
# Run `playwright install chromium` once after pip install to enable it;
# the pipeline works fine without it (just one fewer fallback source).
playwright>=1.47.0
# --- S3 / DigitalOcean Spaces (optional product image storage) ---
boto3>=1.34.162
# --- ML / analytics (app/intelligence/*, incl. nutrition similarity & ---
# --- clustering) - CPU-only estimators only (GradientBoosting/Random- ---
# --- Forest/KMeans/NearestNeighbors), no GPU dependency, per the 8GB ---
# --- RAM / CPU-only environment this project targets. Was already an ---
# --- implicit dependency (app/intelligence/*.py imports these) but was ---
# --- missing from this file - declared explicitly now. ---
scikit-learn>=1.5.2
pandas>=2.2.2
numpy>=1.26.4
scipy>=1.13.1
joblib>=1.4.2
# --- Dev/test tooling ---
pytest>=8.3.3
# Starlette's TestClient deprecates the httpx 0.x backend and emits a
# StarletteDeprecationWarning without this. pytest.ini turns warnings into
# errors, so it is a hard requirement of the suite, not a nicety. Test-only:
# the application itself uses the `httpx` pinned above.
httpx2>=2.10.0