Files
catalogue_backend/README.md
Suriya 5e373ea20c Trim ~350MB of unreachable payload from the backend image
The image was 2.45GB, of which the venv is 1.73GB. The build was already
multi-stage and already installed CPU-only torch, so the remaining weight was
not build tooling - it was payload inside the installed packages that the
running service can never execute.

Removed in the build stage, before the runtime stage copies /opt/venv, so the
bytes never enter the final image:

  - bundled test suites (~237MB; torch/test is 83MB, pandas/tests 40MB)
  - torch/include (62MB), C++ headers for compiling against libtorch
  - torch/bin (50MB), gtest binaries and protoc; torch_shm_manager is kept

pytest and httpx2 move to requirements-dev.txt: the image copies app/, cli/,
scripts/, data/ and serve.py, never tests/, so the test stack was unusable
there regardless.

sympy was checked and deliberately kept - 'import sentence_transformers' does
pull it in through torch.fx, so removing it would break embeddings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 09:52:52 +05:30

210 lines
8.4 KiB
Markdown

# Backend - Brand Product Search Engine (RAG API)
FastAPI service exposing the product catalog through:
- **Browse** - plain listing endpoints (`/api/brands`, `/api/brands/{brand}/products`)
- **Search** - pgvector semantic similarity search, no LLM (`/api/search`)
- **Chat** - full RAG: retrieval + local Ollama generation (`/api/chat`)
- **Admin** - trigger brand ingestion in the background (`/api/catalog/generate`)
Full setup, architecture, and troubleshooting steps are in the project
documentation (`docs/`). This file is just a fast local reference.
## Quick start
### One command (from the project root)
```bash
python run_project.py --backend-only
```
Picks up `backend/venv` if present (else the current interpreter) and starts
uvicorn with autoreload on port 8000. Drop `--backend-only` to run the React
frontend alongside it; `--help` lists the port and reload flags.
It does **not** start Postgres or Ollama for you - it reports them via
`/api/health` and warns if either is unreachable. Bring those up first
(see Manual below, and "Pulling the local LLM").
### Manual
```bash
cd backend
python3 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt -r requirements-dev.txt # -dev is pytest only
cp .env.example .env # then edit DB_PASSWORD etc.
# Auth is required: the app will not start without AUTH_SECRET_KEY and the two
# password hashes. This prints them, plus the sign-in passwords (shown once).
python scripts/make_auth_secrets.py
# Option A: already have a Postgres+pgvector catalog from the old project?
# Just point .env at it (DB_HOST/DB_PORT/DB_NAME/DB_USER/DB_PASSWORD) - done.
# Option B: starting fresh locally?
docker compose -f docker-compose.yml up -d
python scripts/seed_sample_data.py # loads bundled sample catalogs instantly
uvicorn app.main:app --reload --port 8000
```
Then open http://localhost:8000/docs for interactive API docs, or run the
frontend (`../frontend/README.md`) to use the React UI.
## Authentication
Reads are public; the 18 write/compute endpoints require a credential, enforced
by a dependency on each route (`app/api/deps.py`). Sign in for a bearer token:
```bash
curl -X POST localhost:8000/api/auth/login \
-H 'Content-Type: application/json' \
-d '{"username":"admin","password":"<from make_auth_secrets.py>"}'
```
Send it as `Authorization: Bearer <token>`, or use an `X-API-Key` from the
`API_KEYS` setting for server-to-server callers. `admin` passes every
permission check; `user` holds the product/store/inventory permissions.
`AUTH_ENABLED=false` disables all of it for local work — never in a deployment.
See the Authentication section of `../DEPLOYMENT.md` for the full endpoint map.
## MCP server
The catalog is exposed to AI clients over the Model Context Protocol at `/mcp`,
mounted onto this same FastAPI app (`app/mcp_server.py`) - no separate process
or container, so it deploys with the API and answers on the same host.
**15 tools, all read-only.** Catalog search and browse, nutrition facts and
health scores, healthier alternatives, per-store inventory and pricing,
discounts, trending and sales analytics. None of the write or compute endpoints
are reachable through MCP: a tool list is chosen from by a model rather than by
a person, and "retrain the models" is not something to leave one tool call away.
**Auth is the app's own access token** - there is no separate MCP credential.
Clients send `Authorization: Bearer <token>` from `POST /api/auth/login`, and
the same `Principal` and expiry apply as on the REST side. Note the practical
consequence: tokens expire after `AUTH_TOKEN_TTL_MINUTES` (12h by default), so a
long-running client has to refresh. Raise the TTL if that is a problem.
```jsonc
{
"mcpServers": {
"nearle-catalogue": {
"type": "http",
"url": "https://mcp.nearle.ai.in/mcp",
"headers": { "Authorization": "Bearer <token>" }
}
}
}
```
The admin UI has an inspector at `/mcp` (React route, admin-only) listing every
tool with its arguments and a console to run one against live data. It is backed
by `GET /api/mcp/info` and `POST /api/mcp/tools/{name}` rather than by the MCP
endpoint itself - the browser would otherwise need a full MCP client to render a
tool list.
`fastmcp` requires Python 3.10+. The image is 3.11; on an older interpreter the
import fails, `app/main.py` logs a warning, and the REST API starts without the
MCP endpoint rather than failing outright.
## Ports
The container answers on **3000 and 8000 at the same time**, the same way the
frontend image answers on 80 and 3000. 3000 is what Dokploy routes a domain to;
8000 is what this README, the vite dev proxy and `docker-compose.yml` use. Both
being live means the deployment works whichever one it is pointed at, instead of
returning 502 from a healthy container.
`serve.py` is what makes that possible - uvicorn's CLI binds a single `--port`,
but `Server.run()` accepts a list of pre-bound sockets, so it is still one
process. If one port is unavailable it logs and carries on with the other; it
exits non-zero only when nothing is listening.
```bash
python serve.py # 3000 and 8000
PORT=8080 python serve.py # only 8080 (PORT pins a single port)
PORTS=80,3000 python serve.py # a different pair
```
For local development `uvicorn app.main:app --reload --port 8000` is still the
normal thing to run - one port is all you need, and it gives you autoreload.
## Persistence: the two volumes a deployment needs
Most state lives in Postgres, but three things are written to the filesystem,
and in a container those live inside the image - so a redeploy rebuilds the
image and silently discards them:
| Path | Written by |
|---|---|
| `/app/data/seed_catalogs` | `POST /api/user/products/add`, `/upload-file` - every product added through the UI is appended to the brand's JSON |
| `/app/app/intelligence/artifacts` | the training endpoints - every retrained `*.joblib` model |
| `/app/data` | catalogs saved by the ingestion pipeline |
Mount a volume on each (the first is inside the third, so two mounts cover all
three):
```
/app/data
/app/app/intelligence/artifacts
```
In Dokploy, add both under the service's **Volumes**. Named volume or bind
mount, either is fine - `docker compose --profile full up -d` shows the same
two mounts as named volumes.
Bind mounts normally break this pattern, because they start empty and hide the
seed catalogs and pre-trained models the image ships with. They are safe here:
the image keeps read-only copies at `/app/.bundled`, and on startup
`app/infrastructure/persistence.py` copies in whatever the mounted directory is
missing. It never overwrites an existing file, so a user-added product always
survives the next redeploy rather than being reverted to the bundled catalog.
If you leave the volumes off, the API still runs and logs a warning naming the
directories that will be lost.
Override the locations with `DATA_DIR`, `SEED_CATALOG_DIR` and
`MODEL_ARTIFACTS_DIR` if the writable data belongs somewhere else.
## Pulling the local LLM (one-time)
```bash
ollama pull qwen2.5:1.5b
ollama serve # if not already running as a service
```
## Project layout
```
backend/
├── app/
│ ├── main.py # FastAPI app + router wiring
│ ├── infrastructure/settings.py
│ ├── api/
│ │ ├── schemas.py # Pydantic models
│ │ ├── job_store.py # in-memory background-job tracker
│ │ └── routers/ # health, brands, search, chat, catalog
│ ├── core/
│ │ ├── catalog_engine.py # discovery + enrichment + image pipeline
│ │ └── ingestion.py # thin wrapper used by API + CLI
│ └── services/
│ ├── embeddings_service.py # sentence-transformers (lazy-loaded)
│ ├── ollama_service.py # local LLM calls (catalog + RAG answers)
│ ├── vector_store.py # pgvector reads/writes/SEMANTIC SEARCH
│ ├── rag_service.py # RAG orchestration (NEW)
│ ├── image_search.py / s3_service.py / price_estimator.py / brand_registry.py
├── cli/ingest_brand.py # CLI: ingest one brand end-to-end
├── scripts/seed_sample_data.py # load bundled sample catalogs (no LLM needed)
├── data/seed_catalogs/*.json # bundled sample catalogs (Parle, Cadbury, ...)
├── requirements.txt
└── requirements-dev.txt
```
## Running tests
```bash
pytest -q
```