304 lines
12 KiB
Markdown
304 lines
12 KiB
Markdown
# Backend - Brand Product Search Engine (RAG API)
|
|
|
|
FastAPI service exposing the product catalog through:
|
|
|
|
- **Browse** - plain listing endpoints (`/api/brands`, `/api/brands/{brand}/products`)
|
|
- **Search** - pgvector semantic similarity search, no LLM (`/api/search`)
|
|
- **Chat** - full RAG: retrieval + local Ollama generation (`/api/chat`)
|
|
- **Admin** - trigger brand ingestion in the background (`/api/catalog/generate`)
|
|
|
|
Full setup, architecture, and troubleshooting steps are in the project
|
|
documentation (`docs/`). This file is just a fast local reference.
|
|
|
|
## Quick start
|
|
|
|
### One command (from the project root)
|
|
|
|
```bash
|
|
python run_project.py --backend-only
|
|
```
|
|
|
|
Picks up `backend/venv` if present (else the current interpreter) and starts
|
|
uvicorn with autoreload on port 8000. Drop `--backend-only` to run the React
|
|
frontend alongside it; `--help` lists the port and reload flags.
|
|
|
|
It does **not** start Postgres or Ollama for you - it reports them via
|
|
`/api/health` and warns if either is unreachable. Bring those up first
|
|
(see Manual below, and "Pulling the local LLM").
|
|
|
|
### Manual
|
|
|
|
```bash
|
|
cd backend
|
|
python3 -m venv .venv
|
|
source .venv/bin/activate # Windows: .venv\Scripts\activate
|
|
pip install -r requirements.txt -r requirements-dev.txt # -dev is pytest only
|
|
|
|
cp .env.example .env # then edit DB_PASSWORD etc.
|
|
|
|
# Auth is required: the app will not start without AUTH_SECRET_KEY and the two
|
|
# password hashes. This prints them, plus the sign-in passwords (shown once).
|
|
python scripts/make_auth_secrets.py
|
|
|
|
# Option A: already have a Postgres+pgvector catalog from the old project?
|
|
# Just point .env at it (DB_HOST/DB_PORT/DB_NAME/DB_USER/DB_PASSWORD) - done.
|
|
# Option B: starting fresh locally?
|
|
docker compose -f docker-compose.yml up -d
|
|
python scripts/seed_sample_data.py # loads bundled sample catalogs instantly
|
|
|
|
uvicorn app.main:app --reload --port 8000
|
|
```
|
|
|
|
Then open http://localhost:8000/docs for interactive API docs, or run the
|
|
frontend (`../frontend/README.md`) to use the React UI.
|
|
|
|
## Active brands (development working set)
|
|
|
|
One setting decides which brands the application and every pipeline work on:
|
|
|
|
```
|
|
# backend/.env
|
|
ACTIVE_BRANDS=Amul,Cadbury,Hindustan Unilever
|
|
```
|
|
|
|
**Blank or unset means every brand is active** - that is the production default
|
|
and the way to turn the feature off.
|
|
|
|
Nothing is deleted when it is set. The other `brand_*` tables and their
|
|
embeddings stay in Postgres untouched; they simply stop being discovered.
|
|
Everything downstream inherits it, because the whole app funnels through two
|
|
functions in `app/services/vector_store.py`
|
|
(`list_available_brands` and `_list_brand_table_suffixes`):
|
|
|
|
```
|
|
ACTIVE_BRANDS
|
|
|
|
|
+-- /api/brands, /api/products, /api/search, /api/suggest
|
|
+-- RAG chat and the query_intent brand index
|
|
+-- MCP tools
|
|
+-- nutrition enrichment, store intelligence, analytics
|
|
+-- the boot auto-seed and the 300s brand reconcile
|
|
+-- every Dagster asset and partition
|
|
```
|
|
|
|
Going from 3 brands to 5, 10 or all of them is an edit to this one line - no
|
|
code changes. Catalogs for inactive brands live in
|
|
`data/seed_catalogs/archive/`, which the loader still reads, so re-activating a
|
|
brand does not require moving files back.
|
|
|
|
Names resolve through the brand aliases, so `ACTIVE_BRANDS=Tata` activates
|
|
`brand_hindustan_unilever` - the same table an ingest of Tata products targets.
|
|
|
|
## Orchestration (Dagster)
|
|
|
|
Ingestion, enrichment, embedding and ML training can be run as a Dagster asset
|
|
graph, with lineage, retries and run history. It is a **development tool** -
|
|
FastAPI still serves every request and nothing in a request path touches it.
|
|
|
|
```
|
|
pip install -r requirements-orchestration.txt
|
|
DAGSTER_HOME="$(pwd)/orchestration/.dagster_home" dagster dev -m orchestration.definitions -p 3030
|
|
```
|
|
|
|
See `orchestration/README.md`. Note that it pins itself to the **local**
|
|
database and refuses to write to a remote one, because `backend/.env` points at
|
|
production.
|
|
|
|
## Authentication
|
|
|
|
Reads are public; the 18 write/compute endpoints require a credential, enforced
|
|
by a dependency on each route (`app/api/deps.py`). Sign in for a bearer token:
|
|
|
|
```bash
|
|
curl -X POST localhost:8000/api/auth/login \
|
|
-H 'Content-Type: application/json' \
|
|
-d '{"username":"admin","password":"<from make_auth_secrets.py>"}'
|
|
```
|
|
|
|
Send it as `Authorization: Bearer <token>`, or use an `X-API-Key` from the
|
|
`API_KEYS` setting for server-to-server callers. `admin` passes every
|
|
permission check; `user` holds the product/store/inventory permissions;
|
|
`uploader` holds exactly one, `upload_catalog` (see below).
|
|
`AUTH_ENABLED=false` disables all of it for local work — never in a deployment.
|
|
See the Authentication section of `../DEPLOYMENT.md` for the full endpoint map.
|
|
|
|
## Catalog ingestion for API clients
|
|
|
|
An outside party can send spreadsheets straight into the catalog pipeline
|
|
without an admin in the loop. Issue them an `uploader` key:
|
|
|
|
```bash
|
|
python scripts/make_auth_secrets.py --api-key catalog-drop:uploader
|
|
# -> API_KEYS=catalog-drop:uploader:<secret> (add to .env, redeploy)
|
|
```
|
|
|
|
Name the key for its function, not its holder: `/api/health` publicly reports
|
|
every key's name, role and fingerprint (never the secret).
|
|
|
|
They then POST files and poll the batch:
|
|
|
|
```bash
|
|
curl -X POST https://mcp.nearle.ai.in/api/uploads/catalog \
|
|
-H 'X-API-Key: <secret>' \
|
|
-F 'files=@store-catalog.xlsx' -F 'files=@second-store.xlsx'
|
|
# -> 202 {"batch_id": "...", "status": "queued", "files": [...], "message": "..."}
|
|
|
|
curl https://mcp.nearle.ai.in/api/uploads/catalog/<batch_id> -H 'X-API-Key: <secret>'
|
|
# -> {"status": "running", "files_done": 1, "files": [{"stage_name": "...", ...}]}
|
|
```
|
|
|
|
`.xlsx`, `.xls` and `.csv` are accepted, up to 10MB / 2000 rows per file and
|
|
`BATCH_MAX_FILES` files per request. Each file runs the same 11 stages as the
|
|
admin route (`app/core/store_catalog_pipeline.py`). A sheet that cannot be
|
|
parsed — or that has no product-name column — is rejected during the request
|
|
with a 400 naming the problem, so the sender finds out while they can still fix
|
|
it; a bad file alongside good ones comes back in `files` as `status: "failed"`
|
|
while the rest still run.
|
|
|
|
Two things this credential cannot do. It cannot see anything but its own
|
|
submissions — every read is filtered by `submitted_by`, so it reaches neither
|
|
the catalog nor another caller's batches — and it cannot multiply the work:
|
|
all ingestion, from every source, goes through one worker thread behind a queue
|
|
of `BATCH_QUEUE_MAX`, past which the endpoint answers 429. Cancel, resume and
|
|
the full batch list stay on the admin router
|
|
(`/api/admin/catalog-batch/...`, `require_admin`).
|
|
|
|
## MCP server
|
|
|
|
The catalog is exposed to AI clients over the Model Context Protocol at `/mcp`,
|
|
mounted onto this same FastAPI app (`app/mcp_server.py`) - no separate process
|
|
or container, so it deploys with the API and answers on the same host.
|
|
|
|
**15 tools, all read-only.** Catalog search and browse, nutrition facts and
|
|
health scores, healthier alternatives, per-store inventory and pricing,
|
|
discounts, trending and sales analytics. None of the write or compute endpoints
|
|
are reachable through MCP: a tool list is chosen from by a model rather than by
|
|
a person, and "retrain the models" is not something to leave one tool call away.
|
|
|
|
**Auth is the app's own access token** - there is no separate MCP credential.
|
|
Clients send `Authorization: Bearer <token>` from `POST /api/auth/login`, and
|
|
the same `Principal` and expiry apply as on the REST side. Note the practical
|
|
consequence: tokens expire after `AUTH_TOKEN_TTL_MINUTES` (12h by default), so a
|
|
long-running client has to refresh. Raise the TTL if that is a problem.
|
|
|
|
```jsonc
|
|
{
|
|
"mcpServers": {
|
|
"nearle-catalogue": {
|
|
"type": "http",
|
|
"url": "https://mcp.nearle.ai.in/mcp",
|
|
"headers": { "Authorization": "Bearer <token>" }
|
|
}
|
|
}
|
|
}
|
|
```
|
|
|
|
The admin UI has an inspector at `/mcp` (React route, admin-only) listing every
|
|
tool with its arguments and a console to run one against live data. It is backed
|
|
by `GET /api/mcp/info` and `POST /api/mcp/tools/{name}` rather than by the MCP
|
|
endpoint itself - the browser would otherwise need a full MCP client to render a
|
|
tool list.
|
|
|
|
`fastmcp` requires Python 3.10+. The image is 3.11; on an older interpreter the
|
|
import fails, `app/main.py` logs a warning, and the REST API starts without the
|
|
MCP endpoint rather than failing outright.
|
|
|
|
## Ports
|
|
|
|
The container answers on **3000 and 8000 at the same time**, the same way the
|
|
frontend image answers on 80 and 3000. 3000 is what Dokploy routes a domain to;
|
|
8000 is what this README, the vite dev proxy and `docker-compose.yml` use. Both
|
|
being live means the deployment works whichever one it is pointed at, instead of
|
|
returning 502 from a healthy container.
|
|
|
|
`serve.py` is what makes that possible - uvicorn's CLI binds a single `--port`,
|
|
but `Server.run()` accepts a list of pre-bound sockets, so it is still one
|
|
process. If one port is unavailable it logs and carries on with the other; it
|
|
exits non-zero only when nothing is listening.
|
|
|
|
```bash
|
|
python serve.py # 3000 and 8000
|
|
PORT=8080 python serve.py # only 8080 (PORT pins a single port)
|
|
PORTS=80,3000 python serve.py # a different pair
|
|
```
|
|
|
|
For local development `uvicorn app.main:app --reload --port 8000` is still the
|
|
normal thing to run - one port is all you need, and it gives you autoreload.
|
|
|
|
## Persistence: the two volumes a deployment needs
|
|
|
|
Most state lives in Postgres, but three things are written to the filesystem,
|
|
and in a container those live inside the image - so a redeploy rebuilds the
|
|
image and silently discards them:
|
|
|
|
| Path | Written by |
|
|
|---|---|
|
|
| `/app/data/seed_catalogs` | `POST /api/user/products/add`, `/upload-file` - every product added through the UI is appended to the brand's JSON |
|
|
| `/app/app/intelligence/artifacts` | the training endpoints - every retrained `*.joblib` model |
|
|
| `/app/data` | catalogs saved by the ingestion pipeline |
|
|
|
|
Mount a volume on each (the first is inside the third, so two mounts cover all
|
|
three):
|
|
|
|
```
|
|
/app/data
|
|
/app/app/intelligence/artifacts
|
|
```
|
|
|
|
In Dokploy, add both under the service's **Volumes**. Named volume or bind
|
|
mount, either is fine - `docker compose --profile full up -d` shows the same
|
|
two mounts as named volumes.
|
|
|
|
Bind mounts normally break this pattern, because they start empty and hide the
|
|
seed catalogs and pre-trained models the image ships with. They are safe here:
|
|
the image keeps read-only copies at `/app/.bundled`, and on startup
|
|
`app/infrastructure/persistence.py` copies in whatever the mounted directory is
|
|
missing. It never overwrites an existing file, so a user-added product always
|
|
survives the next redeploy rather than being reverted to the bundled catalog.
|
|
|
|
If you leave the volumes off, the API still runs and logs a warning naming the
|
|
directories that will be lost.
|
|
|
|
Override the locations with `DATA_DIR`, `SEED_CATALOG_DIR` and
|
|
`MODEL_ARTIFACTS_DIR` if the writable data belongs somewhere else.
|
|
|
|
## Pulling the local LLM (one-time)
|
|
|
|
```bash
|
|
ollama pull qwen2.5:1.5b
|
|
ollama serve # if not already running as a service
|
|
```
|
|
|
|
## Project layout
|
|
|
|
```
|
|
backend/
|
|
├── app/
|
|
│ ├── main.py # FastAPI app + router wiring
|
|
│ ├── infrastructure/settings.py
|
|
│ ├── api/
|
|
│ │ ├── schemas.py # Pydantic models
|
|
│ │ ├── job_store.py # in-memory background-job tracker
|
|
│ │ └── routers/ # health, brands, search, chat, catalog
|
|
│ ├── core/
|
|
│ │ ├── catalog_engine.py # discovery + enrichment + image pipeline
|
|
│ │ └── ingestion.py # thin wrapper used by API + CLI
|
|
│ └── services/
|
|
│ ├── embeddings_service.py # sentence-transformers (lazy-loaded)
|
|
│ ├── ollama_service.py # local LLM calls (catalog + RAG answers)
|
|
│ ├── vector_store.py # pgvector reads/writes/SEMANTIC SEARCH
|
|
│ ├── rag_service.py # RAG orchestration (NEW)
|
|
│ ├── image_search.py / s3_service.py / price_estimator.py / brand_registry.py
|
|
├── cli/ingest_brand.py # CLI: ingest one brand end-to-end
|
|
├── scripts/seed_sample_data.py # load bundled sample catalogs (no LLM needed)
|
|
├── data/seed_catalogs/*.json # bundled sample catalogs (Parle, Cadbury, ...)
|
|
├── requirements.txt
|
|
└── requirements-dev.txt
|
|
```
|
|
|
|
## Running tests
|
|
|
|
```bash
|
|
pytest -q
|
|
```
|