Suriyakumarvijayanayagam d81c3ea18e Fill .env with the real deployment configuration
Adds the live database, DigitalOcean Spaces and Google CSE settings supplied by
the repo owner, so the container needs nothing set in the Dokploy UI.

Three deliberate departures from the development .env this came from:

AUTH_ALLOW_ANY_LOGIN is false, not true. In development it is a convenience -
the password field is not checked, so any username signs in and `admin` reaches
the admin pages. On a host published to the internet it means anyone who finds
mcp.nearle.ai.in becomes admin by typing anything at all. The development file's
own comment says to turn it off before the backend leaves the laptop.

The auth secrets are the freshly generated ones, not the development values.
Those hashes are for the passwords DevAdmin!2026 and DevUser!2026, which sit in
plaintext in test_login_fix.py in this same repository - committing them would
have published working admin credentials alongside the hash that accepts them.
Verified: DevAdmin!2026 is now rejected with a 401.

USE_OLLAMA is false. The development value http://localhost:11434 cannot work
from inside a container, where localhost is the container rather than the VPS
host. Left on with nothing listening, /api/chat fails and every healthcheck
takes ~3s longer, because the health handler probes Ollama with a 3s timeout.
Enable it by pointing OLLAMA_BASE_URL at something the container can reach.

DB_NAME is stated explicitly rather than relying on settings.py's default,
which is what the development file was leaning on.

Verified booting from this file alone, with no environment variables: binds
3000 and 8000, mounts MCP, correct password returns a token, and both the
any-password bypass and the old development password return 401.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:44:26 +05:30
2026-08-11 19:16:01 +05:30
2026-08-11 19:16:01 +05:30
2026-08-12 16:37:28 +05:30
2026-08-12 16:37:28 +05:30
2026-08-12 16:37:28 +05:30

Backend - Brand Product Search Engine (RAG API)

FastAPI service exposing the product catalog through:

  • Browse - plain listing endpoints (/api/brands, /api/brands/{brand}/products)
  • Search - pgvector semantic similarity search, no LLM (/api/search)
  • Chat - full RAG: retrieval + local Ollama generation (/api/chat)
  • Admin - trigger brand ingestion in the background (/api/catalog/generate)

Full setup, architecture, and troubleshooting steps are in the project documentation (docs/). This file is just a fast local reference.

Quick start

One command (from the project root)

python run_project.py --backend-only

Picks up backend/venv if present (else the current interpreter) and starts uvicorn with autoreload on port 8000. Drop --backend-only to run the React frontend alongside it; --help lists the port and reload flags.

It does not start Postgres or Ollama for you - it reports them via /api/health and warns if either is unreachable. Bring those up first (see Manual below, and "Pulling the local LLM").

Manual

cd backend
python3 -m venv .venv
source .venv/bin/activate          # Windows: .venv\Scripts\activate
pip install -r requirements.txt

cp .env.example .env               # then edit DB_PASSWORD etc.

# Auth is required: the app will not start without AUTH_SECRET_KEY and the two
# password hashes. This prints them, plus the sign-in passwords (shown once).
python scripts/make_auth_secrets.py

# Option A: already have a Postgres+pgvector catalog from the old project?
#   Just point .env at it (DB_HOST/DB_PORT/DB_NAME/DB_USER/DB_PASSWORD) - done.
# Option B: starting fresh locally?
docker compose -f docker-compose.yml up -d
python scripts/seed_sample_data.py     # loads bundled sample catalogs instantly

uvicorn app.main:app --reload --port 8000

Then open http://localhost:8000/docs for interactive API docs, or run the frontend (../frontend/README.md) to use the React UI.

Authentication

Reads are public; the 18 write/compute endpoints require a credential, enforced by a dependency on each route (app/api/deps.py). Sign in for a bearer token:

curl -X POST localhost:8000/api/auth/login \
  -H 'Content-Type: application/json' \
  -d '{"username":"admin","password":"<from make_auth_secrets.py>"}'

Send it as Authorization: Bearer <token>, or use an X-API-Key from the API_KEYS setting for server-to-server callers. admin passes every permission check; user holds the product/store/inventory permissions. AUTH_ENABLED=false disables all of it for local work — never in a deployment. See the Authentication section of ../DEPLOYMENT.md for the full endpoint map.

MCP server

The catalog is exposed to AI clients over the Model Context Protocol at /mcp, mounted onto this same FastAPI app (app/mcp_server.py) - no separate process or container, so it deploys with the API and answers on the same host.

15 tools, all read-only. Catalog search and browse, nutrition facts and health scores, healthier alternatives, per-store inventory and pricing, discounts, trending and sales analytics. None of the write or compute endpoints are reachable through MCP: a tool list is chosen from by a model rather than by a person, and "retrain the models" is not something to leave one tool call away.

Auth is the app's own access token - there is no separate MCP credential. Clients send Authorization: Bearer <token> from POST /api/auth/login, and the same Principal and expiry apply as on the REST side. Note the practical consequence: tokens expire after AUTH_TOKEN_TTL_MINUTES (12h by default), so a long-running client has to refresh. Raise the TTL if that is a problem.

{
  "mcpServers": {
    "nearle-catalogue": {
      "type": "http",
      "url": "https://mcp.nearle.ai.in/mcp",
      "headers": { "Authorization": "Bearer <token>" }
    }
  }
}

The admin UI has an inspector at /mcp (React route, admin-only) listing every tool with its arguments and a console to run one against live data. It is backed by GET /api/mcp/info and POST /api/mcp/tools/{name} rather than by the MCP endpoint itself - the browser would otherwise need a full MCP client to render a tool list.

fastmcp requires Python 3.10+. The image is 3.11; on an older interpreter the import fails, app/main.py logs a warning, and the REST API starts without the MCP endpoint rather than failing outright.

Ports

The container answers on 3000 and 8000 at the same time, the same way the frontend image answers on 80 and 3000. 3000 is what Dokploy routes a domain to; 8000 is what this README, the vite dev proxy and docker-compose.yml use. Both being live means the deployment works whichever one it is pointed at, instead of returning 502 from a healthy container.

serve.py is what makes that possible - uvicorn's CLI binds a single --port, but Server.run() accepts a list of pre-bound sockets, so it is still one process. If one port is unavailable it logs and carries on with the other; it exits non-zero only when nothing is listening.

python serve.py               # 3000 and 8000
PORT=8080 python serve.py     # only 8080 (PORT pins a single port)
PORTS=80,3000 python serve.py # a different pair

For local development uvicorn app.main:app --reload --port 8000 is still the normal thing to run - one port is all you need, and it gives you autoreload.

Persistence: the two volumes a deployment needs

Most state lives in Postgres, but three things are written to the filesystem, and in a container those live inside the image - so a redeploy rebuilds the image and silently discards them:

Path Written by
/app/data/seed_catalogs POST /api/user/products/add, /upload-file - every product added through the UI is appended to the brand's JSON
/app/app/intelligence/artifacts the training endpoints - every retrained *.joblib model
/app/data catalogs saved by the ingestion pipeline

Mount a volume on each (the first is inside the third, so two mounts cover all three):

/app/data
/app/app/intelligence/artifacts

In Dokploy, add both under the service's Volumes. Named volume or bind mount, either is fine - docker compose --profile full up -d shows the same two mounts as named volumes.

Bind mounts normally break this pattern, because they start empty and hide the seed catalogs and pre-trained models the image ships with. They are safe here: the image keeps read-only copies at /app/.bundled, and on startup app/infrastructure/persistence.py copies in whatever the mounted directory is missing. It never overwrites an existing file, so a user-added product always survives the next redeploy rather than being reverted to the bundled catalog.

If you leave the volumes off, the API still runs and logs a warning naming the directories that will be lost.

Override the locations with DATA_DIR, SEED_CATALOG_DIR and MODEL_ARTIFACTS_DIR if the writable data belongs somewhere else.

Pulling the local LLM (one-time)

ollama pull qwen2.5:1.5b
ollama serve   # if not already running as a service

Project layout

backend/
├── app/
│   ├── main.py                  # FastAPI app + router wiring
│   ├── infrastructure/settings.py
│   ├── api/
│   │   ├── schemas.py           # Pydantic models
│   │   ├── job_store.py         # in-memory background-job tracker
│   │   └── routers/             # health, brands, search, chat, catalog
│   ├── core/
│   │   ├── catalog_engine.py    # discovery + enrichment + image pipeline
│   │   └── ingestion.py         # thin wrapper used by API + CLI
│   └── services/
│       ├── embeddings_service.py  # sentence-transformers (lazy-loaded)
│       ├── ollama_service.py      # local LLM calls (catalog + RAG answers)
│       ├── vector_store.py        # pgvector reads/writes/SEMANTIC SEARCH
│       ├── rag_service.py         # RAG orchestration (NEW)
│       ├── image_search.py / s3_service.py / price_estimator.py / brand_registry.py
├── cli/ingest_brand.py           # CLI: ingest one brand end-to-end
├── scripts/seed_sample_data.py   # load bundled sample catalogs (no LLM needed)
├── data/seed_catalogs/*.json     # bundled sample catalogs (Parle, Cadbury, ...)
└── requirements.txt

Running tests

pytest -q
Description
No description provided
Readme 22 MiB
Languages
Python 99.4%
Dockerfile 0.3%
Shell 0.3%