RAG-Powered Brand Product Search Engine
This is the v2.0 upgrade of the Indian FMCG product catalog project: the same discovery/enrichment pipeline (Ollama + pgvector + S3 image storage), now with a Retrieval-Augmented Generation (RAG) layer and a React UI in place of the old Streamlit app.
Two independent pieces:
backend/- FastAPI service. Browse/search/chat endpoints over a pgvector-backed catalog; a RAG chat endpoint that retrieves relevant products and asks a local Ollama model to answer grounded in them.frontend/- React + Vite single-page app: product search/browse grid, a conversational "Ask AI" panel with source citations, and an admin page to ingest new brands.
See docs/ for the full architecture write-up, setup guide, and API
reference. The two component READMEs (backend/README.md,
frontend/README.md) are quick local references once you've read the
main docs once.
v3.0 update: this project now also includes a Multi-Store Intelligence layer - 5 simulated stores with independent pricing and inventory, ML-based dynamic discounts, trending-product detection, a hybrid recommendation engine, and a full analytics dashboard. See
docs/CHANGES.docxfor the complete write-up (architecture decisions, database schema, API reference, and setup guide), or jump straight tobackend/scripts/seed_store_intelligence.pyandbackend/scripts/train_ml_models.pyto try it.
v3.1 update: added an AI Nutritional Intelligence module - verified nutrition facts retrieved from Open Food Facts (never LLM- generated), transparent/configurable health & nutrition scoring, allergen and diet-compatibility detection, ML-based nutritional similarity and clustering, healthier-alternative suggestions, personalized nutrition recommendations, and a nutrition analytics dashboard. See
docs/NUTRITION_MODULE.docxfor the complete write-up, or jump straight tobackend/scripts/enrich_nutrition.pyandbackend/scripts/train_nutrition_models.pyto try it.
Fastest Path to Execute Project (Instant 1-Command Startup)
The project now includes a unified single-command launcher and automated background initialization. You no longer need to manually run multiple seed scripts or manage separate terminals:
# Single command from project root (starts Docker, Backend, and Frontend):
python run_project.py
# Or on Windows, double-click:
start_app.bat
The system automatically:
- Starts the PostgreSQL container via Docker Desktop.
- Checks database status and skips redundant seeding for instant boot (< 2 seconds).
- Launches Backend API (
http://localhost:8000) and Frontend UI (http://localhost:5173) concurrently. - Auto-seeds missing catalogs or store intelligence in the background if the database is empty.
Individual Commands (Optional / Advanced)
- Backend:
start_backend.batorpython -m uvicorn app.main:app --reload --port 8000 - Frontend:
start_frontend.batornpm run dev - System Status API:
http://localhost:8000/api/system/status
Why this exists (what changed from v1.0)
The original project (Project_LLM) could discover, enrich, and store
products with embeddings in pgvector - but nothing ever read those
embeddings back out. There was no similarity-search function, no API
layer (app/api/ was an empty folder), and the UI was a Streamlit app
that only ever queried the catalog by exact brand/category, never by
meaning. This project adds the missing retrieval layer
(vector_store.semantic_search), a RAG orchestration service
(rag_service.py), a full FastAPI surface, and a React UI to use it.
It also fixes a real security issue: app/infrastructure/settings.py
used to hard-code a live database host/port/password as Python literal
fallbacks. Every credential now comes only from environment
variables/.env - see backend/app/infrastructure/settings.py for
details, and docs/ for the full write-up.