Initial commit

This commit is contained in:
sriram
2026-08-10 10:54:58 +05:30
commit d97f14f7a1
147 changed files with 548306 additions and 0 deletions

78
README.md Normal file
View File

@@ -0,0 +1,78 @@
# Kirana AI — RAG-Powered Brand Product Search Engine
This is the v2.0 upgrade of the Indian FMCG product catalog project: the
same discovery/enrichment pipeline (Ollama + pgvector + S3 image
storage), now with a **Retrieval-Augmented Generation (RAG) layer** and a
**React UI** in place of the old Streamlit app.
Two independent pieces:
- **`backend/`** - FastAPI service. Browse/search/chat endpoints over a
pgvector-backed catalog; a RAG chat endpoint that retrieves relevant
products and asks a local Ollama model to answer grounded in them.
- **`frontend/`** - React + Vite single-page app: product search/browse
grid, a conversational "Ask AI" panel with source citations, and an
admin page to ingest new brands.
See **`docs/`** for the full architecture write-up, setup guide, and API
reference. The two component READMEs (`backend/README.md`,
`frontend/README.md`) are quick local references once you've read the
main docs once.
> **v3.0 update:** this project now also includes a **Multi-Store
> Intelligence** layer - 5 simulated stores with independent pricing
> and inventory, ML-based dynamic discounts, trending-product
> detection, a hybrid recommendation engine, and a full analytics
> dashboard. See **`docs/CHANGES.docx`** for the complete write-up
> (architecture decisions, database schema, API reference, and setup
> guide), or jump straight to `backend/scripts/seed_store_intelligence.py`
> and `backend/scripts/train_ml_models.py` to try it.
> **v3.1 update:** added an **AI Nutritional Intelligence** module -
> verified nutrition facts retrieved from Open Food Facts (never LLM-
> generated), transparent/configurable health & nutrition scoring,
> allergen and diet-compatibility detection, ML-based nutritional
> similarity and clustering, healthier-alternative suggestions,
> personalized nutrition recommendations, and a nutrition analytics
> dashboard. See **`docs/NUTRITION_MODULE.docx`** for the complete
> write-up, or jump straight to `backend/scripts/enrich_nutrition.py`
> and `backend/scripts/train_nutrition_models.py` to try it.
## Fastest path to a working demo
```bash
# 1. Backend
cd backend
python3 -m venv .venv && source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env # edit DB_PASSWORD, see docs
docker compose -f ../docker-compose.yml up -d # local Postgres+pgvector (skip if you already have one)
python scripts/seed_sample_data.py # loads bundled sample catalogs - no LLM needed
ollama pull qwen2.5:1.5b # one-time
uvicorn app.main:app --reload --port 8000
# 2. Frontend (new terminal)
cd frontend
npm install
npm run dev
```
Open http://localhost:5173, pick a brand, try the search bar, then switch
to **Ask AI** and ask something like "recommend a low sugar biscuit".
## Why this exists (what changed from v1.0)
The original project (`Project_LLM`) could discover, enrich, and store
products with embeddings in pgvector - but nothing ever read those
embeddings back out. There was no similarity-search function, no API
layer (`app/api/` was an empty folder), and the UI was a Streamlit app
that only ever queried the catalog by exact brand/category, never by
meaning. This project adds the missing retrieval layer
(`vector_store.semantic_search`), a RAG orchestration service
(`rag_service.py`), a full FastAPI surface, and a React UI to use it.
It also fixes a real security issue: `app/infrastructure/settings.py`
used to hard-code a live database host/port/password as Python literal
fallbacks. Every credential now comes only from environment
variables/`.env` - see `backend/app/infrastructure/settings.py` for
details, and `docs/` for the full write-up.