imag vector generation with dimentionality reduction

This commit is contained in:
sriram
2026-09-18 15:26:25 +05:30
parent afa0bfa743
commit d6296bd1f0
16 changed files with 1724 additions and 48 deletions

141
docs/IMAGE_SEARCH_API.md Normal file
View File

@@ -0,0 +1,141 @@
# Search the catalogue by photo
Base: `https://mcp.nearle.ai.in` · Auth: **none** (public, like `GET /api/search`) · Read-only
The Nearle app photographs a pack, crops it, embeds it on-device with
MobileNetV3-Small (1024 floats, L2-normalised) and reads the label with OCR.
Every catalogue row carries the same kind of vector in `img_vector`
(same model, same OpenCV preprocessing, `vector(1024)` with an hnsw cosine
index). These two endpoints turn the app's vector - or a photo - into
product cards.
```
POST /api/search/image-vector JSON {vector[1024], text?, brand?, category?, top_k?, min_score?}
POST /api/search/image multipart file + the same optional fields as form fields
```
Both run the same ranking. Use the first from the app (it already has the
vector); use the second from anything without the model, or for testing.
## How a match is found
1. **Scope.** An explicit `brand` searches that one brand table, full stop.
Otherwise the OCR `text` is checked for a brand name ("Britannia Marie
Gold 300 g" → `brand_britannia`); if none is recognised, every brand table
is searched. A brand recognised from OCR that yields nothing above
`min_score` is retried across every brand (`scope_fallback: true`) - OCR
misreads happen; an explicit `brand` never falls back.
2. **Rank by cosine.** `score = 1 - (img_vector <=> vector)`. Both sides are
unit vectors, so this is cosine similarity; 1.0 is the identical picture.
3. **Break ties with the label.** Every pack size of one product usually
shares one catalogue photo, so "Marie Gold 89g / 300g / 1kg" tie exactly.
Size tokens in `text` ("300 g", "300gm", "300G" all read as `300g`) count
3 points each, other words 1 point, matched against the product name and
size variants. The row with the most points wins the tie; then name order.
4. Return `top_k` cards.
`score` is the honest number: a phone photo of a pack against the
catalogue's render of it measured **0.63** in testing; identical files give
0.99+. The app team's "above 0.7 means the same product" is a phone-vs-phone
rule of thumb. The server does not impose it - pass `min_score` if you want
a floor, and read `score` on each result.
## Request fields
| Field | Where | Default | Notes |
|---|---|---|---|
| `vector` | JSON only | required | exactly 1024 finite floats, not all zero; re-normalised if not unit length |
| `file` | multipart only | required | JPEG/PNG/WebP, ≤ 8 MB, ideally cropped to the pack; transparency is flattened on white |
| `text` | both | – | OCR text from the label, ≤ 500 chars |
| `brand` | both | – | hard filter to one brand |
| `category` | both | – | `ILIKE` filter on the category column |
| `top_k` | both | 10 | 1–50 |
| `min_score` | both | 0.0 | −1…1; drop matches below it |
## Examples
```bash
# The app: vector + OCR text
curl -s -X POST https://mcp.nearle.ai.in/api/search/image-vector \
-H 'Content-Type: application/json' \
-d '{"vector":[0.0312,0.0682,-0.0223, ... 1024 values ...],
"text":"Britannia Marie Gold 300 g","top_k":5}'
# A photo
curl -s -X POST https://mcp.nearle.ai.in/api/search/image \
-F 'file=@marie_gold.jpg' -F 'text=Britannia Marie Gold 300 g' -F 'top_k=5'
```
Captured response (a simulated phone shot of Marie Gold, trimmed to two results):
```json
{
"results": [
{
"image_id": "britannia_marie_gold_300g",
"image_url": "https://www.britannia.co.in/_next/image?url=...",
"image_urls": ["https://www.britannia.co.in/_next/image?url=..."],
"brand": "Britannia",
"product_name": "Britannia Marie Gold 300g",
"title": "Britannia Marie Gold 300g",
"category": "General",
"description": "...",
"price_range": "₹45 - ₹55",
"size_variants": ["300g"],
"providers": [],
"highlights": [],
"nutrients": [],
"fssai_license": null,
"product_sku": "BRI-0342",
"sku_source": "internal",
"hsn_code": "1905",
"final_selling_price": 51.0,
"selling_price": 48.0,
"barcode": "8901063023949",
"barcode_type": "EAN13",
"nutrition_score": 6.2,
"health_score": 5.8,
"score": 0.631,
"text_overlap": 6.0
},
{ "product_name": "Britannia Marie Gold 117g", "score": 0.631, "text_overlap": 3.0, "...": "..." }
],
"total": 5,
"detected_brand": "Britannia",
"scoped_to_brand": true,
"scope_fallback": false,
"min_score": 0.0,
"top_k": 5,
"query_text": "Britannia Marie Gold 300 g"
}
```
Every result is the full product card (`ProductOut`) plus `score` and
`text_overlap`. `detected_brand` is the brand the search was scoped to,
whether it came from `brand` or from the text.
## Errors
| Code | Cause | What to do |
|---|---|---|
| 400 | `/image`: the upload is empty | send the file |
| 413 | `/image`: file over 8 MB | crop or downscale |
| 422 | wrong vector length, NaN, all zeros, `top_k` out of 1–50, `min_score` out of −1…1, undecodable image | `detail` names the field |
| 503 | `/image`: this deployment has no embedding model | embed client-side and use `/image-vector`; `GET /api/health` → `image_vectors.model_present` says whether this can happen |
## Good to know
- The brand switch `ACTIVE_BRANDS` (blank in production = all brands) limits
the *unscoped* search exactly as it limits `GET /api/search`; an explicit or
OCR-recognised brand is searched even if inactive.
- Rows with no usable photo have no vector and cannot be found this way
(~92% of image-bearing rows are covered; the rest are dead image hosts).
- The unscoped search runs one small index query per brand table (≈60) and
then reads full cards only for the winners; expect tens of milliseconds
with the database co-located, more over a WAN.
- Nothing here writes. The ingestion pipeline, upserts and the image-vector
worker are untouched.
Implementation: `app/services/image_match.py` (ranking), `vector_store.image_vector_search`
/ `fetch_products_by_image_ids` (reads), `app/api/routers/search.py` (routes),
`tests/test_image_match.py`, `tests/test_image_search_api.py`.