Image vector to product details
This commit is contained in:
@@ -10,13 +10,16 @@ index). These two endpoints turn the app's vector - or a photo - into
|
||||
product cards.
|
||||
|
||||
```
|
||||
POST /api/search/image-vector JSON {vector[1024], text?, brand?, category?, top_k?, min_score?}
|
||||
POST /api/search/image-vector JSON {vector[1024], text?, brand?, category?, top_k?, min_score?, text_fallback?}
|
||||
GET /api/search/image-vector query string: vector=<base64 or csv>&text=...&top_k=... (see "GET variant")
|
||||
POST /api/search/image multipart file + the same optional fields as form fields
|
||||
POST /api/search/identify multipart file + the same fields; image first, label text second (see "Identify")
|
||||
```
|
||||
|
||||
Both run the same ranking. Use the first from the app (it already has the
|
||||
vector); use the second from anything without the model, or for testing.
|
||||
The first three run the same ranking. Use the first from the app (it already
|
||||
has the vector); use `/image` from anything without the model, or for
|
||||
testing. `/identify` is for a phone photo that the vector alone cannot
|
||||
confirm - see the section below.
|
||||
|
||||
## How a match is found
|
||||
|
||||
@@ -159,14 +162,97 @@ final b64 = base64Url.encode(vector.buffer.asUint8List()).replaceAll('=', '');
|
||||
A malformed `vector` (wrong count, not a number, bad base64, all zeros)
|
||||
is a 422 whose `detail` says which value or what length was wrong.
|
||||
|
||||
## Identify — image first, label text second
|
||||
|
||||
`POST /api/search/identify` exists because a **phone photo of a pack scores
|
||||
only ~0.63 against the catalogue's render of the same pack** (measured), so
|
||||
`img_vector` alone cannot confirm which product it is. This route runs a
|
||||
ladder and tells you which rung answered:
|
||||
|
||||
```
|
||||
photo ─► img_vector search ─► best score ≥ IMAGE_IDENTIFY_MIN_IMAGE_SCORE (0.70)?
|
||||
│ yes → matched_by "image_vector", fallback_reason null (confirmed)
|
||||
▼ no
|
||||
label text = your `text` (ocr_source "client")
|
||||
else the server reads it off the photo (ocr_source "server")
|
||||
▼
|
||||
resolve the label: brand from the text, MiniLM embedding vs the rows'
|
||||
`embedding` + product-name match, ranked by shared size/word tokens
|
||||
first, cosine second
|
||||
│ confident → matched_by "text" (confirmed)
|
||||
▼ not → the low image rows, if any, with fallback_reason "text_no_match"
|
||||
```
|
||||
|
||||
Same multipart fields as `/image`. `text` is optional: send it when your
|
||||
client already OCR'd the label (it is also used to scope and tie-break the
|
||||
image search); leave it out and the server reads the label itself.
|
||||
|
||||
```bash
|
||||
curl -s -X POST https://mcp.nearle.ai.in/api/search/identify \
|
||||
-F "file=@marie_gold_phone.jpg" -F top_k=5
|
||||
```
|
||||
|
||||
```jsonc
|
||||
{
|
||||
"matched_by": "text", // "image_vector" | "text" | "none"
|
||||
"fallback_reason": "image_below_threshold", // why the image rung was not the answer
|
||||
"ocr_source": "server", // "client" | "server" | null
|
||||
"ocr_text": "Britannia MARIE GOLD Original Tea Time Biscuit 300 g", // the cleaned label used
|
||||
"image_top_score": 0.6312, // the image side, always reported
|
||||
"detected_brand": "britannia", "scoped_to_brand": true, "scope_fallback": false,
|
||||
"results": [ { "product_name": "Britannia Marie Gold 300g", "score": 0.71, "text_overlap": 5.0, ... } ],
|
||||
"total": 5, "top_k": 5, "min_score": 0.0, "query_text": "Britannia MARIE GOLD ... 300 g"
|
||||
}
|
||||
```
|
||||
|
||||
**Reading the answer.** It is confirmed when `matched_by` is `"text"`, or
|
||||
`"image_vector"` with `fallback_reason` null. Anything else is a best effort:
|
||||
show it as such, or ask for another shot.
|
||||
|
||||
| `fallback_reason` | Meaning |
|
||||
|---|---|
|
||||
| `null` | the image match cleared the floor |
|
||||
| `image_below_threshold` | image ran, best score under the floor → the label decided (on a `"text"` answer) |
|
||||
| `no_image_match` | image ran and found nothing → the label decided |
|
||||
| `image_embedder_unavailable` | this deployment has no image model → text only |
|
||||
| `no_text` | `/image-vector` with `text_fallback` but no `text` |
|
||||
| `ocr_unavailable` | no `text` sent and server OCR is off / not installed (`GET /api/health` → `ocr`) |
|
||||
| `ocr_empty` | server OCR ran and read nothing usable |
|
||||
| `text_no_match` | the label was resolved but nothing was confident; the image rows are returned as they were |
|
||||
|
||||
**`score` is not one number.** On an `"image_vector"` answer it is the
|
||||
1024-d photo cosine as on `/image`. On a `"text"` answer it is
|
||||
`1 - (embedding <=> q)` in the 384-d MiniLM space - a different scale, not
|
||||
comparable to the first. `image_top_score` always carries the image side so
|
||||
you can see both. On text answers `text_overlap` (3 per shared pack-size
|
||||
token, 1 per shared word) is what ranked the rows, then - among rows that
|
||||
explain the label equally - the fewest name words the label never said, so
|
||||
"Dairy Milk 50g" outranks "Dairy Milk Fruit & Nut 50g" on a plain label and
|
||||
the reverse on a label that says "Fruit & Nut"; cosine only breaks what is
|
||||
left.
|
||||
|
||||
**Server OCR** is rapidocr (PP-OCR models on onnxruntime, CPU). It is loaded
|
||||
on the first photo that needs it (~2 s), then costs ~1–2 s per photo at the
|
||||
1280 px it downscales to; it runs only when the image rung failed *and* no
|
||||
`text` was sent. `GET /api/health` → `ocr` reports whether it is installed
|
||||
and loaded. `ENABLE_SERVER_OCR=false` turns it off; the route then relies on
|
||||
client `text`.
|
||||
|
||||
**The app's route.** `POST /api/search/image-vector` accepts
|
||||
`"text_fallback": true`. Off (the default) the response is exactly what it
|
||||
has always been. On, the same ladder runs with the app's vector and `text`
|
||||
(no photo, so never server OCR), and the response is the identify shape
|
||||
above. Not on the GET variant.
|
||||
|
||||
## Errors
|
||||
|
||||
| Code | Cause | What to do |
|
||||
|---|---|---|
|
||||
| 400 | `/image`: the upload is empty | send the file |
|
||||
| 413 | `/image`: file over 8 MB | crop or downscale |
|
||||
| 400 | `/image`, `/identify`: the upload is empty | send the file |
|
||||
| 413 | `/image`, `/identify`: file over 8 MB | crop or downscale |
|
||||
| 422 | wrong vector length, NaN, all zeros, unparseable GET `vector`, `top_k` out of 1–50, `min_score` out of −1…1, undecodable image | `detail` names the field |
|
||||
| 503 | `/image`: this deployment has no embedding model | embed client-side and use `/image-vector`; `GET /api/health` → `image_vectors.model_present` says whether this can happen |
|
||||
| 503 | `/identify`: no embedding model AND no server OCR AND no `text` | send `text`, or a vector to `/image-vector`; `GET /api/health` → `image_vectors`, `ocr` |
|
||||
|
||||
## Good to know
|
||||
|
||||
@@ -183,4 +269,8 @@ is a 422 whose `detail` says which value or what length was wrong.
|
||||
|
||||
Implementation: `app/services/image_match.py` (ranking), `vector_store.image_vector_search`
|
||||
/ `fetch_products_by_image_ids` (reads), `app/api/routers/search.py` (routes),
|
||||
`tests/test_image_match.py`, `tests/test_image_search_api.py`.
|
||||
`tests/test_image_match.py`, `tests/test_image_search_api.py`. Identify:
|
||||
`app/services/product_identify.py` (the ladder), `app/services/label_match.py`
|
||||
(label → rows), `app/services/ocr_service.py` (server OCR; `requirements-ocr.txt`),
|
||||
`tests/test_product_identify.py`, `tests/test_label_match.py`,
|
||||
`tests/test_ocr_service.py`, `tests/test_identify_api.py`.
|
||||
|
||||
Reference in New Issue
Block a user