Updates on Image search using vectors

This commit is contained in:
sriram
2026-09-28 15:44:01 +05:30
parent 6628207810
commit c0489d89d6
16 changed files with 1184 additions and 37 deletions

View File

@@ -9,6 +9,24 @@ Every catalogue row carries the same kind of vector in `img_vector`
index). These two endpoints turn the app's vector - or a photo - into
product cards.
## Recommended client flow (read this first)
A colleague photographed the **Cadbury Dairy Milk Lickables** card and got
"Milk Toned"; **Aachi Sambar Powder 100g** came back as Sakthi's "Sambar
powder 50g". A photo of a card on a screen scores only ~0.5 against its own
product, and a stranger can sit a few hundredths behind it. So:
1. **Send the photo to `POST /api/search/identify`** (preferred). The server
embeds it with the same model as the catalogue, reads the card/label text
with OCR, and lets the text decide when the image cannot. No model-parity
risk. *Or*, keeping the on-device vector, **always send the full OCR read
as `text`** to `/image-vector`: the label then decides (see `text_fallback`).
2. **Crop to the pack** (a framing guide in the camera UI). Measured: a card
photographed whole scored 0.51 for the right product; cropped to the pack
image, 0.94.
3. **Honour `match_confidence`.** Only `"confirmed"` is an answer. On `"low"`
show the results as a list to pick from - never auto-select the first.
```
POST /api/search/image-vector JSON {vector[1024], text?, brand?, category?, top_k?, min_score?, text_fallback?}
GET /api/search/image-vector query string: vector=<base64 or csv>&text=...&top_k=... (see "GET variant")
@@ -55,6 +73,14 @@ a floor, and read `score` on each result.
| `category` | both | – | `ILIKE` filter on the category column |
| `top_k` | both | 10 | 1–50 |
| `min_score` | both | 0.0 | −1…1; drop matches below it |
| `text_fallback` | `/image-vector` POST only | on when `text` is sent | run the identify ladder (below); `false` = image-only ranking, `text` breaks ties only |
Every response also carries:
| Field | Meaning |
|---|---|
| `match_confidence` | `"confirmed"`: the best match scored ≥ `IMAGE_IDENTIFY_MIN_IMAGE_SCORE` (0.70) **and** leads the best *different* photo by ≥ `IMAGE_SEARCH_MIN_MARGIN` (0.05) - or, on the identify shape, the ladder confirmed it. `"low"`: a list to pick from, not an answer. `"none"`: no results. |
| `margin` | best image score minus the best different photo's. Pack sizes sharing one photo are not rivals. `null` when there was no rival, or when the rows came from the label text. |
## Examples
@@ -170,7 +196,8 @@ only ~0.63 against the catalogue's render of the same pack** (measured), so
ladder and tells you which rung answered:
```
photo ─► img_vector search ─► best score ≥ IMAGE_IDENTIFY_MIN_IMAGE_SCORE (0.70)?
photo ─► img_vector search ─► best score ≥ IMAGE_IDENTIFY_MIN_IMAGE_SCORE (0.70)
│ and ≥ IMAGE_SEARCH_MIN_MARGIN (0.05) ahead of the next different photo?
│ yes → matched_by "image_vector", fallback_reason null (confirmed)
▼ no
label text = your `text` (ocr_source "client")
@@ -213,6 +240,7 @@ show it as such, or ask for another shot.
|---|---|
| `null` | the image match cleared the floor |
| `image_below_threshold` | image ran, best score under the floor → the label decided (on a `"text"` answer) |
| `image_ambiguous` | best score cleared the floor, but a different photo was within `IMAGE_SEARCH_MIN_MARGIN` |
| `no_image_match` | image ran and found nothing → the label decided |
| `image_embedder_unavailable` | this deployment has no image model → text only |
| `no_text` | `/image-vector` with `text_fallback` but no `text` |
@@ -238,11 +266,22 @@ on the first photo that needs it (~2 s), then costs ~1–2 s per photo at the
and loaded. `ENABLE_SERVER_OCR=false` turns it off; the route then relies on
client `text`.
**The app's route.** `POST /api/search/image-vector` accepts
`"text_fallback": true`. Off (the default) the response is exactly what it
has always been. On, the same ladder runs with the app's vector and `text`
(no photo, so never server OCR), and the response is the identify shape
above. Not on the GET variant.
**The app's route.** `POST /api/search/image-vector` runs the same ladder
with the app's vector and `text` (no photo, so never server OCR) whenever
`text` is sent, and answers with the identify shape above. Before
2026-09-28 it did so only with `"text_fallback": true`, and otherwise the
label only broke exact ties - a label naming the product lost to any photo
that scored 0.01 higher. Send `"text_fallback": false` to keep that old
image-only ranking; without `text` nothing changes. Not on the GET variant.
**Diagnosing a wrong match.** Every search-by-photo request logs one
`[IMAGE_SEARCH]` line (route, text, brand, a fingerprint of the vector, top 3
with scores, margin). With `IMAGE_SEARCH_CAPTURE_DIR` set the full request -
and the photo, on `/image` and `/identify` - is saved for
`python -m scripts.replay_image_query <file> [--photo shot.jpg]`, which
replays it and, given the photo, checks that the app's on-device vector
matches the server's (cosine ≥ 0.99). `python -m scripts.eval_identify`
measures accuracy and the thresholds on synthetic and real card photos.
## When the product is not in the catalogue (capture-to-catalog)