Ask instead of guessing when a label fits several products
`"britannia"` is a substring of all 258 Britannia product names, and textScore returned 0.95 for any product whose name contained the label. So every one of them tied, the tie broke alphabetically, and the customer was shown one arbitrary biscuit with "confidence": 0.95 and a price. Lens hands back a bare wordmark often — it is usually the biggest thing printed on a packet — so this was the common case, not an edge one. Found via the example request in the mobile team's own proposal. Scoring now asks both questions. A hit carries `score` (ranks) and `text` (how specifically the label names THIS product: the harmonic mean of how much of the label the product explains and how much of the product's name the label explains, pack sizes dropped from both sides). A brand name scores its products ~0.33 equally instead of 0.95 arbitrarily. The "vector and text agree" bonus is now proportional to the text score, so a weak match can no longer inflate a whole brand. isAmbiguous reads that: the leader is a guess if anything is level with it (margin) or if the label names no one product (specificity), and then the response carries `ambiguous: true` with `candidates` — distinct products, not pack sizes, at most ten, each marked with whether one of the customer's stores has it in stock, available ones first. `match` is nil and `stores` empty on that path: no price for a product nobody chose. Erring towards asking is deliberate — a tap versus the wrong biscuit. To act on a pick, /lookup now accepts `brand` + `catalogueid` instead of a label and skips recognition entirely (also serves deep links and re-order). New: ScanRepository.CatalogueRef, resolving via the brand tables discovered from information_schema, never a name built from the request. Also: scratch/cataloguedims now reports every vector column, not just `embedding` — which is how we learned the catalogue also carries img_vector(1024), filled on 1885 of 2124 rows. SCAN_TO_ORDER.md records why that column stays unread for now and what would change it, alongside why the app is not asked to compute vectors on the phone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -8,6 +8,10 @@ recommend. When the customer taps a store and a size, a second call confirms
|
||||
the shelf still has it — and if it does not, names the next-nearest store
|
||||
that does.
|
||||
|
||||
When the label fits several products — `"britannia"` names 258 of them — it
|
||||
answers with a short "did you mean?" list instead of picking one, because a
|
||||
confident price on the wrong biscuit is worse than one extra tap.
|
||||
|
||||
Base path: `/live/api/v1/mob/scan`. Every response uses the usual envelope
|
||||
`{ code, status, message, details }`; the shapes below are `details`.
|
||||
|
||||
@@ -17,7 +21,13 @@ Base path: `/live/api/v1/mob/scan`. Every response uses the usual envelope
|
||||
photo ──Lens──▶ label
|
||||
│
|
||||
▼
|
||||
POST /lookup ───▶ match + stores[] (recommended first)
|
||||
POST /lookup ───▶ ambiguous:true + candidates[] "did you mean?"
|
||||
│ │
|
||||
│ customer taps one candidate
|
||||
│ │
|
||||
│ POST /lookup { brand, catalogueid }
|
||||
│ │
|
||||
└───▶ match + stores[] (recommended first) ◀──┘
|
||||
│
|
||||
customer taps a store + a size
|
||||
│
|
||||
@@ -26,11 +36,21 @@ photo ──Lens──▶ label
|
||||
ok:false + alternative → offer the other store
|
||||
```
|
||||
|
||||
**`/lookup` has two possible answers and the app must handle both.** A label
|
||||
that names one product comes back with `match` + `stores`. A label that fits
|
||||
several — a bare brand name like `"britannia"`, a generic word like
|
||||
`"biscuits"` — comes back with `ambiguous: true` and `candidates`, and the
|
||||
app asks the customer which one before any price is shown. Lens returns a
|
||||
bare wordmark often, because it is usually the biggest thing printed on a
|
||||
packet, so this is a normal path and not an error case.
|
||||
|
||||
`GET /stores` is for the "choose another shop" sheet: the customer's
|
||||
registered stores, nearest first, independent of any product.
|
||||
|
||||
## `POST /lookup`
|
||||
|
||||
Note the `//` notes below are annotations, not JSON — strip them.
|
||||
|
||||
```json
|
||||
{
|
||||
"customerid": 5123,
|
||||
@@ -39,16 +59,25 @@ registered stores, nearest first, independent of any product.
|
||||
"longitude": 77.0290,
|
||||
"tenantids": [1135, 1140], // optional: what the app THINKS the customer joined
|
||||
"limit": 0 // optional: max stores, 0 = all
|
||||
|
||||
// Instead of a label: name the product outright. This is how you resolve
|
||||
// a candidate the customer tapped, and how a deep link or a "buy again"
|
||||
// skips recognition. With both set, `label` is ignored.
|
||||
// "brand": "britannia", "catalogueid": 7
|
||||
}
|
||||
```
|
||||
|
||||
`label` is required **unless** `brand` and `catalogueid` are both given.
|
||||
|
||||
`tenantids` is verified, never trusted: the server intersects it with the
|
||||
`tenantcustomers` table. Ids the customer is not actually registered with
|
||||
come back in `unregistered_tenantids` — treat that as "refresh the local
|
||||
list". A list that matches nothing at all is treated as stale and all
|
||||
registered stores are used.
|
||||
|
||||
Response:
|
||||
### Response A — one product identified
|
||||
|
||||
`ambiguous: false`, `match` set, `candidates` empty.
|
||||
|
||||
```json
|
||||
{
|
||||
@@ -59,6 +88,8 @@ Response:
|
||||
"image": "https://…", "score": 0.94, "method": "vector+text"
|
||||
},
|
||||
"catalogue_variants": [ { "…same shape…": "100 g" }, { "…": "200 g" } ],
|
||||
"ambiguous": false,
|
||||
"candidates": [],
|
||||
"confidence": 0.94,
|
||||
"available": true,
|
||||
"recommended_locationid": 20,
|
||||
@@ -83,12 +114,52 @@ Response:
|
||||
}
|
||||
```
|
||||
|
||||
How to read it:
|
||||
### Response B — several products fit, none clearly
|
||||
|
||||
- `match == null` → nothing recognised; show `message` and let them retry.
|
||||
`confidence` below ~0.5 → recognised but unsure; confirm the name with the
|
||||
customer before showing prices. `method: "text"` means no embedding model
|
||||
was involved (not configured, or it timed out) — be a little more cautious.
|
||||
`ambiguous: true`, `match: null`, `stores: []`. Show a "did you mean?" list.
|
||||
|
||||
```json
|
||||
{
|
||||
"label": "britannia",
|
||||
"match": null,
|
||||
"ambiguous": true,
|
||||
"candidates": [
|
||||
{ "brand": "britannia", "catalogueid": 23, "product_name": "Britannia Marie Gold",
|
||||
"size": "250 g", "image": "https://…", "score": 0.5, "method": "text", "available": true },
|
||||
{ "brand": "britannia", "catalogueid": 22, "product_name": "Britannia Good Day Butter Cookies",
|
||||
"image": "https://…", "score": 0.333, "method": "text" },
|
||||
{ "brand": "britannia", "catalogueid": 21, "product_name": "Britannia Good Day Cashew Cookies",
|
||||
"image": "https://…", "score": 0.333, "method": "text" }
|
||||
],
|
||||
"confidence": 0.5,
|
||||
"available": false,
|
||||
"stores": [],
|
||||
"catalogue_variants": [],
|
||||
"message": "Which one is it? 1 of these 3 are in stock near you."
|
||||
}
|
||||
```
|
||||
|
||||
- **`available` on a candidate** means at least one of the customer's
|
||||
registered stores has it in stock right now. Candidates are ordered
|
||||
available-first, so the list can show what is buyable before what is not
|
||||
— and the field is absent (not `false`) when unavailable, so read it as
|
||||
falsy, not as a required key.
|
||||
- **To resolve a pick**, call `/lookup` again with that candidate's `brand`
|
||||
and `catalogueid` and no label. You get Response A for that exact product,
|
||||
with `method: "direct"` and `confidence: 1`.
|
||||
- At most 10 candidates come back.
|
||||
|
||||
### How to read either response
|
||||
|
||||
- `match == null && !ambiguous` → nothing recognised; show `message` and let
|
||||
them retry with a clearer photo.
|
||||
- `ambiguous: true` → ask, do not guess. Never show a price on this path;
|
||||
`stores` is deliberately empty.
|
||||
- `confidence` below ~0.5 with a `match` → recognised but unsure; worth
|
||||
confirming the name before showing prices. `method: "text"` means no
|
||||
embedding model was involved (not configured, or it timed out) — be a
|
||||
little more cautious. `method: "direct"` means the caller named the
|
||||
product, so nothing was recognised at all.
|
||||
- `stores` is ordered **in-stock first, then nearest**. Exactly one store has
|
||||
`recommended: true` — the nearest with stock — and only when `available`
|
||||
is true. Stores that sell it but have nothing on the shelf are still listed
|
||||
@@ -187,6 +258,59 @@ Same `ScanStore` shape as inside `stores[]` above, without options.
|
||||
- **Identity** is the `customerid` in the body, like every other mobile
|
||||
endpoint here — there is no auth layer yet (see `SECURITY_HANDOFF.md`).
|
||||
|
||||
## Two decisions, and why
|
||||
|
||||
Both come from a proposal (2026-09-23) to have the app send vectors it
|
||||
computed on the phone. Recorded here because the next person will ask.
|
||||
|
||||
### The app does not send `textvector`
|
||||
|
||||
An on-device MiniLM vector is only comparable to the catalogue's if the app
|
||||
ships the identical model *and* tokenizer *and* pooling *and* normalisation;
|
||||
a quantised tflite build usually drifts, and the failure is silent — the
|
||||
ranking just gets worse. There is also nothing to gain: the server-side
|
||||
embed is ~30 ms warm and the result is cached in Redis by label, so one
|
||||
model call serves every customer who scans that product. A client-supplied
|
||||
vector *defeats* that cache (the key would have to be the vector, not the
|
||||
label), and 384 floats is ~5 KB of upload against ~12 bytes for
|
||||
`"Milk Bikis"`. If the field ever arrives it can be accepted and validated,
|
||||
but the app should not be asked to compute it.
|
||||
|
||||
**Send the full OCR text instead** if you want to give the server more to
|
||||
work with — ~100 bytes, no model coupling, strictly more information than a
|
||||
single label.
|
||||
|
||||
### The app does not send `imagevector` — yet
|
||||
|
||||
The catalogue *does* carry image vectors: every `brand_*` table has
|
||||
`img_vector vector(1024)`, filled on 1885 of 2124 rows (empty in
|
||||
`brand_haldirams`, `brand_kaleesuwari`, `brand_mdh`, `brand_zzsmoketest`).
|
||||
That matches the proposed MobileNetV3-Small embedder, so the idea is
|
||||
coherent and half-built — this flow simply does not read that column.
|
||||
|
||||
It stays unread for now because **Google Lens is already the image
|
||||
recogniser, and a far better one**: photo → Lens → label is Google's product
|
||||
recognition, trained on billions of images. Putting a 137M-parameter
|
||||
ImageNet backbone searching 1885 vectors *behind* that adds little where
|
||||
Lens succeeds, and MobileNetV3-Small — which struggles to tell one blue
|
||||
biscuit wrapper from another — is unlikely to rescue the cases where Lens
|
||||
fails. There is also an unverified dependency: the preprocessing the app
|
||||
would use (BGR → centre crop → 224×224 INTER_AREA → RGB → `/255.0`) has to
|
||||
match whatever the catalogue pipeline actually ran, or the search returns
|
||||
confidently-ranked noise.
|
||||
|
||||
**What would change this:** the field data. Once live, count how often
|
||||
`/lookup` returns `ambiguous: true` or nothing recognised. If Lens labels are
|
||||
reliable, image search is polish; if that number is high, it becomes the
|
||||
priority — and the first task is the cosine check (embed a known catalogue
|
||||
product's image through the app's exact pipeline, compare with its stored
|
||||
`img_vector`; ≈0.99 means the contract holds), not writing the query.
|
||||
|
||||
There is one non-recognition argument for it worth remembering: on-device
|
||||
inference is free and needs no Google dependency, which matters if Cloud
|
||||
Vision costs start to bite at volume. That is a business reason, not a
|
||||
quality one.
|
||||
|
||||
## For backend developers
|
||||
|
||||
### Where the code is
|
||||
@@ -201,7 +325,7 @@ Same `ScanStore` shape as inside `stores[]` above, without options.
|
||||
| `utils/embedding.go` | `Embedder` interface, OpenAI-compatible and Gemini clients |
|
||||
| `utils/geo.go` | coordinate parsing, haversine, opening hours, label tokenising |
|
||||
| `config/config.go` | `EmbeddingConfig` and its validation |
|
||||
| `scratch/cataloguedims` | read-only check of the catalogue's embedding width / fill |
|
||||
| `scratch/cataloguedims` | read-only check of every catalogue vector column's width and fill |
|
||||
|
||||
### Try it locally
|
||||
|
||||
@@ -220,11 +344,13 @@ anything; a schema-only dump does not.
|
||||
|
||||
### Tests
|
||||
|
||||
`go test ./services -run 'Lookup|Confirm|Stores|CatalogueFamily'` drives
|
||||
the whole pipeline through a fake repository (`services/scan_test.go`); no
|
||||
database. `go test ./utils` covers both HTTP clients against `httptest`
|
||||
servers, and the geo helpers. Add a case to `scan_test.go`'s fixture when
|
||||
you change ranking — it is the spec.
|
||||
`go test ./services -run 'Lookup|Confirm|Stores|Brand|Ambiguous|Specific|TextScore|Distinct|Naming'`
|
||||
drives the whole pipeline through a fake repository
|
||||
(`services/scan_test.go`); no database. `go test ./utils` covers both HTTP
|
||||
clients against `httptest` servers, and the geo helpers. Add a case to
|
||||
`scan_test.go`'s fixture when you change ranking — it is the spec, and
|
||||
`newBrandLabelFixture` in particular is the regression guard for the
|
||||
brand-name bug described under Scoring.
|
||||
|
||||
### Knobs (constants in `scanService.go`)
|
||||
|
||||
@@ -233,12 +359,33 @@ you change ranking — it is the spec.
|
||||
| `scanLookupTimeout` | 5 s | whole lookup, including the model call |
|
||||
| `scanCatalogueTopK` | 15 | rows taken from each brand table and from the merge |
|
||||
| `scanMinScore` | 0.30 | below this the best hit is not shown as a match |
|
||||
| `scanAmbiguityMargin` | 0.06 | how close the runner-up may be before the answer becomes a question |
|
||||
| `scanSpecificEnough` | 0.55 | text coverage the leader needs before it counts as identified |
|
||||
| `scanMaxCandidates` | 10 | longest "did you mean?" list |
|
||||
| `embedTimeout` (`utils/embedding.go`) | 4 s | one model call |
|
||||
| `scanVectorTTL` / `scanHitsTTL` (`scanRepository.go`) | 7 d / 30 min | cache lifetimes |
|
||||
|
||||
Scores: vector = `1 − cosine distance`; text = 0.95 for the whole label
|
||||
inside the name, else `0.8 × (label words found / label words)`; combined =
|
||||
`max(vector, text) + 0.10` when both hit, capped at 1.
|
||||
**Scoring.** Each hit carries two numbers, and they answer different
|
||||
questions:
|
||||
|
||||
- `score` ranks. Vector = `1 − cosine distance`. Combined =
|
||||
`max(vector, text) + 0.10 × text`, capped at 1 — the confirmation bonus is
|
||||
proportional, so only a text match that actually names the product
|
||||
strengthens a vector hit.
|
||||
- `text` says how *specifically* the label names this product, and is the
|
||||
harmonic mean of two coverages: how much of the label the product accounts
|
||||
for, and how much of the product's name the label accounts for. Pack sizes
|
||||
are dropped from both sides.
|
||||
|
||||
Why both: `"britannia"` is a substring of all 258 Britannia product names.
|
||||
Judging on overlap alone scored every one of them 0.95, the tie broke
|
||||
alphabetically, and one arbitrary biscuit came back with a price. Now they
|
||||
score ~0.33 *equally*, which `isAmbiguous` reads as "ask, don't guess" — via
|
||||
the margin test (something is level with the leader) or the specificity test
|
||||
(the leader may rank first on vector similarity while the label names no one
|
||||
product). Erring towards asking is deliberate: asking costs one tap on a
|
||||
picture, guessing wrong costs the customer's belief that the scanner works.
|
||||
An exact product name still scores ~1.0, so the common case is untouched.
|
||||
|
||||
### Changing the embedding model
|
||||
|
||||
|
||||
Reference in New Issue
Block a user