Merge origin/main: keep the substring rule, read its tie

main had moved on with retrieval work validated against real queries —
minTokenHits (the word match needs two thirds of the label, not all of
it), separator folding so "Parle G"/"Parle-G"/"ParleG" all reach Parle-G,
the floor at 0.50 after "Paracetamol" came back as "Paneer Makhni 500ml"
at 0.304, and ties broken on cosine distance instead of name. All of that
is kept exactly as it was.

The conflict was in textScore: this branch replaced the substring rule
with a coverage formula to stop a bare brand name resolving to one
arbitrary product. That is the wrong half to change. The substring rule
scores every product of a brand 0.95 IDENTICALLY, and that tie is not the
bug — it is the signal. isAmbiguous reads it, so the branch's coverage
rewrite is dropped and the ambiguity layer alone does the work:

  "britannia" → all 258 rows tie at 0.95 → ambiguous: true + candidates
  "Parle G"   → folding and the single-character token still land it
  a real name → runner-up far behind → match, unchanged

Dropped with it: scanSpecificEnough, the per-hit text score, and the
proportional confirmation bonus — the flat +0.10 is back. Simpler, and it
leaves main's tuning untouched.

TestTextScoreRewardsSpecificityNotJustOverlap tested the removed formula
and is replaced by TestABrandNameScoresItsProductsIdentically, which
guards the tie itself: a formula that broke it on name length or word
count would bring the bug back.

Docs carry both rationales, and now say plainly that confidence stays
high on the ambiguous path — gate on `ambiguous`, never on `confidence`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
2026-09-23 11:05:03 +05:30
25 changed files with 1661 additions and 228 deletions

View File

@@ -125,13 +125,13 @@ registered stores are used.
"ambiguous": true,
"candidates": [
{ "brand": "britannia", "catalogueid": 23, "product_name": "Britannia Marie Gold",
"size": "250 g", "image": "https://…", "score": 0.5, "method": "text", "available": true },
"size": "250 g", "image": "https://…", "score": 0.95, "method": "text", "available": true },
{ "brand": "britannia", "catalogueid": 22, "product_name": "Britannia Good Day Butter Cookies",
"image": "https://…", "score": 0.333, "method": "text" },
"image": "https://…", "score": 0.95, "method": "text" },
{ "brand": "britannia", "catalogueid": 21, "product_name": "Britannia Good Day Cashew Cookies",
"image": "https://…", "score": 0.333, "method": "text" }
"image": "https://…", "score": 0.95, "method": "text" }
],
"confidence": 0.5,
"confidence": 0.95,
"available": false,
"stores": [],
"catalogue_variants": [],
@@ -139,6 +139,11 @@ registered stores are used.
}
```
- **`confidence` is not low here, and that is not a bug.** "britannia" really
does appear in all three names, so relevance is high — what is missing is
*identification*. Gate on `ambiguous`, never on `confidence`: an app that
reads 0.95 as "sure enough to show a price" reintroduces the exact bug this
path exists to prevent.
- **`available` on a candidate** means at least one of the customer's
registered stores has it in stock right now. Candidates are ordered
available-first, so the list can show what is buyable before what is not
@@ -171,6 +176,10 @@ registered stores are used.
cart/order calls exactly as you would from the catalogue screen.
- `distance_km: -1` means the distance is unknown (no fix from the phone and
no saved address, or the store has no coordinates). Do not render it as 0.
Send `latitude`/`longitude` on `/confirm` too if you display distance from
its reply: the saved address is only consulted there when the shelf is
empty and alternatives have to be ranked, so without a fix the store you
tapped comes back `-1`.
## `POST /confirm`
@@ -186,7 +195,7 @@ nothing is cached on this path.
{
"ok": false,
"reason": "out_of_stock", // in_stock | insufficient_stock | out_of_stock | not_sold_here | store_not_registered
"store": { "…the store they tapped…" },
"store": { "…the store they tapped…" }, // distance_km filled from the fix you send
"option": { "productid": 100, "stock": 0, "…": "…" },
"requested": 2,
"alternative": { // absent when nobody has enough
@@ -225,6 +234,18 @@ Same `ScanStore` shape as inside `stores[]` above, without options.
`EMBEDDING_PROVIDER/MODEL/API_KEY` and **must** be the one that indexed
the catalogue — the first search checks the vector width and refuses a
mismatch by name.
- **The word match asks for most of the label, not all of it**
(`minTokenHits`: two thirds, rounded up, and both of a two-word label).
Requiring every word meant one word the catalogue does not use took the
right product out of the running entirely — "Dettol bottle pack" retrieved
no Dettol, "Parle G biscuit pack" retrieved no Parle-G — and the vector
search then answered alone, confidently and wrongly, at a score the floor
could not catch. Each brand's rows are ordered by how much of the label
they carry (the whole label as a substring outranks any number of loose
words) so that the per-brand `LIMIT` keeps the best rows and not merely the
first ones the planner reached. Packaging words — "pack", "bottle", "jar",
"sachet" and friends, see `utils.isPackaging` — are dropped before any of
this, like pack sizes, unless the label is nothing else.
- **The catalogue's model** (verified 2026-09-15 by cosine against a stored
row: 1.0000): `all-MiniLM-L6-v2`, 384-d, unit-normalised, embedding the
`search_query` column (brand + name + category + blurb + price range).
@@ -238,10 +259,12 @@ Same `ScanStore` shape as inside `stores[]` above, without options.
EMBEDDING_DIMENSIONS=384
```
A bare label ("Milk Bikis") scores ~0.92 against its product's stored
vector and ~0.23 against an unrelated one, which is what the 0.30 floor in
`scanService.go` is set against. If the catalogue team ever re-embeds
with another model, change `EMBEDDING_MODEL`/`DIMENSIONS` here and
nothing else.
vector and ~0.23 against an unrelated one, which is what the 0.50 floor in
`scanService.go` is set against — the middle of that split, not the edge of
the noise. It was 0.30 until a near-miss got through in production
("Paracetamol" → "Paneer Makhni 500ml", 0.304). If the catalogue team ever
re-embeds with another model, change `EMBEDDING_MODEL`/`DIMENSIONS` here
and nothing else.
- **Speed**: the label's vector (7 days) and the ranked catalogue hits
(30 min) are cached in Redis and in-process, so a popular product costs
one model call platform-wide. Customer, stores and catalogue are read
@@ -358,34 +381,44 @@ brand-name bug described under Scoring.
|---|---|---|
| `scanLookupTimeout` | 5 s | whole lookup, including the model call |
| `scanCatalogueTopK` | 15 | rows taken from each brand table and from the merge |
| `scanMinScore` | 0.30 | below this the best hit is not shown as a match |
| `scanMinScore` | 0.50 | below this the best hit is not shown as a match |
| `scanAmbiguityMargin` | 0.06 | how close the runner-up may be before the answer becomes a question |
| `scanSpecificEnough` | 0.55 | text coverage the leader needs before it counts as identified |
| `scanMaxCandidates` | 10 | longest "did you mean?" list |
| `embedTimeout` (`utils/embedding.go`) | 4 s | one model call |
| `scanVectorTTL` / `scanHitsTTL` (`scanRepository.go`) | 7 d / 30 min | cache lifetimes |
**Scoring.** Each hit carries two numbers, and they answer different
questions:
Scores: vector = `1 − cosine distance`; text = 0.95 for the whole label
inside the name, else `0.8 × (label words found / label words)`; combined =
`max(vector, text) + 0.10` when both hit, capped at 1. Ties are broken by
cosine distance — nearest first, a text-only row last — and only then by
name.
- `score` ranks. Vector = `1 − cosine distance`. Combined =
`max(vector, text) + 0.10 × text`, capped at 1 — the confirmation bonus is
proportional, so only a text match that actually names the product
strengthens a vector hit.
- `text` says how *specifically* the label names this product, and is the
harmonic mean of two coverages: how much of the label the product accounts
for, and how much of the product's name the label accounts for. Pack sizes
are dropped from both sides.
The label and the product name are both separator-folded before that
substring test (`utils.FoldSeparators`), and compared again with separators
removed (`utils.TightenLabel`, labels of 4+ characters), so the brand's own
punctuation does not decide the match: "Parle G", "Parle-G" and "ParleG" all
reach *Parle-G Original Glucose Biscuits*. A single-character token survives
tokenising when it follows a word, because it is often the whole name — the
"G" of Parle-G, the "K" of Special K. It is still dropped when it stands
alone or is a pack multiplier.
Why both: `"britannia"` is a substring of all 258 Britannia product names.
Judging on overlap alone scored every one of them 0.95, the tie broke
alphabetically, and one arbitrary biscuit came back with a price. Now they
score ~0.33 *equally*, which `isAmbiguous` reads as "ask, don't guess" — via
the margin test (something is level with the leader) or the specificity test
(the leader may rank first on vector similarity while the label names no one
product). Erring towards asking is deliberate: asking costs one tap on a
picture, guessing wrong costs the customer's belief that the scanner works.
An exact product name still scores ~1.0, so the common case is untouched.
All three mattered at once: before this, "Parle G" tied with *Parle Monaco
Classic* at 0.9 (the "G" was dropped, so only "parle" matched either row),
and the name tie-break handed it to Monaco because a space precedes a hyphen
in ASCII. A confident, wrong answer — the kind no score floor can catch.
**When the substring rule ties, that tie is the answer.** A bare brand name
is a substring of every one of that brand's names, so all of them score 0.95
— identically, at a high score no floor would ever catch. Rather than
scoring around it, `isAmbiguous` reads it: if the runner-up is within
`scanAmbiguityMargin` of the leader, the reply becomes `ambiguous: true`
with `candidates` instead of a match (see Response B). Erring towards asking
is deliberate — one tap on a picture against the wrong biscuit. A label that
names one product leaves the runner-up far behind, so the common case is
untouched, and `services/scan_test.go`'s
`TestABrandNameScoresItsProductsIdentically` guards the tie itself: a
formula that broke it on name length or word count would bring the bug
back.
### Changing the embedding model