Merge origin/main: keep the substring rule, read its tie
main had moved on with retrieval work validated against real queries — minTokenHits (the word match needs two thirds of the label, not all of it), separator folding so "Parle G"/"Parle-G"/"ParleG" all reach Parle-G, the floor at 0.50 after "Paracetamol" came back as "Paneer Makhni 500ml" at 0.304, and ties broken on cosine distance instead of name. All of that is kept exactly as it was. The conflict was in textScore: this branch replaced the substring rule with a coverage formula to stop a bare brand name resolving to one arbitrary product. That is the wrong half to change. The substring rule scores every product of a brand 0.95 IDENTICALLY, and that tie is not the bug — it is the signal. isAmbiguous reads it, so the branch's coverage rewrite is dropped and the ambiguity layer alone does the work: "britannia" → all 258 rows tie at 0.95 → ambiguous: true + candidates "Parle G" → folding and the single-character token still land it a real name → runner-up far behind → match, unchanged Dropped with it: scanSpecificEnough, the per-hit text score, and the proportional confirmation bonus — the flat +0.10 is back. Simpler, and it leaves main's tuning untouched. TestTextScoreRewardsSpecificityNotJustOverlap tested the removed formula and is replaced by TestABrandNameScoresItsProductsIdentically, which guards the tie itself: a formula that broke it on name length or word count would bring the bug back. Docs carry both rationales, and now say plainly that confidence stays high on the ambiguous path — gate on `ambiguous`, never on `confidence`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -125,13 +125,13 @@ registered stores are used.
|
||||
"ambiguous": true,
|
||||
"candidates": [
|
||||
{ "brand": "britannia", "catalogueid": 23, "product_name": "Britannia Marie Gold",
|
||||
"size": "250 g", "image": "https://…", "score": 0.5, "method": "text", "available": true },
|
||||
"size": "250 g", "image": "https://…", "score": 0.95, "method": "text", "available": true },
|
||||
{ "brand": "britannia", "catalogueid": 22, "product_name": "Britannia Good Day Butter Cookies",
|
||||
"image": "https://…", "score": 0.333, "method": "text" },
|
||||
"image": "https://…", "score": 0.95, "method": "text" },
|
||||
{ "brand": "britannia", "catalogueid": 21, "product_name": "Britannia Good Day Cashew Cookies",
|
||||
"image": "https://…", "score": 0.333, "method": "text" }
|
||||
"image": "https://…", "score": 0.95, "method": "text" }
|
||||
],
|
||||
"confidence": 0.5,
|
||||
"confidence": 0.95,
|
||||
"available": false,
|
||||
"stores": [],
|
||||
"catalogue_variants": [],
|
||||
@@ -139,6 +139,11 @@ registered stores are used.
|
||||
}
|
||||
```
|
||||
|
||||
- **`confidence` is not low here, and that is not a bug.** "britannia" really
|
||||
does appear in all three names, so relevance is high — what is missing is
|
||||
*identification*. Gate on `ambiguous`, never on `confidence`: an app that
|
||||
reads 0.95 as "sure enough to show a price" reintroduces the exact bug this
|
||||
path exists to prevent.
|
||||
- **`available` on a candidate** means at least one of the customer's
|
||||
registered stores has it in stock right now. Candidates are ordered
|
||||
available-first, so the list can show what is buyable before what is not
|
||||
@@ -171,6 +176,10 @@ registered stores are used.
|
||||
cart/order calls exactly as you would from the catalogue screen.
|
||||
- `distance_km: -1` means the distance is unknown (no fix from the phone and
|
||||
no saved address, or the store has no coordinates). Do not render it as 0.
|
||||
Send `latitude`/`longitude` on `/confirm` too if you display distance from
|
||||
its reply: the saved address is only consulted there when the shelf is
|
||||
empty and alternatives have to be ranked, so without a fix the store you
|
||||
tapped comes back `-1`.
|
||||
|
||||
## `POST /confirm`
|
||||
|
||||
@@ -186,7 +195,7 @@ nothing is cached on this path.
|
||||
{
|
||||
"ok": false,
|
||||
"reason": "out_of_stock", // in_stock | insufficient_stock | out_of_stock | not_sold_here | store_not_registered
|
||||
"store": { "…the store they tapped…" },
|
||||
"store": { "…the store they tapped…" }, // distance_km filled from the fix you send
|
||||
"option": { "productid": 100, "stock": 0, "…": "…" },
|
||||
"requested": 2,
|
||||
"alternative": { // absent when nobody has enough
|
||||
@@ -225,6 +234,18 @@ Same `ScanStore` shape as inside `stores[]` above, without options.
|
||||
`EMBEDDING_PROVIDER/MODEL/API_KEY` and **must** be the one that indexed
|
||||
the catalogue — the first search checks the vector width and refuses a
|
||||
mismatch by name.
|
||||
- **The word match asks for most of the label, not all of it**
|
||||
(`minTokenHits`: two thirds, rounded up, and both of a two-word label).
|
||||
Requiring every word meant one word the catalogue does not use took the
|
||||
right product out of the running entirely — "Dettol bottle pack" retrieved
|
||||
no Dettol, "Parle G biscuit pack" retrieved no Parle-G — and the vector
|
||||
search then answered alone, confidently and wrongly, at a score the floor
|
||||
could not catch. Each brand's rows are ordered by how much of the label
|
||||
they carry (the whole label as a substring outranks any number of loose
|
||||
words) so that the per-brand `LIMIT` keeps the best rows and not merely the
|
||||
first ones the planner reached. Packaging words — "pack", "bottle", "jar",
|
||||
"sachet" and friends, see `utils.isPackaging` — are dropped before any of
|
||||
this, like pack sizes, unless the label is nothing else.
|
||||
- **The catalogue's model** (verified 2026-09-15 by cosine against a stored
|
||||
row: 1.0000): `all-MiniLM-L6-v2`, 384-d, unit-normalised, embedding the
|
||||
`search_query` column (brand + name + category + blurb + price range).
|
||||
@@ -238,10 +259,12 @@ Same `ScanStore` shape as inside `stores[]` above, without options.
|
||||
EMBEDDING_DIMENSIONS=384
|
||||
```
|
||||
A bare label ("Milk Bikis") scores ~0.92 against its product's stored
|
||||
vector and ~0.23 against an unrelated one, which is what the 0.30 floor in
|
||||
`scanService.go` is set against. If the catalogue team ever re-embeds
|
||||
with another model, change `EMBEDDING_MODEL`/`DIMENSIONS` here and
|
||||
nothing else.
|
||||
vector and ~0.23 against an unrelated one, which is what the 0.50 floor in
|
||||
`scanService.go` is set against — the middle of that split, not the edge of
|
||||
the noise. It was 0.30 until a near-miss got through in production
|
||||
("Paracetamol" → "Paneer Makhni 500ml", 0.304). If the catalogue team ever
|
||||
re-embeds with another model, change `EMBEDDING_MODEL`/`DIMENSIONS` here
|
||||
and nothing else.
|
||||
- **Speed**: the label's vector (7 days) and the ranked catalogue hits
|
||||
(30 min) are cached in Redis and in-process, so a popular product costs
|
||||
one model call platform-wide. Customer, stores and catalogue are read
|
||||
@@ -358,34 +381,44 @@ brand-name bug described under Scoring.
|
||||
|---|---|---|
|
||||
| `scanLookupTimeout` | 5 s | whole lookup, including the model call |
|
||||
| `scanCatalogueTopK` | 15 | rows taken from each brand table and from the merge |
|
||||
| `scanMinScore` | 0.30 | below this the best hit is not shown as a match |
|
||||
| `scanMinScore` | 0.50 | below this the best hit is not shown as a match |
|
||||
| `scanAmbiguityMargin` | 0.06 | how close the runner-up may be before the answer becomes a question |
|
||||
| `scanSpecificEnough` | 0.55 | text coverage the leader needs before it counts as identified |
|
||||
| `scanMaxCandidates` | 10 | longest "did you mean?" list |
|
||||
| `embedTimeout` (`utils/embedding.go`) | 4 s | one model call |
|
||||
| `scanVectorTTL` / `scanHitsTTL` (`scanRepository.go`) | 7 d / 30 min | cache lifetimes |
|
||||
|
||||
**Scoring.** Each hit carries two numbers, and they answer different
|
||||
questions:
|
||||
Scores: vector = `1 − cosine distance`; text = 0.95 for the whole label
|
||||
inside the name, else `0.8 × (label words found / label words)`; combined =
|
||||
`max(vector, text) + 0.10` when both hit, capped at 1. Ties are broken by
|
||||
cosine distance — nearest first, a text-only row last — and only then by
|
||||
name.
|
||||
|
||||
- `score` ranks. Vector = `1 − cosine distance`. Combined =
|
||||
`max(vector, text) + 0.10 × text`, capped at 1 — the confirmation bonus is
|
||||
proportional, so only a text match that actually names the product
|
||||
strengthens a vector hit.
|
||||
- `text` says how *specifically* the label names this product, and is the
|
||||
harmonic mean of two coverages: how much of the label the product accounts
|
||||
for, and how much of the product's name the label accounts for. Pack sizes
|
||||
are dropped from both sides.
|
||||
The label and the product name are both separator-folded before that
|
||||
substring test (`utils.FoldSeparators`), and compared again with separators
|
||||
removed (`utils.TightenLabel`, labels of 4+ characters), so the brand's own
|
||||
punctuation does not decide the match: "Parle G", "Parle-G" and "ParleG" all
|
||||
reach *Parle-G Original Glucose Biscuits*. A single-character token survives
|
||||
tokenising when it follows a word, because it is often the whole name — the
|
||||
"G" of Parle-G, the "K" of Special K. It is still dropped when it stands
|
||||
alone or is a pack multiplier.
|
||||
|
||||
Why both: `"britannia"` is a substring of all 258 Britannia product names.
|
||||
Judging on overlap alone scored every one of them 0.95, the tie broke
|
||||
alphabetically, and one arbitrary biscuit came back with a price. Now they
|
||||
score ~0.33 *equally*, which `isAmbiguous` reads as "ask, don't guess" — via
|
||||
the margin test (something is level with the leader) or the specificity test
|
||||
(the leader may rank first on vector similarity while the label names no one
|
||||
product). Erring towards asking is deliberate: asking costs one tap on a
|
||||
picture, guessing wrong costs the customer's belief that the scanner works.
|
||||
An exact product name still scores ~1.0, so the common case is untouched.
|
||||
All three mattered at once: before this, "Parle G" tied with *Parle Monaco
|
||||
Classic* at 0.9 (the "G" was dropped, so only "parle" matched either row),
|
||||
and the name tie-break handed it to Monaco because a space precedes a hyphen
|
||||
in ASCII. A confident, wrong answer — the kind no score floor can catch.
|
||||
|
||||
**When the substring rule ties, that tie is the answer.** A bare brand name
|
||||
is a substring of every one of that brand's names, so all of them score 0.95
|
||||
— identically, at a high score no floor would ever catch. Rather than
|
||||
scoring around it, `isAmbiguous` reads it: if the runner-up is within
|
||||
`scanAmbiguityMargin` of the leader, the reply becomes `ambiguous: true`
|
||||
with `candidates` instead of a match (see Response B). Erring towards asking
|
||||
is deliberate — one tap on a picture against the wrong biscuit. A label that
|
||||
names one product leaves the runner-up far behind, so the common case is
|
||||
untouched, and `services/scan_test.go`'s
|
||||
`TestABrandNameScoresItsProductsIdentically` guards the tie itself: a
|
||||
formula that broke it on name length or word count would bring the bug
|
||||
back.
|
||||
|
||||
### Changing the embedding model
|
||||
|
||||
|
||||
Reference in New Issue
Block a user