main had moved on with retrieval work validated against real queries —
minTokenHits (the word match needs two thirds of the label, not all of
it), separator folding so "Parle G"/"Parle-G"/"ParleG" all reach Parle-G,
the floor at 0.50 after "Paracetamol" came back as "Paneer Makhni 500ml"
at 0.304, and ties broken on cosine distance instead of name. All of that
is kept exactly as it was.
The conflict was in textScore: this branch replaced the substring rule
with a coverage formula to stop a bare brand name resolving to one
arbitrary product. That is the wrong half to change. The substring rule
scores every product of a brand 0.95 IDENTICALLY, and that tie is not the
bug — it is the signal. isAmbiguous reads it, so the branch's coverage
rewrite is dropped and the ambiguity layer alone does the work:
"britannia" → all 258 rows tie at 0.95 → ambiguous: true + candidates
"Parle G" → folding and the single-character token still land it
a real name → runner-up far behind → match, unchanged
Dropped with it: scanSpecificEnough, the per-hit text score, and the
proportional confirmation bonus — the flat +0.10 is back. Simpler, and it
leaves main's tuning untouched.
TestTextScoreRewardsSpecificityNotJustOverlap tested the removed formula
and is replaced by TestABrandNameScoresItsProductsIdentically, which
guards the tie itself: a formula that broke it on name length or word
count would bring the bug back.
Docs carry both rationales, and now say plainly that confidence stays
high on the ambiguous path — gate on `ambiguous`, never on `confidence`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`"britannia"` is a substring of all 258 Britannia product names, and
textScore returned 0.95 for any product whose name contained the label.
So every one of them tied, the tie broke alphabetically, and the customer
was shown one arbitrary biscuit with "confidence": 0.95 and a price. Lens
hands back a bare wordmark often — it is usually the biggest thing printed
on a packet — so this was the common case, not an edge one. Found via the
example request in the mobile team's own proposal.
Scoring now asks both questions. A hit carries `score` (ranks) and `text`
(how specifically the label names THIS product: the harmonic mean of how
much of the label the product explains and how much of the product's name
the label explains, pack sizes dropped from both sides). A brand name
scores its products ~0.33 equally instead of 0.95 arbitrarily. The
"vector and text agree" bonus is now proportional to the text score, so a
weak match can no longer inflate a whole brand.
isAmbiguous reads that: the leader is a guess if anything is level with it
(margin) or if the label names no one product (specificity), and then the
response carries `ambiguous: true` with `candidates` — distinct products,
not pack sizes, at most ten, each marked with whether one of the
customer's stores has it in stock, available ones first. `match` is nil
and `stores` empty on that path: no price for a product nobody chose.
Erring towards asking is deliberate — a tap versus the wrong biscuit.
To act on a pick, /lookup now accepts `brand` + `catalogueid` instead of a
label and skips recognition entirely (also serves deep links and re-order).
New: ScanRepository.CatalogueRef, resolving via the brand tables discovered
from information_schema, never a name built from the request.
Also: scratch/cataloguedims now reports every vector column, not just
`embedding` — which is how we learned the catalogue also carries
img_vector(1024), filled on 1885 of 2124 rows. SCAN_TO_ORDER.md records
why that column stays unread for now and what would change it, alongside
why the app is not asked to compute vectors on the phone.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
POST /v1/mob/scan/lookup label + customer → catalogue match, sizes, and
every registered store that sells it with live
stock, in-stock first / nearest first, one
recommended
POST /v1/mob/scan/confirm chosen store + size + qty → re-read the ledger;
ok, or the next-nearest store with enough of the
same product
GET /v1/mob/scan/stores registered stores nearest first
Recognition is pgvector cosine search over every brand_* table (each
with its own index, merged) plus a word match that settles near-ties
and works alone when no model is configured. The embedder is chosen by
EMBEDDING_PROVIDER (OpenAI-compatible or Gemini) and must be the model
that indexed the catalogue: verified 2026-09-15 as all-MiniLM-L6-v2 over
search_query, served by the cluster's Ollama as `all-minilm`; the first
search refuses a width mismatch by name.
Customer, stores and catalogue are read concurrently under a 5 s cap; a
slow model degrades to a text answer. Vectors and ranked hits are cached
in Redis and in-process; live stock never is. Availability uses the same
rules as the customer catalogue (approve, publishedat, ledger balance,
outlet price else retail). No stock reservation: confirm re-reads.
scratch/cataloguedims reports the catalogue's embedding width and fill.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>