feat/env-login-scan-to-order #5
Reference in New Issue
Block a user
Delete Branch "feat/env-login-scan-to-order"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Scan-to-order: ask when a label fits several products
A bare brand label resolved to one arbitrary product and quoted a price for it.
This makes
/lookupreturn a "did you mean?" list instead.textScoreisuntouched — the fix reads the tie it already produces.
Found from the example request in the mobile team's own proposal:
{"label": "britannia"}.The bug
"britannia"is a substring of all 258 Britannia product names, sotextScorescored every one of them 0.95. Nothing downstream read that tie, so the sort
order picked a winner and the customer was shown one arbitrary biscuit with
"confidence": 0.95and a price on it.Lens returns a bare wordmark often — it is usually the biggest thing printed on
a packet — so this was the common path, not an edge case.
What changed
isAmbiguous: if the runner-up is withinscanAmbiguityMargin(0.06) of theleader, the reply becomes
ambiguous: truewithcandidatesinstead of amatch.
matchis nil andstoresempty on that path — no price for a productnobody chose.
Candidates are distinct products, not pack sizes (
distinctProducts, keyedon
variant_keyelse name within a brand), capped at 10, and each is markedwith whether one of that customer's registered stores has it in stock —
available ones first. That availability read reuses the same
StoreOptionsquery the confident path already runs, widened to every candidate, so the list
can show what is buyable before what is not.
To act on a pick,
/lookupnow acceptsbrand+catalogueidinstead of alabel and skips recognition entirely (
method: "direct",confidence: 1).Also serves deep links and "buy again". New:
ScanRepository.CatalogueRef,resolving the brand through the tables discovered from
information_schema,never a name built from the request.
What deliberately did NOT change
textScore, the 0.50 floor,minTokenHits, separator folding, and thecosine tie-break are all exactly as they are on main.
This branch originally replaced the substring rule with a coverage formula
(harmonic mean of label-coverage and name-coverage) to stop the arbitrary pick.
That was the wrong half to change, and the merge drops it:
it is the signal.
isAmbiguousreads it, so a scoring change adds nothing.SearchTokensdrops the singlecharacter, so coverage alone cannot see
Parle-G Original Glucose Biscuitsat all. The folding on main handles it properly.
Dropped with it:
scanSpecificEnough, the per-hit text score, and aproportional confirmation bonus — the flat
+0.10is back. Net result is lesscode than the branch started with, and none of main's tuning is touched.
API change (mobile)
/lookuphas two possible answers now and the app must handle both. Fullyspecified in
docs/SCAN_TO_ORDER.md(Response A / Response B).One thing worth flagging to the app developer:
confidencestays high on theambiguous path — 0.95, because
"britannia"genuinely does appear in allthose names. What is missing is identification, not relevance. Gate on
ambiguous, never onconfidence; an app that reads 0.95 as "safe to show aprice" reintroduces exactly this bug. The doc says so in those words.
Tests
25 scan tests pass together, both sides' included. Full suite,
go buildandgo vetclean.TestTextScoreRewardsSpecificityNotJustOverlaptested the removed formula andis replaced by
TestABrandNameScoresItsProductsIdentically, which guards thetie itself — a formula that broke it on name length or word count would bring
the bug back.
newBrandLabelFixtureguards the end-to-end behaviour.Notes for review
@abhishek — the merge commit (
24339a8) is the interesting one; it explains whyyour
textScoresurvived unchanged. Your two production findings are bothpreserved and now documented in
docs/SCAN_TO_ORDER.md: the 0.304"Paracetamol" → "Paneer Makhni 500ml" near-miss behind the 0.50 floor, and
"Parle G" losing to "Parle Monaco Classic" on an ASCII tie-break.
Two things found on the side, neither acted on in this PR:
img_vector vector(1024)— filled on 1885 of2124 rows, empty in
brand_haldirams,brand_kaleesuwari,brand_mdh,brand_zzsmoketest. This flow does not read it. The reasoning, and thespecific field measurement that would reverse the decision, is under "Two
decisions, and why" in the doc — worth reading before that thread is
reopened, since the mobile team proposed sending image vectors from the
phone.
scratch/cataloguedimsnow reports every vector column rather than onlyembedding, which is how (1) came to light.View command line instructions
Checkout
From your project repository, check out a new branch and test the changes.