main had moved on with retrieval work validated against real queries — minTokenHits (the word match needs two thirds of the label, not all of it), separator folding so "Parle G"/"Parle-G"/"ParleG" all reach Parle-G, the floor at 0.50 after "Paracetamol" came back as "Paneer Makhni 500ml" at 0.304, and ties broken on cosine distance instead of name. All of that is kept exactly as it was. The conflict was in textScore: this branch replaced the substring rule with a coverage formula to stop a bare brand name resolving to one arbitrary product. That is the wrong half to change. The substring rule scores every product of a brand 0.95 IDENTICALLY, and that tie is not the bug — it is the signal. isAmbiguous reads it, so the branch's coverage rewrite is dropped and the ambiguity layer alone does the work: "britannia" → all 258 rows tie at 0.95 → ambiguous: true + candidates "Parle G" → folding and the single-character token still land it a real name → runner-up far behind → match, unchanged Dropped with it: scanSpecificEnough, the per-hit text score, and the proportional confirmation bonus — the flat +0.10 is back. Simpler, and it leaves main's tuning untouched. TestTextScoreRewardsSpecificityNotJustOverlap tested the removed formula and is replaced by TestABrandNameScoresItsProductsIdentically, which guards the tie itself: a formula that broke it on name length or word count would bring the bug back. Docs carry both rationales, and now say plainly that confidence stays high on the ambiguous path — gate on `ambiguous`, never on `confidence`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
31 KiB
31 KiB