image search test

This commit is contained in:
2026-09-16 11:52:35 +05:30
parent 42ea007fe7
commit 28af3e05f2
5 changed files with 181 additions and 6 deletions

View File

@@ -244,7 +244,23 @@ you change ranking — it is the spec.
Scores: vector = `1 − cosine distance`; text = 0.95 for the whole label
inside the name, else `0.8 × (label words found / label words)`; combined =
`max(vector, text) + 0.10` when both hit, capped at 1.
`max(vector, text) + 0.10` when both hit, capped at 1. Ties are broken by
cosine distance — nearest first, a text-only row last — and only then by
name.
The label and the product name are both separator-folded before that
substring test (`utils.FoldSeparators`), and compared again with separators
removed (`utils.TightenLabel`, labels of 4+ characters), so the brand's own
punctuation does not decide the match: "Parle G", "Parle-G" and "ParleG" all
reach *Parle-G Original Glucose Biscuits*. A single-character token survives
tokenising when it follows a word, because it is often the whole name — the
"G" of Parle-G, the "K" of Special K. It is still dropped when it stands
alone or is a pack multiplier.
All three mattered at once: before this, "Parle G" tied with *Parle Monaco
Classic* at 0.9 (the "G" was dropped, so only "parle" matched either row),
and the name tie-break handed it to Monaco because a space precedes a hyphen
in ASCII. A confident, wrong answer — the kind no score floor can catch.
### Changing the embedding model