text search

This commit is contained in:
2026-09-16 12:17:39 +05:30
parent 28af3e05f2
commit c06b029cb2
5 changed files with 133 additions and 11 deletions

View File

@@ -158,6 +158,18 @@ Same `ScanStore` shape as inside `stores[]` above, without options.
`EMBEDDING_PROVIDER/MODEL/API_KEY` and **must** be the one that indexed
the catalogue — the first search checks the vector width and refuses a
mismatch by name.
- **The word match asks for most of the label, not all of it**
(`minTokenHits`: two thirds, rounded up, and both of a two-word label).
Requiring every word meant one word the catalogue does not use took the
right product out of the running entirely — "Dettol bottle pack" retrieved
no Dettol, "Parle G biscuit pack" retrieved no Parle-G — and the vector
search then answered alone, confidently and wrongly, at a score the floor
could not catch. Each brand's rows are ordered by how much of the label
they carry (the whole label as a substring outranks any number of loose
words) so that the per-brand `LIMIT` keeps the best rows and not merely the
first ones the planner reached. Packaging words — "pack", "bottle", "jar",
"sachet" and friends, see `utils.isPackaging` — are dropped before any of
this, like pack sizes, unless the label is nothing else.
- **The catalogue's model** (verified 2026-09-15 by cosine against a stored
row: 1.0000): `all-MiniLM-L6-v2`, 384-d, unit-normalised, embedding the
`search_query` column (brand + name + category + blurb + price range).