main had moved on with retrieval work validated against real queries — minTokenHits (the word match needs two thirds of the label, not all of it), separator folding so "Parle G"/"Parle-G"/"ParleG" all reach Parle-G, the floor at 0.50 after "Paracetamol" came back as "Paneer Makhni 500ml" at 0.304, and ties broken on cosine distance instead of name. All of that is kept exactly as it was. The conflict was in textScore: this branch replaced the substring rule with a coverage formula to stop a bare brand name resolving to one arbitrary product. That is the wrong half to change. The substring rule scores every product of a brand 0.95 IDENTICALLY, and that tie is not the bug — it is the signal. isAmbiguous reads it, so the branch's coverage rewrite is dropped and the ambiguity layer alone does the work: "britannia" → all 258 rows tie at 0.95 → ambiguous: true + candidates "Parle G" → folding and the single-character token still land it a real name → runner-up far behind → match, unchanged Dropped with it: scanSpecificEnough, the per-hit text score, and the proportional confirmation bonus — the flat +0.10 is back. Simpler, and it leaves main's tuning untouched. TestTextScoreRewardsSpecificityNotJustOverlap tested the removed formula and is replaced by TestABrandNameScoresItsProductsIdentically, which guards the tie itself: a formula that broke it on name length or word count would bring the bug back. Docs carry both rationales, and now say plainly that confidence stays high on the ambiguous path — gate on `ambiguous`, never on `confidence`. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
442 lines
21 KiB
Markdown
442 lines
21 KiB
Markdown
# Scan-to-order — mobile integration
|
||
|
||
A customer photographs a product. Google Lens (on the phone) turns the photo
|
||
into a label — `"Milk Bikis"`, `"Dabur Honey 500g"`. The app sends that label
|
||
here and gets back: what the product is, which of the customer's stores sell
|
||
it, in which sizes, with live stock, nearest first, and which store we
|
||
recommend. When the customer taps a store and a size, a second call confirms
|
||
the shelf still has it — and if it does not, names the next-nearest store
|
||
that does.
|
||
|
||
When the label fits several products — `"britannia"` names 258 of them — it
|
||
answers with a short "did you mean?" list instead of picking one, because a
|
||
confident price on the wrong biscuit is worse than one extra tap.
|
||
|
||
Base path: `/live/api/v1/mob/scan`. Every response uses the usual envelope
|
||
`{ code, status, message, details }`; the shapes below are `details`.
|
||
|
||
## The flow
|
||
|
||
```
|
||
photo ──Lens──▶ label
|
||
│
|
||
▼
|
||
POST /lookup ───▶ ambiguous:true + candidates[] "did you mean?"
|
||
│ │
|
||
│ customer taps one candidate
|
||
│ │
|
||
│ POST /lookup { brand, catalogueid }
|
||
│ │
|
||
└───▶ match + stores[] (recommended first) ◀──┘
|
||
│
|
||
customer taps a store + a size
|
||
│
|
||
▼
|
||
POST /confirm ───▶ ok:true → add to basket with existing order APIs
|
||
ok:false + alternative → offer the other store
|
||
```
|
||
|
||
**`/lookup` has two possible answers and the app must handle both.** A label
|
||
that names one product comes back with `match` + `stores`. A label that fits
|
||
several — a bare brand name like `"britannia"`, a generic word like
|
||
`"biscuits"` — comes back with `ambiguous: true` and `candidates`, and the
|
||
app asks the customer which one before any price is shown. Lens returns a
|
||
bare wordmark often, because it is usually the biggest thing printed on a
|
||
packet, so this is a normal path and not an error case.
|
||
|
||
`GET /stores` is for the "choose another shop" sheet: the customer's
|
||
registered stores, nearest first, independent of any product.
|
||
|
||
## `POST /lookup`
|
||
|
||
Note the `//` notes below are annotations, not JSON — strip them.
|
||
|
||
```json
|
||
{
|
||
"customerid": 5123,
|
||
"label": "Milk Bikis",
|
||
"latitude": 11.0290, // phone fix; optional — saved address is used without it
|
||
"longitude": 77.0290,
|
||
"tenantids": [1135, 1140], // optional: what the app THINKS the customer joined
|
||
"limit": 0 // optional: max stores, 0 = all
|
||
|
||
// Instead of a label: name the product outright. This is how you resolve
|
||
// a candidate the customer tapped, and how a deep link or a "buy again"
|
||
// skips recognition. With both set, `label` is ignored.
|
||
// "brand": "britannia", "catalogueid": 7
|
||
}
|
||
```
|
||
|
||
`label` is required **unless** `brand` and `catalogueid` are both given.
|
||
|
||
`tenantids` is verified, never trusted: the server intersects it with the
|
||
`tenantcustomers` table. Ids the customer is not actually registered with
|
||
come back in `unregistered_tenantids` — treat that as "refresh the local
|
||
list". A list that matches nothing at all is treated as stale and all
|
||
registered stores are used.
|
||
|
||
### Response A — one product identified
|
||
|
||
`ambiguous: false`, `match` set, `candidates` empty.
|
||
|
||
```json
|
||
{
|
||
"label": "Milk Bikis",
|
||
"match": {
|
||
"brand": "britannia", "catalogueid": 7, "imageid": "britannia_milk_bikis_100g",
|
||
"product_name": "Milk Bikis", "size": "100 g", "variant_key": "milk_bikis",
|
||
"image": "https://…", "score": 0.94, "method": "vector+text"
|
||
},
|
||
"catalogue_variants": [ { "…same shape…": "100 g" }, { "…": "200 g" } ],
|
||
"ambiguous": false,
|
||
"candidates": [],
|
||
"confidence": 0.94,
|
||
"available": true,
|
||
"recommended_locationid": 20,
|
||
"stores": [
|
||
{
|
||
"tenantid": 2, "tenantname": "R Mart", "locationid": 20, "locationname": "Hopes",
|
||
"latitude": 11.01, "longitude": 77.0, "distance_km": 3.8, "open": true,
|
||
"deliveryradius": 5, "deliverymins": 30,
|
||
"recommended": true, "available": true,
|
||
"options": [
|
||
{ "productid": 200, "productname": "Milk Bikis 100g", "size": "100 g", "price": 12, "stock": 6,
|
||
"available": true, "is_variant": false, "matched_by": "imageid", "image": "…" },
|
||
{ "productid": 201, "productname": "Milk Bikis 200g", "size": "200 g", "price": 22, "stock": 3,
|
||
"available": true, "is_variant": true, "variantname": "200 g", "matched_by": "variant-of:200" }
|
||
]
|
||
},
|
||
{ "locationid": 10, "locationname": "Peelamedu", "distance_km": 0.9, "available": false, "recommended": false,
|
||
"options": [ { "productid": 100, "stock": 0, "available": false, "…": "…" } ] }
|
||
],
|
||
"unregistered_tenantids": [],
|
||
"message": "Available at 1 of your stores."
|
||
}
|
||
```
|
||
|
||
### Response B — several products fit, none clearly
|
||
|
||
`ambiguous: true`, `match: null`, `stores: []`. Show a "did you mean?" list.
|
||
|
||
```json
|
||
{
|
||
"label": "britannia",
|
||
"match": null,
|
||
"ambiguous": true,
|
||
"candidates": [
|
||
{ "brand": "britannia", "catalogueid": 23, "product_name": "Britannia Marie Gold",
|
||
"size": "250 g", "image": "https://…", "score": 0.95, "method": "text", "available": true },
|
||
{ "brand": "britannia", "catalogueid": 22, "product_name": "Britannia Good Day Butter Cookies",
|
||
"image": "https://…", "score": 0.95, "method": "text" },
|
||
{ "brand": "britannia", "catalogueid": 21, "product_name": "Britannia Good Day Cashew Cookies",
|
||
"image": "https://…", "score": 0.95, "method": "text" }
|
||
],
|
||
"confidence": 0.95,
|
||
"available": false,
|
||
"stores": [],
|
||
"catalogue_variants": [],
|
||
"message": "Which one is it? 1 of these 3 are in stock near you."
|
||
}
|
||
```
|
||
|
||
- **`confidence` is not low here, and that is not a bug.** "britannia" really
|
||
does appear in all three names, so relevance is high — what is missing is
|
||
*identification*. Gate on `ambiguous`, never on `confidence`: an app that
|
||
reads 0.95 as "sure enough to show a price" reintroduces the exact bug this
|
||
path exists to prevent.
|
||
- **`available` on a candidate** means at least one of the customer's
|
||
registered stores has it in stock right now. Candidates are ordered
|
||
available-first, so the list can show what is buyable before what is not
|
||
— and the field is absent (not `false`) when unavailable, so read it as
|
||
falsy, not as a required key.
|
||
- **To resolve a pick**, call `/lookup` again with that candidate's `brand`
|
||
and `catalogueid` and no label. You get Response A for that exact product,
|
||
with `method: "direct"` and `confidence: 1`.
|
||
- At most 10 candidates come back.
|
||
|
||
### How to read either response
|
||
|
||
- `match == null && !ambiguous` → nothing recognised; show `message` and let
|
||
them retry with a clearer photo.
|
||
- `ambiguous: true` → ask, do not guess. Never show a price on this path;
|
||
`stores` is deliberately empty.
|
||
- `confidence` below ~0.5 with a `match` → recognised but unsure; worth
|
||
confirming the name before showing prices. `method: "text"` means no
|
||
embedding model was involved (not configured, or it timed out) — be a
|
||
little more cautious. `method: "direct"` means the caller named the
|
||
product, so nothing was recognised at all.
|
||
- `stores` is ordered **in-stock first, then nearest**. Exactly one store has
|
||
`recommended: true` — the nearest with stock — and only when `available`
|
||
is true. Stores that sell it but have nothing on the shelf are still listed
|
||
(so the customer understands why they are not recommended); stores that do
|
||
not sell it are not.
|
||
- `options` are the things that can actually go in a basket at that store —
|
||
the matched product and each of its sizes — each a real product with its
|
||
own `productid`, price and live `stock`. Use `productid` in the existing
|
||
cart/order calls exactly as you would from the catalogue screen.
|
||
- `distance_km: -1` means the distance is unknown (no fix from the phone and
|
||
no saved address, or the store has no coordinates). Do not render it as 0.
|
||
Send `latitude`/`longitude` on `/confirm` too if you display distance from
|
||
its reply: the saved address is only consulted there when the shelf is
|
||
empty and alternatives have to be ranked, so without a fix the store you
|
||
tapped comes back `-1`.
|
||
|
||
## `POST /confirm`
|
||
|
||
Sent when the customer taps a store and an option. Re-reads live stock —
|
||
nothing is cached on this path.
|
||
|
||
```json
|
||
{ "customerid": 5123, "tenantid": 1, "locationid": 10, "productid": 100, "quantity": 2,
|
||
"latitude": 11.029, "longitude": 77.029 }
|
||
```
|
||
|
||
```json
|
||
{
|
||
"ok": false,
|
||
"reason": "out_of_stock", // in_stock | insufficient_stock | out_of_stock | not_sold_here | store_not_registered
|
||
"store": { "…the store they tapped…" }, // distance_km filled from the fix you send
|
||
"option": { "productid": 100, "stock": 0, "…": "…" },
|
||
"requested": 2,
|
||
"alternative": { // absent when nobody has enough
|
||
"locationid": 20, "locationname": "Hopes", "distance_km": 3.8, "recommended": true, "available": true,
|
||
"options": [ { "productid": 200, "stock": 6, "price": 12, "…": "…" } ]
|
||
},
|
||
"message": "Out of stock at Peelamedu. Hopes has it (3.8 km away)."
|
||
}
|
||
```
|
||
|
||
`ok: true` → proceed to the basket. `ok: false` → show `message`; if
|
||
`alternative` is present offer it as a one-tap switch (it is the **same
|
||
product**, not another size — the customer chose a size and we do not
|
||
substitute). These are HTTP 200s: they are answers, not errors.
|
||
|
||
## `GET /stores?customerid=5123&latitude=11.029&longitude=77.029`
|
||
|
||
The customer's registered stores, nearest first, `distance_km: -1` last.
|
||
Same `ScanStore` shape as inside `stores[]` above, without options.
|
||
|
||
## Errors (HTTP status ≠ 200)
|
||
|
||
| Status | When |
|
||
|---|---|
|
||
| 400 | Missing `customerid`/`label`/ids, or a body that is not JSON. `message` says which. |
|
||
| 404 | `customerid` does not exist. |
|
||
| 503 | The catalogue database is not reachable. Retry later; the rest of the app is unaffected. |
|
||
| 500 | Anything else. Logged server-side. |
|
||
|
||
## Behind the curtain (for whoever operates it)
|
||
|
||
- **Recognition** = pgvector cosine search over every `brand_*` table in the
|
||
catalogue (each with its own index, merged), plus a word match on
|
||
`product_name`/`title`/`search_query` that settles near-ties and works on
|
||
its own when no embedding model is configured. The model is set by
|
||
`EMBEDDING_PROVIDER/MODEL/API_KEY` and **must** be the one that indexed
|
||
the catalogue — the first search checks the vector width and refuses a
|
||
mismatch by name.
|
||
- **The word match asks for most of the label, not all of it**
|
||
(`minTokenHits`: two thirds, rounded up, and both of a two-word label).
|
||
Requiring every word meant one word the catalogue does not use took the
|
||
right product out of the running entirely — "Dettol bottle pack" retrieved
|
||
no Dettol, "Parle G biscuit pack" retrieved no Parle-G — and the vector
|
||
search then answered alone, confidently and wrongly, at a score the floor
|
||
could not catch. Each brand's rows are ordered by how much of the label
|
||
they carry (the whole label as a substring outranks any number of loose
|
||
words) so that the per-brand `LIMIT` keeps the best rows and not merely the
|
||
first ones the planner reached. Packaging words — "pack", "bottle", "jar",
|
||
"sachet" and friends, see `utils.isPackaging` — are dropped before any of
|
||
this, like pack sizes, unless the label is nothing else.
|
||
- **The catalogue's model** (verified 2026-09-15 by cosine against a stored
|
||
row: 1.0000): `all-MiniLM-L6-v2`, 384-d, unit-normalised, embedding the
|
||
`search_query` column (brand + name + category + blurb + price range).
|
||
Ollama ships it as `all-minilm`; the cluster's `ollama.krow` service serves
|
||
it, so production is:
|
||
```
|
||
EMBEDDING_PROVIDER=openai
|
||
EMBEDDING_BASE_URL=http://ollama.krow.svc.cluster.local:11434/v1
|
||
EMBEDDING_MODEL=all-minilm
|
||
EMBEDDING_API_KEY=ollama # any non-empty value; Ollama ignores it
|
||
EMBEDDING_DIMENSIONS=384
|
||
```
|
||
A bare label ("Milk Bikis") scores ~0.92 against its product's stored
|
||
vector and ~0.23 against an unrelated one, which is what the 0.50 floor in
|
||
`scanService.go` is set against — the middle of that split, not the edge of
|
||
the noise. It was 0.30 until a near-miss got through in production
|
||
("Paracetamol" → "Paneer Makhni 500ml", 0.304). If the catalogue team ever
|
||
re-embeds with another model, change `EMBEDDING_MODEL`/`DIMENSIONS` here
|
||
and nothing else.
|
||
- **Speed**: the label's vector (7 days) and the ranked catalogue hits
|
||
(30 min) are cached in Redis and in-process, so a popular product costs
|
||
one model call platform-wide. Customer, stores and catalogue are read
|
||
concurrently; the whole lookup is capped at 5 s and a slow model degrades
|
||
to a text answer instead of a spinner. Live stock is one indexed query and
|
||
is never cached.
|
||
- **Availability** is the same rule the app's catalogue screen uses:
|
||
`products.approve = 1`, `productlocations.publishedat IS NOT NULL`, stock =
|
||
live `SUM(in) − SUM(out)` of `productstocks` at that outlet, price = the
|
||
outlet's own price else the tenant's retail price.
|
||
- **No reservation.** Confirm re-reads the ledger; a hold would give the
|
||
same answer with a timer to babysit. If contention becomes real, a
|
||
Redis-backed short hold slots in at `Confirm` without changing the API.
|
||
- **Identity** is the `customerid` in the body, like every other mobile
|
||
endpoint here — there is no auth layer yet (see `SECURITY_HANDOFF.md`).
|
||
|
||
## Two decisions, and why
|
||
|
||
Both come from a proposal (2026-09-23) to have the app send vectors it
|
||
computed on the phone. Recorded here because the next person will ask.
|
||
|
||
### The app does not send `textvector`
|
||
|
||
An on-device MiniLM vector is only comparable to the catalogue's if the app
|
||
ships the identical model *and* tokenizer *and* pooling *and* normalisation;
|
||
a quantised tflite build usually drifts, and the failure is silent — the
|
||
ranking just gets worse. There is also nothing to gain: the server-side
|
||
embed is ~30 ms warm and the result is cached in Redis by label, so one
|
||
model call serves every customer who scans that product. A client-supplied
|
||
vector *defeats* that cache (the key would have to be the vector, not the
|
||
label), and 384 floats is ~5 KB of upload against ~12 bytes for
|
||
`"Milk Bikis"`. If the field ever arrives it can be accepted and validated,
|
||
but the app should not be asked to compute it.
|
||
|
||
**Send the full OCR text instead** if you want to give the server more to
|
||
work with — ~100 bytes, no model coupling, strictly more information than a
|
||
single label.
|
||
|
||
### The app does not send `imagevector` — yet
|
||
|
||
The catalogue *does* carry image vectors: every `brand_*` table has
|
||
`img_vector vector(1024)`, filled on 1885 of 2124 rows (empty in
|
||
`brand_haldirams`, `brand_kaleesuwari`, `brand_mdh`, `brand_zzsmoketest`).
|
||
That matches the proposed MobileNetV3-Small embedder, so the idea is
|
||
coherent and half-built — this flow simply does not read that column.
|
||
|
||
It stays unread for now because **Google Lens is already the image
|
||
recogniser, and a far better one**: photo → Lens → label is Google's product
|
||
recognition, trained on billions of images. Putting a 137M-parameter
|
||
ImageNet backbone searching 1885 vectors *behind* that adds little where
|
||
Lens succeeds, and MobileNetV3-Small — which struggles to tell one blue
|
||
biscuit wrapper from another — is unlikely to rescue the cases where Lens
|
||
fails. There is also an unverified dependency: the preprocessing the app
|
||
would use (BGR → centre crop → 224×224 INTER_AREA → RGB → `/255.0`) has to
|
||
match whatever the catalogue pipeline actually ran, or the search returns
|
||
confidently-ranked noise.
|
||
|
||
**What would change this:** the field data. Once live, count how often
|
||
`/lookup` returns `ambiguous: true` or nothing recognised. If Lens labels are
|
||
reliable, image search is polish; if that number is high, it becomes the
|
||
priority — and the first task is the cosine check (embed a known catalogue
|
||
product's image through the app's exact pipeline, compare with its stored
|
||
`img_vector`; ≈0.99 means the contract holds), not writing the query.
|
||
|
||
There is one non-recognition argument for it worth remembering: on-device
|
||
inference is free and needs no Google dependency, which matters if Cloud
|
||
Vision costs start to bite at volume. That is a business reason, not a
|
||
quality one.
|
||
|
||
## For backend developers
|
||
|
||
### Where the code is
|
||
|
||
| File | Holds |
|
||
|---|---|
|
||
| `models/scan.go` | request/response shapes (`ScanLookupRequest`, `ScanStoreOffer`, `ScanOption`, …) |
|
||
| `repositories/scanRepository.go` | all SQL: registered stores, live options, vector + text search, the two-tier cache |
|
||
| `services/scanService.go` | the pipeline: parallel reads, scoring, family grouping, ranking, confirm fallback |
|
||
| `controllers/scanController.go` | the three handlers and the error → status mapping |
|
||
| `routes/scanroutes.go` | `/v1/mob/scan/*` |
|
||
| `utils/embedding.go` | `Embedder` interface, OpenAI-compatible and Gemini clients |
|
||
| `utils/geo.go` | coordinate parsing, haversine, opening hours, label tokenising |
|
||
| `config/config.go` | `EmbeddingConfig` and its validation |
|
||
| `scratch/cataloguedims` | read-only check of every catalogue vector column's width and fill |
|
||
|
||
### Try it locally
|
||
|
||
```sh
|
||
go run . # with the local compose stack; EMBEDDING_* unset → text-only, still works
|
||
curl -s localhost:1122/live/api/v1/mob/scan/lookup -H 'Content-Type: application/json' \
|
||
-d '{"customerid":1,"label":"Milk Bikis","latitude":11.03,"longitude":77.03}' | jq .details
|
||
```
|
||
|
||
To exercise the vector path locally, run Ollama on your Mac
|
||
(`ollama pull all-minilm`) and set `EMBEDDING_PROVIDER=openai`,
|
||
`EMBEDDING_BASE_URL=http://localhost:11434/v1`, `EMBEDDING_MODEL=all-minilm`,
|
||
`EMBEDDING_API_KEY=ollama`, `EMBEDDING_DIMENSIONS=384` in `.env.local`. The
|
||
local catalogue must carry vectors from the same model for results to mean
|
||
anything; a schema-only dump does not.
|
||
|
||
### Tests
|
||
|
||
`go test ./services -run 'Lookup|Confirm|Stores|Brand|Ambiguous|Specific|TextScore|Distinct|Naming'`
|
||
drives the whole pipeline through a fake repository
|
||
(`services/scan_test.go`); no database. `go test ./utils` covers both HTTP
|
||
clients against `httptest` servers, and the geo helpers. Add a case to
|
||
`scan_test.go`'s fixture when you change ranking — it is the spec, and
|
||
`newBrandLabelFixture` in particular is the regression guard for the
|
||
brand-name bug described under Scoring.
|
||
|
||
### Knobs (constants in `scanService.go`)
|
||
|
||
| Constant | Default | Effect |
|
||
|---|---|---|
|
||
| `scanLookupTimeout` | 5 s | whole lookup, including the model call |
|
||
| `scanCatalogueTopK` | 15 | rows taken from each brand table and from the merge |
|
||
| `scanMinScore` | 0.50 | below this the best hit is not shown as a match |
|
||
| `scanAmbiguityMargin` | 0.06 | how close the runner-up may be before the answer becomes a question |
|
||
| `scanMaxCandidates` | 10 | longest "did you mean?" list |
|
||
| `embedTimeout` (`utils/embedding.go`) | 4 s | one model call |
|
||
| `scanVectorTTL` / `scanHitsTTL` (`scanRepository.go`) | 7 d / 30 min | cache lifetimes |
|
||
|
||
Scores: vector = `1 − cosine distance`; text = 0.95 for the whole label
|
||
inside the name, else `0.8 × (label words found / label words)`; combined =
|
||
`max(vector, text) + 0.10` when both hit, capped at 1. Ties are broken by
|
||
cosine distance — nearest first, a text-only row last — and only then by
|
||
name.
|
||
|
||
The label and the product name are both separator-folded before that
|
||
substring test (`utils.FoldSeparators`), and compared again with separators
|
||
removed (`utils.TightenLabel`, labels of 4+ characters), so the brand's own
|
||
punctuation does not decide the match: "Parle G", "Parle-G" and "ParleG" all
|
||
reach *Parle-G Original Glucose Biscuits*. A single-character token survives
|
||
tokenising when it follows a word, because it is often the whole name — the
|
||
"G" of Parle-G, the "K" of Special K. It is still dropped when it stands
|
||
alone or is a pack multiplier.
|
||
|
||
All three mattered at once: before this, "Parle G" tied with *Parle Monaco
|
||
Classic* at 0.9 (the "G" was dropped, so only "parle" matched either row),
|
||
and the name tie-break handed it to Monaco because a space precedes a hyphen
|
||
in ASCII. A confident, wrong answer — the kind no score floor can catch.
|
||
|
||
**When the substring rule ties, that tie is the answer.** A bare brand name
|
||
is a substring of every one of that brand's names, so all of them score 0.95
|
||
— identically, at a high score no floor would ever catch. Rather than
|
||
scoring around it, `isAmbiguous` reads it: if the runner-up is within
|
||
`scanAmbiguityMargin` of the leader, the reply becomes `ambiguous: true`
|
||
with `candidates` instead of a match (see Response B). Erring towards asking
|
||
is deliberate — one tap on a picture against the wrong biscuit. A label that
|
||
names one product leaves the runner-up far behind, so the common case is
|
||
untouched, and `services/scan_test.go`'s
|
||
`TestABrandNameScoresItsProductsIdentically` guards the tie itself: a
|
||
formula that broke it on name length or word count would bring the bug
|
||
back.
|
||
|
||
### Changing the embedding model
|
||
|
||
1. The catalogue team re-embeds `search_query` with the new model.
|
||
2. Serve it (Ollama pull, or a hosted key).
|
||
3. Change `EMBEDDING_MODEL` / `EMBEDDING_DIMENSIONS` (and provider/URL if
|
||
needed) in the cluster; roll the pods.
|
||
4. Flush the hit cache if you cannot wait 30 min: keys are
|
||
`scan:hits:v1:*` and `scan:emb:v1:*` in Redis (they are also keyed by
|
||
model name, so old entries simply stop being read).
|
||
|
||
Nothing in Go changes. A width mismatch fails the first search with an error
|
||
naming both numbers.
|
||
|
||
### Adding a provider
|
||
|
||
Implement `utils.Embedder` (`Embed(ctx, text) ([]float32, error)` and
|
||
`Model() string`), add a case to `NewEmbedder`, and add the provider name to
|
||
the allow-list in `config.validate`. Keep the HTTP client timeout: the
|
||
customer is holding a phone.
|