Files
backend_fiesta/docs/SCAN_TO_ORDER.md
Suriyakumarvijayanayagam 01bc89ab77 Ask instead of guessing when a label fits several products
`"britannia"` is a substring of all 258 Britannia product names, and
textScore returned 0.95 for any product whose name contained the label.
So every one of them tied, the tie broke alphabetically, and the customer
was shown one arbitrary biscuit with "confidence": 0.95 and a price. Lens
hands back a bare wordmark often — it is usually the biggest thing printed
on a packet — so this was the common case, not an edge one. Found via the
example request in the mobile team's own proposal.

Scoring now asks both questions. A hit carries `score` (ranks) and `text`
(how specifically the label names THIS product: the harmonic mean of how
much of the label the product explains and how much of the product's name
the label explains, pack sizes dropped from both sides). A brand name
scores its products ~0.33 equally instead of 0.95 arbitrarily. The
"vector and text agree" bonus is now proportional to the text score, so a
weak match can no longer inflate a whole brand.

isAmbiguous reads that: the leader is a guess if anything is level with it
(margin) or if the label names no one product (specificity), and then the
response carries `ambiguous: true` with `candidates` — distinct products,
not pack sizes, at most ten, each marked with whether one of the
customer's stores has it in stock, available ones first. `match` is nil
and `stores` empty on that path: no price for a product nobody chose.
Erring towards asking is deliberate — a tap versus the wrong biscuit.

To act on a pick, /lookup now accepts `brand` + `catalogueid` instead of a
label and skips recognition entirely (also serves deep links and re-order).
New: ScanRepository.CatalogueRef, resolving via the brand tables discovered
from information_schema, never a name built from the request.

Also: scratch/cataloguedims now reports every vector column, not just
`embedding` — which is how we learned the catalogue also carries
img_vector(1024), filled on 1885 of 2124 rows. SCAN_TO_ORDER.md records
why that column stays unread for now and what would change it, alongside
why the app is not asked to compute vectors on the phone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-23 10:56:02 +05:30

19 KiB
Raw Blame History

Scan-to-order — mobile integration

A customer photographs a product. Google Lens (on the phone) turns the photo into a label — "Milk Bikis", "Dabur Honey 500g". The app sends that label here and gets back: what the product is, which of the customer's stores sell it, in which sizes, with live stock, nearest first, and which store we recommend. When the customer taps a store and a size, a second call confirms the shelf still has it — and if it does not, names the next-nearest store that does.

When the label fits several products — "britannia" names 258 of them — it answers with a short "did you mean?" list instead of picking one, because a confident price on the wrong biscuit is worse than one extra tap.

Base path: /live/api/v1/mob/scan. Every response uses the usual envelope { code, status, message, details }; the shapes below are details.

The flow

photo ──Lens──▶ label
                 │
                 ▼
   POST /lookup  ───▶  ambiguous:true + candidates[]   "did you mean?"
                 │                        │
                 │      customer taps one candidate
                 │                        │
                 │      POST /lookup { brand, catalogueid }
                 │                        │
                 └───▶  match + stores[] (recommended first)  ◀──┘
                 │
   customer taps a store + a size
                 │
                 ▼
   POST /confirm ───▶  ok:true            → add to basket with existing order APIs
                       ok:false + alternative → offer the other store

/lookup has two possible answers and the app must handle both. A label that names one product comes back with match + stores. A label that fits several — a bare brand name like "britannia", a generic word like "biscuits" — comes back with ambiguous: true and candidates, and the app asks the customer which one before any price is shown. Lens returns a bare wordmark often, because it is usually the biggest thing printed on a packet, so this is a normal path and not an error case.

GET /stores is for the "choose another shop" sheet: the customer's registered stores, nearest first, independent of any product.

POST /lookup

Note the // notes below are annotations, not JSON — strip them.

{
  "customerid": 5123,
  "label": "Milk Bikis",
  "latitude": 11.0290,          // phone fix; optional — saved address is used without it
  "longitude": 77.0290,
  "tenantids": [1135, 1140],    // optional: what the app THINKS the customer joined
  "limit": 0                    // optional: max stores, 0 = all

  // Instead of a label: name the product outright. This is how you resolve
  // a candidate the customer tapped, and how a deep link or a "buy again"
  // skips recognition. With both set, `label` is ignored.
  // "brand": "britannia", "catalogueid": 7
}

label is required unless brand and catalogueid are both given.

tenantids is verified, never trusted: the server intersects it with the tenantcustomers table. Ids the customer is not actually registered with come back in unregistered_tenantids — treat that as "refresh the local list". A list that matches nothing at all is treated as stale and all registered stores are used.

Response A — one product identified

ambiguous: false, match set, candidates empty.

{
  "label": "Milk Bikis",
  "match": {
    "brand": "britannia", "catalogueid": 7, "imageid": "britannia_milk_bikis_100g",
    "product_name": "Milk Bikis", "size": "100 g", "variant_key": "milk_bikis",
    "image": "https://…", "score": 0.94, "method": "vector+text"
  },
  "catalogue_variants": [ { "…same shape…": "100 g" }, { "…": "200 g" } ],
  "ambiguous": false,
  "candidates": [],
  "confidence": 0.94,
  "available": true,
  "recommended_locationid": 20,
  "stores": [
    {
      "tenantid": 2, "tenantname": "R Mart", "locationid": 20, "locationname": "Hopes",
      "latitude": 11.01, "longitude": 77.0, "distance_km": 3.8, "open": true,
      "deliveryradius": 5, "deliverymins": 30,
      "recommended": true, "available": true,
      "options": [
        { "productid": 200, "productname": "Milk Bikis 100g", "size": "100 g", "price": 12, "stock": 6,
          "available": true, "is_variant": false, "matched_by": "imageid", "image": "…" },
        { "productid": 201, "productname": "Milk Bikis 200g", "size": "200 g", "price": 22, "stock": 3,
          "available": true, "is_variant": true, "variantname": "200 g", "matched_by": "variant-of:200" }
      ]
    },
    { "locationid": 10, "locationname": "Peelamedu", "distance_km": 0.9, "available": false, "recommended": false,
      "options": [ { "productid": 100, "stock": 0, "available": false, "…": "…" } ] }
  ],
  "unregistered_tenantids": [],
  "message": "Available at 1 of your stores."
}

Response B — several products fit, none clearly

ambiguous: true, match: null, stores: []. Show a "did you mean?" list.

{
  "label": "britannia",
  "match": null,
  "ambiguous": true,
  "candidates": [
    { "brand": "britannia", "catalogueid": 23, "product_name": "Britannia Marie Gold",
      "size": "250 g", "image": "https://…", "score": 0.5, "method": "text", "available": true },
    { "brand": "britannia", "catalogueid": 22, "product_name": "Britannia Good Day Butter Cookies",
      "image": "https://…", "score": 0.333, "method": "text" },
    { "brand": "britannia", "catalogueid": 21, "product_name": "Britannia Good Day Cashew Cookies",
      "image": "https://…", "score": 0.333, "method": "text" }
  ],
  "confidence": 0.5,
  "available": false,
  "stores": [],
  "catalogue_variants": [],
  "message": "Which one is it? 1 of these 3 are in stock near you."
}
  • available on a candidate means at least one of the customer's registered stores has it in stock right now. Candidates are ordered available-first, so the list can show what is buyable before what is not — and the field is absent (not false) when unavailable, so read it as falsy, not as a required key.
  • To resolve a pick, call /lookup again with that candidate's brand and catalogueid and no label. You get Response A for that exact product, with method: "direct" and confidence: 1.
  • At most 10 candidates come back.

How to read either response

  • match == null && !ambiguous → nothing recognised; show message and let them retry with a clearer photo.
  • ambiguous: true → ask, do not guess. Never show a price on this path; stores is deliberately empty.
  • confidence below ~0.5 with a match → recognised but unsure; worth confirming the name before showing prices. method: "text" means no embedding model was involved (not configured, or it timed out) — be a little more cautious. method: "direct" means the caller named the product, so nothing was recognised at all.
  • stores is ordered in-stock first, then nearest. Exactly one store has recommended: true — the nearest with stock — and only when available is true. Stores that sell it but have nothing on the shelf are still listed (so the customer understands why they are not recommended); stores that do not sell it are not.
  • options are the things that can actually go in a basket at that store — the matched product and each of its sizes — each a real product with its own productid, price and live stock. Use productid in the existing cart/order calls exactly as you would from the catalogue screen.
  • distance_km: -1 means the distance is unknown (no fix from the phone and no saved address, or the store has no coordinates). Do not render it as 0.

POST /confirm

Sent when the customer taps a store and an option. Re-reads live stock — nothing is cached on this path.

{ "customerid": 5123, "tenantid": 1, "locationid": 10, "productid": 100, "quantity": 2,
  "latitude": 11.029, "longitude": 77.029 }
{
  "ok": false,
  "reason": "out_of_stock",           // in_stock | insufficient_stock | out_of_stock | not_sold_here | store_not_registered
  "store":  { "…the store they tapped…" },
  "option": { "productid": 100, "stock": 0, "…": "…" },
  "requested": 2,
  "alternative": {                     // absent when nobody has enough
    "locationid": 20, "locationname": "Hopes", "distance_km": 3.8, "recommended": true, "available": true,
    "options": [ { "productid": 200, "stock": 6, "price": 12, "…": "…" } ]
  },
  "message": "Out of stock at Peelamedu. Hopes has it (3.8 km away)."
}

ok: true → proceed to the basket. ok: false → show message; if alternative is present offer it as a one-tap switch (it is the same product, not another size — the customer chose a size and we do not substitute). These are HTTP 200s: they are answers, not errors.

GET /stores?customerid=5123&latitude=11.029&longitude=77.029

The customer's registered stores, nearest first, distance_km: -1 last. Same ScanStore shape as inside stores[] above, without options.

Errors (HTTP status ≠ 200)

Status When
400 Missing customerid/label/ids, or a body that is not JSON. message says which.
404 customerid does not exist.
503 The catalogue database is not reachable. Retry later; the rest of the app is unaffected.
500 Anything else. Logged server-side.

Behind the curtain (for whoever operates it)

  • Recognition = pgvector cosine search over every brand_* table in the catalogue (each with its own index, merged), plus a word match on product_name/title/search_query that settles near-ties and works on its own when no embedding model is configured. The model is set by EMBEDDING_PROVIDER/MODEL/API_KEY and must be the one that indexed the catalogue — the first search checks the vector width and refuses a mismatch by name.
  • The catalogue's model (verified 2026-09-15 by cosine against a stored row: 1.0000): all-MiniLM-L6-v2, 384-d, unit-normalised, embedding the search_query column (brand + name + category + blurb + price range). Ollama ships it as all-minilm; the cluster's ollama.krow service serves it, so production is:
    EMBEDDING_PROVIDER=openai
    EMBEDDING_BASE_URL=http://ollama.krow.svc.cluster.local:11434/v1
    EMBEDDING_MODEL=all-minilm
    EMBEDDING_API_KEY=ollama        # any non-empty value; Ollama ignores it
    EMBEDDING_DIMENSIONS=384
    
    A bare label ("Milk Bikis") scores ~0.92 against its product's stored vector and ~0.23 against an unrelated one, which is what the 0.30 floor in scanService.go is set against. If the catalogue team ever re-embeds with another model, change EMBEDDING_MODEL/DIMENSIONS here and nothing else.
  • Speed: the label's vector (7 days) and the ranked catalogue hits (30 min) are cached in Redis and in-process, so a popular product costs one model call platform-wide. Customer, stores and catalogue are read concurrently; the whole lookup is capped at 5 s and a slow model degrades to a text answer instead of a spinner. Live stock is one indexed query and is never cached.
  • Availability is the same rule the app's catalogue screen uses: products.approve = 1, productlocations.publishedat IS NOT NULL, stock = live SUM(in) − SUM(out) of productstocks at that outlet, price = the outlet's own price else the tenant's retail price.
  • No reservation. Confirm re-reads the ledger; a hold would give the same answer with a timer to babysit. If contention becomes real, a Redis-backed short hold slots in at Confirm without changing the API.
  • Identity is the customerid in the body, like every other mobile endpoint here — there is no auth layer yet (see SECURITY_HANDOFF.md).

Two decisions, and why

Both come from a proposal (2026-09-23) to have the app send vectors it computed on the phone. Recorded here because the next person will ask.

The app does not send textvector

An on-device MiniLM vector is only comparable to the catalogue's if the app ships the identical model and tokenizer and pooling and normalisation; a quantised tflite build usually drifts, and the failure is silent — the ranking just gets worse. There is also nothing to gain: the server-side embed is ~30 ms warm and the result is cached in Redis by label, so one model call serves every customer who scans that product. A client-supplied vector defeats that cache (the key would have to be the vector, not the label), and 384 floats is ~5 KB of upload against ~12 bytes for "Milk Bikis". If the field ever arrives it can be accepted and validated, but the app should not be asked to compute it.

Send the full OCR text instead if you want to give the server more to work with — ~100 bytes, no model coupling, strictly more information than a single label.

The app does not send imagevector — yet

The catalogue does carry image vectors: every brand_* table has img_vector vector(1024), filled on 1885 of 2124 rows (empty in brand_haldirams, brand_kaleesuwari, brand_mdh, brand_zzsmoketest). That matches the proposed MobileNetV3-Small embedder, so the idea is coherent and half-built — this flow simply does not read that column.

It stays unread for now because Google Lens is already the image recogniser, and a far better one: photo → Lens → label is Google's product recognition, trained on billions of images. Putting a 137M-parameter ImageNet backbone searching 1885 vectors behind that adds little where Lens succeeds, and MobileNetV3-Small — which struggles to tell one blue biscuit wrapper from another — is unlikely to rescue the cases where Lens fails. There is also an unverified dependency: the preprocessing the app would use (BGR → centre crop → 224×224 INTER_AREA → RGB → /255.0) has to match whatever the catalogue pipeline actually ran, or the search returns confidently-ranked noise.

What would change this: the field data. Once live, count how often /lookup returns ambiguous: true or nothing recognised. If Lens labels are reliable, image search is polish; if that number is high, it becomes the priority — and the first task is the cosine check (embed a known catalogue product's image through the app's exact pipeline, compare with its stored img_vector; ≈0.99 means the contract holds), not writing the query.

There is one non-recognition argument for it worth remembering: on-device inference is free and needs no Google dependency, which matters if Cloud Vision costs start to bite at volume. That is a business reason, not a quality one.

For backend developers

Where the code is

File Holds
models/scan.go request/response shapes (ScanLookupRequest, ScanStoreOffer, ScanOption, …)
repositories/scanRepository.go all SQL: registered stores, live options, vector + text search, the two-tier cache
services/scanService.go the pipeline: parallel reads, scoring, family grouping, ranking, confirm fallback
controllers/scanController.go the three handlers and the error → status mapping
routes/scanroutes.go /v1/mob/scan/*
utils/embedding.go Embedder interface, OpenAI-compatible and Gemini clients
utils/geo.go coordinate parsing, haversine, opening hours, label tokenising
config/config.go EmbeddingConfig and its validation
scratch/cataloguedims read-only check of every catalogue vector column's width and fill

Try it locally

go run .    # with the local compose stack; EMBEDDING_* unset → text-only, still works
curl -s localhost:1122/live/api/v1/mob/scan/lookup -H 'Content-Type: application/json' \
  -d '{"customerid":1,"label":"Milk Bikis","latitude":11.03,"longitude":77.03}' | jq .details

To exercise the vector path locally, run Ollama on your Mac (ollama pull all-minilm) and set EMBEDDING_PROVIDER=openai, EMBEDDING_BASE_URL=http://localhost:11434/v1, EMBEDDING_MODEL=all-minilm, EMBEDDING_API_KEY=ollama, EMBEDDING_DIMENSIONS=384 in .env.local. The local catalogue must carry vectors from the same model for results to mean anything; a schema-only dump does not.

Tests

go test ./services -run 'Lookup|Confirm|Stores|Brand|Ambiguous|Specific|TextScore|Distinct|Naming' drives the whole pipeline through a fake repository (services/scan_test.go); no database. go test ./utils covers both HTTP clients against httptest servers, and the geo helpers. Add a case to scan_test.go's fixture when you change ranking — it is the spec, and newBrandLabelFixture in particular is the regression guard for the brand-name bug described under Scoring.

Knobs (constants in scanService.go)

Constant Default Effect
scanLookupTimeout 5 s whole lookup, including the model call
scanCatalogueTopK 15 rows taken from each brand table and from the merge
scanMinScore 0.30 below this the best hit is not shown as a match
scanAmbiguityMargin 0.06 how close the runner-up may be before the answer becomes a question
scanSpecificEnough 0.55 text coverage the leader needs before it counts as identified
scanMaxCandidates 10 longest "did you mean?" list
embedTimeout (utils/embedding.go) 4 s one model call
scanVectorTTL / scanHitsTTL (scanRepository.go) 7 d / 30 min cache lifetimes

Scoring. Each hit carries two numbers, and they answer different questions:

  • score ranks. Vector = 1 − cosine distance. Combined = max(vector, text) + 0.10 × text, capped at 1 — the confirmation bonus is proportional, so only a text match that actually names the product strengthens a vector hit.
  • text says how specifically the label names this product, and is the harmonic mean of two coverages: how much of the label the product accounts for, and how much of the product's name the label accounts for. Pack sizes are dropped from both sides.

Why both: "britannia" is a substring of all 258 Britannia product names. Judging on overlap alone scored every one of them 0.95, the tie broke alphabetically, and one arbitrary biscuit came back with a price. Now they score ~0.33 equally, which isAmbiguous reads as "ask, don't guess" — via the margin test (something is level with the leader) or the specificity test (the leader may rank first on vector similarity while the label names no one product). Erring towards asking is deliberate: asking costs one tap on a picture, guessing wrong costs the customer's belief that the scanner works. An exact product name still scores ~1.0, so the common case is untouched.

Changing the embedding model

  1. The catalogue team re-embeds search_query with the new model.
  2. Serve it (Ollama pull, or a hosted key).
  3. Change EMBEDDING_MODEL / EMBEDDING_DIMENSIONS (and provider/URL if needed) in the cluster; roll the pods.
  4. Flush the hit cache if you cannot wait 30 min: keys are scan:hits:v1:* and scan:emb:v1:* in Redis (they are also keyed by model name, so old entries simply stop being read).

Nothing in Go changes. A width mismatch fails the first search with an error naming both numbers.

Adding a provider

Implement utils.Embedder (Embed(ctx, text) ([]float32, error) and Model() string), add a case to NewEmbedder, and add the provider name to the allow-list in config.validate. Keep the HTTP client timeout: the customer is holding a phone.