Catalog feature updates on column fields

This commit is contained in:
sriram
2026-09-08 15:18:29 +05:30
parent 2749bee1a3
commit 10b24c6348
60 changed files with 9224 additions and 31 deletions

198
docs/BARCODE_ENRICHMENT.md Normal file
View File

@@ -0,0 +1,198 @@
# Barcode enrichment
`app/services/enrichment/barcode/` — how a catalog row acquires a barcode, what
the barcode is then allowed to claim, and which knob to turn.
Referenced from `.env.example` and `app/services/enrichment/pipeline.py`, both
of which pointed at this file for a long time before it existed.
---
## The four ways a row gets a barcode
| # | Path | When | Cost | Confidence |
|---|---|---|---|---|
| 1 | **The sheet** | The merchant typed it | free | highest — they are holding the pack |
| 2 | **Identity derivation** | Always, inline | free, offline | derives `gtin`/`ean13`/`upc`/`barcode_type` from a barcode already present |
| 3 | **Bulk brand corpus** | After upload, per brand | ~5 requests **per brand** | name-matched at ≥ 0.88 |
| 4 | **Per-product cascade** | Only if enabled | 1 request **per product**, 10/min cap | brand + size + name matched |
Paths 1 and 2 are always on. Path 3 is on by default. Path 4 is off by default.
### Why path 4 is off and path 3 is on
They differ in cost, not in appetite for risk.
The Open Food Facts per-product search endpoint is capped at **10 requests per
minute**. A 200-row upload on path 4 is twenty minutes of a held HTTP request,
on a shared 8 GB host — which is what `settings.py:420-423` is about.
Path 3 fetches a brand's *entire* India catalogue in about five requests, caches
it to `data/cache/off_brand_corpus/<slug>.json`, and matches offline. A brand
costs the same whether it has 3 rows or 300.
```
ENABLE_BARCODE_LOOKUP=false # path 4: per product, inline
ENRICH_BARCODES_ON_UPLOAD=true # path 3: per brand, after the upload settles
```
> `.env.example` shipped `ENABLE_BARCODE_LOOKUP=true` while `settings.py`
> defaulted it `false`, for as long as both existed. Anyone copying the example
> got a materially different pipeline from anyone relying on defaults. Fixed;
> if you see the two disagree again, `settings.py` is authoritative.
---
## Two thresholds, and why they are different numbers
| Constant | Value | Direction | Used by |
|---|---|---|---|
| `BARCODE_MIN_NAME_SIMILARITY` | 0.78 | reverse — barcode known, name is a sanity check | `fetch_verified_nutrition_by_barcode` |
| `OFF bulk review_min` | 0.88 | forward — name carries the whole decision | `post_ingest_barcodes`, `backfill_barcodes_from_off` |
`settings.py:449-477` records the measurement behind 0.78: at 0.45 the cascade
accepted 15 candidates of which 8 were wrong; at 0.78 it accepted 2 and none
were wrong. **Do not lower it to fix a coverage complaint.**
### The containment rule
Open Food Facts stores short names. We store long ones. Measured over the
catalog on 2026-09-08, 149 of 300 barcoded rows were refused as "found, wrong
product" when the barcode had resolved perfectly:
```
ours "Nestle Munch 8.9g" OFF "Munch" similarity 0.332
ours "Coca-Cola Maaza 750ml" OFF "Maaza" similarity 0.304
```
`name_similarity` divides overlap by *our* token count, so a one-token candidate
cannot exceed ~0.33 however right it is. But the same run correctly refused:
```
ours "Pepsico Lays 1kg" OFF "Spanish tomato tango" 0.133
ours "Coca-Cola Fanta 750ml" OFF "Orange" 0.089
```
A threshold low enough to admit the first group admits the second. So
`matching.name_is_contained` separates them **structurally**: every candidate
token must be one of ours once brand and size tokens are discounted, and a bare
brand name ("Colgate", "godrej" — both real OFF titles) never matches.
### `barcode_is_identity` — and the gate that actually blocked most of them
The name gate was the *visible* symptom. When the fix was measured it moved only
3 rows to 8, and the reason is that `is_match` applies its rules in order and the
**size** gate rejects first:
```
Nestle Munch 38.5 g <- OFF "Munch" blocked by SIZE (off quantity = None)
Coca-Cola Maaza 750ml <- OFF "Maaza" blocked by SIZE (off quantity = None)
Cadbury Perk 22 g <- OFF "Perk" blocked by SIZE (off quantity = None)
```
`size_matches` returns False whenever *either* side is blank, and Open Food
Facts leaves `quantity` null on a large share of records — 57 of 146 Amul hits.
`off_bulk` had already documented this and worked around it by treating size as
a ranking bonus rather than a veto.
So the flag is `barcode_is_identity`, not `allow_containment`, because it
describes the precondition rather than one of its consequences: **the caller
already knows which product this is, because it fetched by GTIN.** Under it,
name and size stop being evidence of identity and become sanity checks against
our barcode being on the wrong row — and a sanity check cannot fail on
information the source does not have:
| Rule | Normally | Under `barcode_is_identity` |
|---|---|---|
| 1. brand | must match | unchanged |
| 2. size | blank ⇒ reject | **blank ⇒ no information, allowed.** Present-and-different still rejects |
| 3. variant terms | must not conflict | unchanged |
| 4. name similarity | ≥ floor | floor, **or** containment |
**Defaults off.** On the search path many candidates compete and name and size
are the only things telling them apart — "Munch" with no size would match every
Nestle product containing that word. Pass it only where a single candidate was
fetched by barcode: `fetch_verified_nutrition_by_barcode` and
`scripts/backfill_nutrition_from_barcodes`, and nothing else today.
---
## What a barcode is allowed to claim
`barcode_verified = true` means brand, size **and** name were matched against a
source record. It is not a synonym for "we have a barcode".
| `barcode_lookup_status` | Meaning |
|---|---|
| `verified` | the cascade matched brand + size + name |
| `name_matched` | bulk corpus match ≥ 0.88; the pack is **not** confirmed |
| `sheet_validated` | the merchant supplied it and it passes the GTIN checksum |
| `not_found` / `error` / `disabled` | no barcode stored |
A `name_matched` code is a real GS1 barcode for *that brand*, written to every
size variant of a title. Good enough for catalog matching, dedup and nutrition
lookups. **Not** good enough for logistics, invoicing, or anything a scanner
drives. Filter on `barcode_verified = false` to select, correct or revert them.
---
## Validation is not optional and not per-source
`validators.validate_barcode` is the single gate: digits only → a legal GTIN
length (8/12/13/14) → recomputed check digit. A source saying "this is the
barcode" is never sufficient. Trust affects the *order* sources are tried, never
whether validation runs.
Derived fields follow from the digits alone:
- `gtin` — the validated code
- `ean13` — a UPC-A zero-padded to 13. **GTIN-8 is not padded**: an 8-digit GTIN
is its own symbology, not a truncated EAN-13.
- `upc` — 12-digit codes only
> **`upc` will always be 0% in this catalog, and that is correct.** Every code
> here is GS1 India (prefix `890`), which issues EAN-13 and GTIN-8. UPC-A is
> North American. Measured: 0 of 300. The coverage report counts it *not
> applicable*, not missing.
---
## Where the fields go
A field computed by a stage reaches Postgres only if it is named in **all three**
of these. Two of them were missing the barcode fields for a long time, which is
why `upc` sat at 0% and `gtin` at 8.7% while the code that produced them ran on
every ingestion:
1. `store_catalog_pipeline._to_storage_row` — the projection
2. `vector_store.upsert_brand_products` — the INSERT column list
3. `vector_store._ensure_columns` — the only migration mechanism (no Alembic)
Plus `brand_sync.EXPORT_COLUMNS` for the seed-file round trip.
Every enrichment column uses `COALESCE(EXCLUDED.x, table.x)` in the
`ON CONFLICT` clause, including `barcode` itself. Without it a bare re-seed —
which carries no enrichment keys — sets them to NULL. Proven against a live
table: it kept `gtin` and blanked `barcode`, leaving a row claiming a GTIN with
no barcode.
---
## Running it
```bash
# What is filled, and where each value came from
python -m scripts.catalog_coverage --by-provenance
# Derive gtin/ean13/upc/barcode_type from barcodes already stored (offline)
python -m scripts.backfill_barcode_identity --apply
# Bulk-match barcodes from the OFF brand corpora
python -m scripts.backfill_barcodes_from_off --apply
# Upgrade nutrition from a name match to an exact barcode match
python -m scripts.backfill_nutrition_from_barcodes --apply
```
Every script is **dry-run by default** and prints its target database on
startup. `backend/.env` points at production.