# Barcode enrichment `app/services/enrichment/barcode/` — how a catalog row acquires a barcode, what the barcode is then allowed to claim, and which knob to turn. Referenced from `.env.example` and `app/services/enrichment/pipeline.py`, both of which pointed at this file for a long time before it existed. --- ## The four ways a row gets a barcode | # | Path | When | Cost | Confidence | |---|---|---|---|---| | 1 | **The sheet** | The merchant typed it | free | highest — they are holding the pack | | 2 | **Identity derivation** | Always, inline | free, offline | derives `gtin`/`ean13`/`upc`/`barcode_type` from a barcode already present | | 3 | **Bulk brand corpus** | After upload, per brand | ~5 requests **per brand** | name-matched at ≥ 0.88 | | 4 | **Per-product cascade** | Only if enabled | 1 request **per product**, 10/min cap | brand + size + name matched | Paths 1 and 2 are always on. Path 3 is on by default. Path 4 is off by default. ### Why path 4 is off and path 3 is on They differ in cost, not in appetite for risk. The Open Food Facts per-product search endpoint is capped at **10 requests per minute**. A 200-row upload on path 4 is twenty minutes of a held HTTP request, on a shared 8 GB host — which is what `settings.py:420-423` is about. Path 3 fetches a brand's *entire* India catalogue in about five requests, caches it to `data/cache/off_brand_corpus/.json`, and matches offline. A brand costs the same whether it has 3 rows or 300. ``` ENABLE_BARCODE_LOOKUP=false # path 4: per product, inline ENRICH_BARCODES_ON_UPLOAD=true # path 3: per brand, after the upload settles ``` > `.env.example` shipped `ENABLE_BARCODE_LOOKUP=true` while `settings.py` > defaulted it `false`, for as long as both existed. Anyone copying the example > got a materially different pipeline from anyone relying on defaults. Fixed; > if you see the two disagree again, `settings.py` is authoritative. --- ## Two thresholds, and why they are different numbers | Constant | Value | Direction | Used by | |---|---|---|---| | `BARCODE_MIN_NAME_SIMILARITY` | 0.78 | reverse — barcode known, name is a sanity check | `fetch_verified_nutrition_by_barcode` | | `OFF bulk review_min` | 0.88 | forward — name carries the whole decision | `post_ingest_barcodes`, `backfill_barcodes_from_off` | `settings.py:449-477` records the measurement behind 0.78: at 0.45 the cascade accepted 15 candidates of which 8 were wrong; at 0.78 it accepted 2 and none were wrong. **Do not lower it to fix a coverage complaint.** ### The containment rule Open Food Facts stores short names. We store long ones. Measured over the catalog on 2026-09-08, 149 of 300 barcoded rows were refused as "found, wrong product" when the barcode had resolved perfectly: ``` ours "Nestle Munch 8.9g" OFF "Munch" similarity 0.332 ours "Coca-Cola Maaza 750ml" OFF "Maaza" similarity 0.304 ``` `name_similarity` divides overlap by *our* token count, so a one-token candidate cannot exceed ~0.33 however right it is. But the same run correctly refused: ``` ours "Pepsico Lays 1kg" OFF "Spanish tomato tango" 0.133 ours "Coca-Cola Fanta 750ml" OFF "Orange" 0.089 ``` A threshold low enough to admit the first group admits the second. So `matching.name_is_contained` separates them **structurally**: every candidate token must be one of ours once brand and size tokens are discounted, and a bare brand name ("Colgate", "godrej" — both real OFF titles) never matches. ### `barcode_is_identity` — and the gate that actually blocked most of them The name gate was the *visible* symptom. When the fix was measured it moved only 3 rows to 8, and the reason is that `is_match` applies its rules in order and the **size** gate rejects first: ``` Nestle Munch 38.5 g <- OFF "Munch" blocked by SIZE (off quantity = None) Coca-Cola Maaza 750ml <- OFF "Maaza" blocked by SIZE (off quantity = None) Cadbury Perk 22 g <- OFF "Perk" blocked by SIZE (off quantity = None) ``` `size_matches` returns False whenever *either* side is blank, and Open Food Facts leaves `quantity` null on a large share of records — 57 of 146 Amul hits. `off_bulk` had already documented this and worked around it by treating size as a ranking bonus rather than a veto. So the flag is `barcode_is_identity`, not `allow_containment`, because it describes the precondition rather than one of its consequences: **the caller already knows which product this is, because it fetched by GTIN.** Under it, name and size stop being evidence of identity and become sanity checks against our barcode being on the wrong row — and a sanity check cannot fail on information the source does not have: | Rule | Normally | Under `barcode_is_identity` | |---|---|---| | 1. brand | must match | unchanged | | 2. size | blank ⇒ reject | **blank ⇒ no information, allowed.** Present-and-different still rejects | | 3. variant terms | must not conflict | unchanged | | 4. name similarity | ≥ floor | floor, **or** containment | **Defaults off.** On the search path many candidates compete and name and size are the only things telling them apart — "Munch" with no size would match every Nestle product containing that word. Pass it only where a single candidate was fetched by barcode: `fetch_verified_nutrition_by_barcode` and `scripts/backfill_nutrition_from_barcodes`, and nothing else today. --- ## What a barcode is allowed to claim `barcode_verified = true` means brand, size **and** name were matched against a source record. It is not a synonym for "we have a barcode". | `barcode_lookup_status` | Meaning | |---|---| | `verified` | the cascade matched brand + size + name | | `name_matched` | bulk corpus match ≥ 0.88; the pack is **not** confirmed | | `sheet_validated` | the merchant supplied it and it passes the GTIN checksum | | `not_found` / `error` / `disabled` | no barcode stored | A `name_matched` code is a real GS1 barcode for *that brand*, written to every size variant of a title. Good enough for catalog matching, dedup and nutrition lookups. **Not** good enough for logistics, invoicing, or anything a scanner drives. Filter on `barcode_verified = false` to select, correct or revert them. --- ## Validation is not optional and not per-source `validators.validate_barcode` is the single gate: digits only → a legal GTIN length (8/12/13/14) → recomputed check digit. A source saying "this is the barcode" is never sufficient. Trust affects the *order* sources are tried, never whether validation runs. Derived fields follow from the digits alone: - `gtin` — the validated code - `ean13` — a UPC-A zero-padded to 13. **GTIN-8 is not padded**: an 8-digit GTIN is its own symbology, not a truncated EAN-13. - `upc` — 12-digit codes only > **`upc` will always be 0% in this catalog, and that is correct.** Every code > here is GS1 India (prefix `890`), which issues EAN-13 and GTIN-8. UPC-A is > North American. Measured: 0 of 300. The coverage report counts it *not > applicable*, not missing. --- ## Where the fields go A field computed by a stage reaches Postgres only if it is named in **all three** of these. Two of them were missing the barcode fields for a long time, which is why `upc` sat at 0% and `gtin` at 8.7% while the code that produced them ran on every ingestion: 1. `store_catalog_pipeline._to_storage_row` — the projection 2. `vector_store.upsert_brand_products` — the INSERT column list 3. `vector_store._ensure_columns` — the only migration mechanism (no Alembic) Plus `brand_sync.EXPORT_COLUMNS` for the seed-file round trip. Every enrichment column uses `COALESCE(EXCLUDED.x, table.x)` in the `ON CONFLICT` clause, including `barcode` itself. Without it a bare re-seed — which carries no enrichment keys — sets them to NULL. Proven against a live table: it kept `gtin` and blanked `barcode`, leaving a row claiming a GTIN with no barcode. --- ## Running it ```bash # What is filled, and where each value came from python -m scripts.catalog_coverage --by-provenance # Derive gtin/ean13/upc/barcode_type from barcodes already stored (offline) python -m scripts.backfill_barcode_identity --apply # Bulk-match barcodes from the OFF brand corpora python -m scripts.backfill_barcodes_from_off --apply # Upgrade nutrition from a name match to an exact barcode match python -m scripts.backfill_nutrition_from_barcodes --apply ``` Every script is **dry-run by default** and prints its target database on startup. `backend/.env` points at production.