Files
catalogue_backend/docs/BARCODE_ENRICHMENT.md
2026-09-08 15:18:29 +05:30

8.4 KiB

Barcode enrichment

app/services/enrichment/barcode/ — how a catalog row acquires a barcode, what the barcode is then allowed to claim, and which knob to turn.

Referenced from .env.example and app/services/enrichment/pipeline.py, both of which pointed at this file for a long time before it existed.


The four ways a row gets a barcode

# Path When Cost Confidence
1 The sheet The merchant typed it free highest — they are holding the pack
2 Identity derivation Always, inline free, offline derives gtin/ean13/upc/barcode_type from a barcode already present
3 Bulk brand corpus After upload, per brand ~5 requests per brand name-matched at ≥ 0.88
4 Per-product cascade Only if enabled 1 request per product, 10/min cap brand + size + name matched

Paths 1 and 2 are always on. Path 3 is on by default. Path 4 is off by default.

Why path 4 is off and path 3 is on

They differ in cost, not in appetite for risk.

The Open Food Facts per-product search endpoint is capped at 10 requests per minute. A 200-row upload on path 4 is twenty minutes of a held HTTP request, on a shared 8 GB host — which is what settings.py:420-423 is about.

Path 3 fetches a brand's entire India catalogue in about five requests, caches it to data/cache/off_brand_corpus/<slug>.json, and matches offline. A brand costs the same whether it has 3 rows or 300.

ENABLE_BARCODE_LOOKUP=false        # path 4: per product, inline
ENRICH_BARCODES_ON_UPLOAD=true     # path 3: per brand, after the upload settles

.env.example shipped ENABLE_BARCODE_LOOKUP=true while settings.py defaulted it false, for as long as both existed. Anyone copying the example got a materially different pipeline from anyone relying on defaults. Fixed; if you see the two disagree again, settings.py is authoritative.


Two thresholds, and why they are different numbers

Constant Value Direction Used by
BARCODE_MIN_NAME_SIMILARITY 0.78 reverse — barcode known, name is a sanity check fetch_verified_nutrition_by_barcode
OFF bulk review_min 0.88 forward — name carries the whole decision post_ingest_barcodes, backfill_barcodes_from_off

settings.py:449-477 records the measurement behind 0.78: at 0.45 the cascade accepted 15 candidates of which 8 were wrong; at 0.78 it accepted 2 and none were wrong. Do not lower it to fix a coverage complaint.

The containment rule

Open Food Facts stores short names. We store long ones. Measured over the catalog on 2026-09-08, 149 of 300 barcoded rows were refused as "found, wrong product" when the barcode had resolved perfectly:

ours "Nestle Munch 8.9g"      OFF "Munch"    similarity 0.332
ours "Coca-Cola Maaza 750ml"  OFF "Maaza"    similarity 0.304

name_similarity divides overlap by our token count, so a one-token candidate cannot exceed ~0.33 however right it is. But the same run correctly refused:

ours "Pepsico Lays 1kg"       OFF "Spanish tomato tango"  0.133
ours "Coca-Cola Fanta 750ml"  OFF "Orange"                0.089

A threshold low enough to admit the first group admits the second. So matching.name_is_contained separates them structurally: every candidate token must be one of ours once brand and size tokens are discounted, and a bare brand name ("Colgate", "godrej" — both real OFF titles) never matches.

barcode_is_identity — and the gate that actually blocked most of them

The name gate was the visible symptom. When the fix was measured it moved only 3 rows to 8, and the reason is that is_match applies its rules in order and the size gate rejects first:

Nestle Munch 38.5 g   <- OFF "Munch"   blocked by SIZE (off quantity = None)
Coca-Cola Maaza 750ml <- OFF "Maaza"   blocked by SIZE (off quantity = None)
Cadbury Perk 22 g     <- OFF "Perk"    blocked by SIZE (off quantity = None)

size_matches returns False whenever either side is blank, and Open Food Facts leaves quantity null on a large share of records — 57 of 146 Amul hits. off_bulk had already documented this and worked around it by treating size as a ranking bonus rather than a veto.

So the flag is barcode_is_identity, not allow_containment, because it describes the precondition rather than one of its consequences: the caller already knows which product this is, because it fetched by GTIN. Under it, name and size stop being evidence of identity and become sanity checks against our barcode being on the wrong row — and a sanity check cannot fail on information the source does not have:

Rule Normally Under barcode_is_identity
1. brand must match unchanged
2. size blank ⇒ reject blank ⇒ no information, allowed. Present-and-different still rejects
3. variant terms must not conflict unchanged
4. name similarity ≥ floor floor, or containment

Defaults off. On the search path many candidates compete and name and size are the only things telling them apart — "Munch" with no size would match every Nestle product containing that word. Pass it only where a single candidate was fetched by barcode: fetch_verified_nutrition_by_barcode and scripts/backfill_nutrition_from_barcodes, and nothing else today.


What a barcode is allowed to claim

barcode_verified = true means brand, size and name were matched against a source record. It is not a synonym for "we have a barcode".

barcode_lookup_status Meaning
verified the cascade matched brand + size + name
name_matched bulk corpus match ≥ 0.88; the pack is not confirmed
sheet_validated the merchant supplied it and it passes the GTIN checksum
not_found / error / disabled no barcode stored

A name_matched code is a real GS1 barcode for that brand, written to every size variant of a title. Good enough for catalog matching, dedup and nutrition lookups. Not good enough for logistics, invoicing, or anything a scanner drives. Filter on barcode_verified = false to select, correct or revert them.


Validation is not optional and not per-source

validators.validate_barcode is the single gate: digits only → a legal GTIN length (8/12/13/14) → recomputed check digit. A source saying "this is the barcode" is never sufficient. Trust affects the order sources are tried, never whether validation runs.

Derived fields follow from the digits alone:

  • gtin — the validated code
  • ean13 — a UPC-A zero-padded to 13. GTIN-8 is not padded: an 8-digit GTIN is its own symbology, not a truncated EAN-13.
  • upc — 12-digit codes only

upc will always be 0% in this catalog, and that is correct. Every code here is GS1 India (prefix 890), which issues EAN-13 and GTIN-8. UPC-A is North American. Measured: 0 of 300. The coverage report counts it not applicable, not missing.


Where the fields go

A field computed by a stage reaches Postgres only if it is named in all three of these. Two of them were missing the barcode fields for a long time, which is why upc sat at 0% and gtin at 8.7% while the code that produced them ran on every ingestion:

  1. store_catalog_pipeline._to_storage_row — the projection
  2. vector_store.upsert_brand_products — the INSERT column list
  3. vector_store._ensure_columns — the only migration mechanism (no Alembic)

Plus brand_sync.EXPORT_COLUMNS for the seed-file round trip.

Every enrichment column uses COALESCE(EXCLUDED.x, table.x) in the ON CONFLICT clause, including barcode itself. Without it a bare re-seed — which carries no enrichment keys — sets them to NULL. Proven against a live table: it kept gtin and blanked barcode, leaving a row claiming a GTIN with no barcode.


Running it

# What is filled, and where each value came from
python -m scripts.catalog_coverage --by-provenance

# Derive gtin/ean13/upc/barcode_type from barcodes already stored (offline)
python -m scripts.backfill_barcode_identity --apply

# Bulk-match barcodes from the OFF brand corpora
python -m scripts.backfill_barcodes_from_off --apply

# Upgrade nutrition from a name match to an exact barcode match
python -m scripts.backfill_nutrition_from_barcodes --apply

Every script is dry-run by default and prints its target database on startup. backend/.env points at production.