8.4 KiB
Barcode enrichment
app/services/enrichment/barcode/ — how a catalog row acquires a barcode, what
the barcode is then allowed to claim, and which knob to turn.
Referenced from .env.example and app/services/enrichment/pipeline.py, both
of which pointed at this file for a long time before it existed.
The four ways a row gets a barcode
| # | Path | When | Cost | Confidence |
|---|---|---|---|---|
| 1 | The sheet | The merchant typed it | free | highest — they are holding the pack |
| 2 | Identity derivation | Always, inline | free, offline | derives gtin/ean13/upc/barcode_type from a barcode already present |
| 3 | Bulk brand corpus | After upload, per brand | ~5 requests per brand | name-matched at ≥ 0.88 |
| 4 | Per-product cascade | Only if enabled | 1 request per product, 10/min cap | brand + size + name matched |
Paths 1 and 2 are always on. Path 3 is on by default. Path 4 is off by default.
Why path 4 is off and path 3 is on
They differ in cost, not in appetite for risk.
The Open Food Facts per-product search endpoint is capped at 10 requests per
minute. A 200-row upload on path 4 is twenty minutes of a held HTTP request,
on a shared 8 GB host — which is what settings.py:420-423 is about.
Path 3 fetches a brand's entire India catalogue in about five requests, caches
it to data/cache/off_brand_corpus/<slug>.json, and matches offline. A brand
costs the same whether it has 3 rows or 300.
ENABLE_BARCODE_LOOKUP=false # path 4: per product, inline
ENRICH_BARCODES_ON_UPLOAD=true # path 3: per brand, after the upload settles
.env.exampleshippedENABLE_BARCODE_LOOKUP=truewhilesettings.pydefaulted itfalse, for as long as both existed. Anyone copying the example got a materially different pipeline from anyone relying on defaults. Fixed; if you see the two disagree again,settings.pyis authoritative.
Two thresholds, and why they are different numbers
| Constant | Value | Direction | Used by |
|---|---|---|---|
BARCODE_MIN_NAME_SIMILARITY |
0.78 | reverse — barcode known, name is a sanity check | fetch_verified_nutrition_by_barcode |
OFF bulk review_min |
0.88 | forward — name carries the whole decision | post_ingest_barcodes, backfill_barcodes_from_off |
settings.py:449-477 records the measurement behind 0.78: at 0.45 the cascade
accepted 15 candidates of which 8 were wrong; at 0.78 it accepted 2 and none
were wrong. Do not lower it to fix a coverage complaint.
The containment rule
Open Food Facts stores short names. We store long ones. Measured over the catalog on 2026-09-08, 149 of 300 barcoded rows were refused as "found, wrong product" when the barcode had resolved perfectly:
ours "Nestle Munch 8.9g" OFF "Munch" similarity 0.332
ours "Coca-Cola Maaza 750ml" OFF "Maaza" similarity 0.304
name_similarity divides overlap by our token count, so a one-token candidate
cannot exceed ~0.33 however right it is. But the same run correctly refused:
ours "Pepsico Lays 1kg" OFF "Spanish tomato tango" 0.133
ours "Coca-Cola Fanta 750ml" OFF "Orange" 0.089
A threshold low enough to admit the first group admits the second. So
matching.name_is_contained separates them structurally: every candidate
token must be one of ours once brand and size tokens are discounted, and a bare
brand name ("Colgate", "godrej" — both real OFF titles) never matches.
barcode_is_identity — and the gate that actually blocked most of them
The name gate was the visible symptom. When the fix was measured it moved only
3 rows to 8, and the reason is that is_match applies its rules in order and the
size gate rejects first:
Nestle Munch 38.5 g <- OFF "Munch" blocked by SIZE (off quantity = None)
Coca-Cola Maaza 750ml <- OFF "Maaza" blocked by SIZE (off quantity = None)
Cadbury Perk 22 g <- OFF "Perk" blocked by SIZE (off quantity = None)
size_matches returns False whenever either side is blank, and Open Food
Facts leaves quantity null on a large share of records — 57 of 146 Amul hits.
off_bulk had already documented this and worked around it by treating size as
a ranking bonus rather than a veto.
So the flag is barcode_is_identity, not allow_containment, because it
describes the precondition rather than one of its consequences: the caller
already knows which product this is, because it fetched by GTIN. Under it,
name and size stop being evidence of identity and become sanity checks against
our barcode being on the wrong row — and a sanity check cannot fail on
information the source does not have:
| Rule | Normally | Under barcode_is_identity |
|---|---|---|
| 1. brand | must match | unchanged |
| 2. size | blank ⇒ reject | blank ⇒ no information, allowed. Present-and-different still rejects |
| 3. variant terms | must not conflict | unchanged |
| 4. name similarity | ≥ floor | floor, or containment |
Defaults off. On the search path many candidates compete and name and size
are the only things telling them apart — "Munch" with no size would match every
Nestle product containing that word. Pass it only where a single candidate was
fetched by barcode: fetch_verified_nutrition_by_barcode and
scripts/backfill_nutrition_from_barcodes, and nothing else today.
What a barcode is allowed to claim
barcode_verified = true means brand, size and name were matched against a
source record. It is not a synonym for "we have a barcode".
barcode_lookup_status |
Meaning |
|---|---|
verified |
the cascade matched brand + size + name |
name_matched |
bulk corpus match ≥ 0.88; the pack is not confirmed |
sheet_validated |
the merchant supplied it and it passes the GTIN checksum |
not_found / error / disabled |
no barcode stored |
A name_matched code is a real GS1 barcode for that brand, written to every
size variant of a title. Good enough for catalog matching, dedup and nutrition
lookups. Not good enough for logistics, invoicing, or anything a scanner
drives. Filter on barcode_verified = false to select, correct or revert them.
Validation is not optional and not per-source
validators.validate_barcode is the single gate: digits only → a legal GTIN
length (8/12/13/14) → recomputed check digit. A source saying "this is the
barcode" is never sufficient. Trust affects the order sources are tried, never
whether validation runs.
Derived fields follow from the digits alone:
gtin— the validated codeean13— a UPC-A zero-padded to 13. GTIN-8 is not padded: an 8-digit GTIN is its own symbology, not a truncated EAN-13.upc— 12-digit codes only
upcwill always be 0% in this catalog, and that is correct. Every code here is GS1 India (prefix890), which issues EAN-13 and GTIN-8. UPC-A is North American. Measured: 0 of 300. The coverage report counts it not applicable, not missing.
Where the fields go
A field computed by a stage reaches Postgres only if it is named in all three
of these. Two of them were missing the barcode fields for a long time, which is
why upc sat at 0% and gtin at 8.7% while the code that produced them ran on
every ingestion:
store_catalog_pipeline._to_storage_row— the projectionvector_store.upsert_brand_products— the INSERT column listvector_store._ensure_columns— the only migration mechanism (no Alembic)
Plus brand_sync.EXPORT_COLUMNS for the seed-file round trip.
Every enrichment column uses COALESCE(EXCLUDED.x, table.x) in the
ON CONFLICT clause, including barcode itself. Without it a bare re-seed —
which carries no enrichment keys — sets them to NULL. Proven against a live
table: it kept gtin and blanked barcode, leaving a row claiming a GTIN with
no barcode.
Running it
# What is filled, and where each value came from
python -m scripts.catalog_coverage --by-provenance
# Derive gtin/ean13/upc/barcode_type from barcodes already stored (offline)
python -m scripts.backfill_barcode_identity --apply
# Bulk-match barcodes from the OFF brand corpora
python -m scripts.backfill_barcodes_from_off --apply
# Upgrade nutrition from a name match to an exact barcode match
python -m scripts.backfill_nutrition_from_barcodes --apply
Every script is dry-run by default and prints its target database on
startup. backend/.env points at production.