Backend catalog recent updates

This commit is contained in:
sriram
2026-09-01 18:02:11 +05:30
parent 68bff007ea
commit 7b3fbc47b4
11 changed files with 1457 additions and 179 deletions

View File

@@ -3,26 +3,28 @@
Reply to the seven findings dated 31 August 2026 (drop
`8e1448e176d843d08183d387ad724f95` → run `0ed4c77b0ca14e03b1aaf6b5b77d1994`).
Thank you for this. It is the most useful bug report this project has had, and
two of the seven led us to faults we had not found ourselves — one of them
affecting rows well outside the ones you named.
Every figure below was measured against the same live deployment, not read off
source. Where we disagree with a finding, the evidence is included so you can
check it rather than take our word for it.
**Summary:** four items are fixed and ship in the next backend deploy. One
(#02) was already in the API and we had failed to document it — that is our
fault and the docs are now corrected. #01 is diagnosed, and the cause is not
what either of us assumed. #06 is confirmed but carries a trap that means we
should agree an approach before touching it.
| # | Finding | Status |
| --- | --- | --- |
| 01 | Pack sizes replaced between scrapes | **Diagnosed** — three causes, not one. Fix needs your input |
| 02 | `rejected` is a bare count | **Already shipped, now documented** — plus `row` added |
| 01 | Pack sizes replaced between scrapes | **Diagnosed** — three causes, not one. One is fixed; the other two need your input |
| 02 | `rejected` is a bare count | **Was already in the API — our docs failed you.** `row` added |
| 03 | Manifest brands are not catalogue keys | **Fixed** — `brand_key` published |
| 04 | Run files carry no `from_drop` | **Fixed** |
| 05 | Manifest carries no `source_row` | **Fixed** |
| 06 | Duplicate brands and products | **Confirmed.** Read the trap below before we act |
| 06 | Duplicate brands and products | **Brand merged. Stray numbers diagnosed — there are 18, not 1** |
| 07 | No loose-produce coverage | **Fixed** — 159-row base list, and the upload path now handles produce |
> **Everything below ships in the next backend deploy and is not live yet.** The
> only change already applied to the live database is the Haldiram merge in #06.
> We will confirm when the deploy lands; please do not re-test until then.
---
## 01 — Pack sizes and names are replaced between scrapes
@@ -31,11 +33,12 @@ You asked us to confirm whether a pack size that once existed is meant to
survive a re-scrape. **It is not, today** — but that is only the last of three
causes, and fixing it alone would not have saved your links.
### Cause 1: pack sizes are invented when a scrape does not supply them
### Cause 1: the pack sizes were never scraped facts. Some were generated.
`_sizes_for()` falls back to `default_size_variants(category, name)` when a
product declares no size. That fallback is **keyed on the resolved category**,
and the resolved category is not stable between runs. Measured today:
When a product declares no size, `_sizes_for()` falls back to
`default_size_variants(category, name)`. That fallback is **keyed on the
resolved category**, and the resolved category is not stable between runs.
Measured today:
```
category "Snacks" -> ['55g', '150g', '200g']
@@ -46,18 +49,25 @@ category "Breakfast Cereal" -> ['250g', '500g', '1kg']
Now compare your table. Cheetos id 25 is **55g** — the *Snacks* set. Ids 26 and
27 are **100g** and **250g** — the *unresolved* set. The same product was
ingested once with its category resolved and once without, and produced two
disjoint sets of pack sizes. Your Cheerios ids (100g, 250g, 500g) sit across the
Breakfast Cereal set and the unresolved set the same way.
disjoint sets of pack sizes. Your Cheerios ids sit across the Breakfast Cereal
set and the unresolved set the same way.
So the pack sizes were never scraped facts that changed. Some of them were
generated, and the generator's input moved.
So a 100 g bag and a 250 g bag are indeed different SKUs, and you are right not
to re-point one at the other — but in these cases neither number came off a
pack. They were both guesses, from two different guesses about the category.
**This one is now fixed for the class of product where it does most damage.**
Unbranded and loose goods no longer receive invented sizes at all (see #07). For
branded packaged goods the fallback still runs, because a brand really does sell
a small/medium/large range and omitting it entirely loses more than it saves.
That is the part we would like your view on — see *What we need from you*.
### Cause 2: the product name gains or loses a brand prefix
Your own #06 has the evidence: `Hot Heads 30g` became `Nestle Hot Heads`, and
`PepsiCo Kurkure Masala Munch 90g` coexists with `Kurkure Masala Munch 90g`.
`image_id` is derived from the name, so a prefix appearing or disappearing moves
the id even when the product is identical.
the id even when the product is identical. Still open; see #06.
### Cause 3: the write then deletes whatever is not in the new set
@@ -71,54 +81,55 @@ re-scrape**, which matches your finding that zero of eleven could be repaired.
### Not the upload path
Worth stating plainly, because it affects how much you need to worry: the
Worth stating plainly, because it changes how much you need to worry: the
**upload** path — everything reached through `POST /api/uploads/catalog` — uses
`cleanup=False` and has always done so. A sheet you send can never delete a row
it does not mention. The deletions came from brand scraping only.
### What we need from you
The real fix is to stop causes 1 and 2 (do not invent sizes for a product
already in the catalogue; settle the naming convention), and to soft-retire
rather than delete for cause 3. That third part changes how the production
catalogue is written and we would rather agree it with you than spring it:
- Would a `retired_at` timestamp plus exclusion from the default read work for
you, instead of the row being deleted? That preserves the `image_id` so your
stored link resolves to something, and lets us give you the `superseded_by`
and per-run changelog you asked for.
- If so, do you want retired rows visible through an explicit query, or gone
from the API entirely?
`cleanup=False` and always has. A sheet you send can never delete a row it does
not mention. The deletions came from brand scraping only.
### Your two direct questions
**Is `image_id` stable across re-scrapes for a product whose name and pack size
have not changed?** Yes. It is a pure deterministic function of brand, product
name and pack size, with no clock, counter or run id in it. Verified:
name and pack size, with no clock, counter or run id in it:
```
build_image_id('pepsico', 'Cheetos Chips', '100g') -> pepsico_cheetos_chips_100g
```
Storing it rather than our row id is the right call and we have documented the
guarantee so it does not quietly change. **The caveat is #06:** the guarantee is
only as good as the stability of the name, and inconsistent brand prefixing
breaks exactly that.
Storing it rather than our row id is the right call, and we have now documented
the guarantee in `INGESTION_API.md` so it cannot quietly change. **The caveat is
#06:** the guarantee is only as good as the stability of the name, and
inconsistent brand prefixing breaks exactly that.
**Do you want to know when a product is dropped or renamed?** Yes, and we agree
it should exist. It falls out of the retirement model above rather than being a
separate feature, which is why we would like to settle that first.
it should exist. It falls out of the retirement model below rather than being a
separate feature.
### What we need from you
- Would a `retired_at` timestamp plus exclusion from the default read work
instead of the row being deleted? That preserves the `image_id` so your stored
link resolves to *something*, and gives us somewhere to hang the
`superseded_by` and per-run changelog you asked for.
- If so: should retired rows be reachable through an explicit query, or absent
from the API entirely?
- On cause 1: would you rather we **stopped inventing sizes altogether** for
branded goods too? It would shrink the catalogue and lose some genuine
variants, but every remaining row would be a size somebody actually saw on a
pack. We can go either way and would rather match how you consume it.
---
## 02 — `rejected` is a count with no reason
**This is our documentation failure, not a missing feature.** `rejections[]` has
been in every response — single-batch read and list endpoint both, since
`to_out(slim=True)` strips only `products` — carrying `product_name`, `size` and
`reason` per refused row. It was absent from `INGESTION_API.md`, which documents
`"rejected": 0` and never mentions the array, so there was no way for you to
know it was there. Sorry — that is a straightforwardly bad docs bug.
**This one is our documentation failing you, not a missing feature, and we are
sorry for the time it cost.** `rejections[]` has been in every response — the
single-batch read and the list endpoint both, since `to_out(slim=True)` strips
only `products` — carrying `product_name`, `size` and `reason` per refused row.
It was absent from `INGESTION_API.md`, which documents `"rejected": 0` and never
mentions the array, so there was no way for you to know it was there. You were
diffing 19 against 17 because our docs told you that was all you had.
The genuine gap was the row number, which is now added:
@@ -133,6 +144,8 @@ The genuine gap was the row number, which is now added:
convention as the `422` responses, so it matches what the operator sees on
screen. `null` only when the row cannot be located. Capped at 50 per file.
The array is now documented, with a field table.
---
## 03 — Manifest brands are not catalogue keys
@@ -143,13 +156,13 @@ Fixed. Every entry in `products[]` now carries `brand_key` beside `brand`:
{ "brand": "24 Mantra", "brand_key": "24_mantra", ... }
```
This is generated by the same function the storage layer uses to name the table,
It is generated by the same function the storage layer uses to name the table,
so it cannot drift from the key the catalogue is actually addressed by. Your
normalisation is correct as far as we can tell, but it is a guess, and the
failure mode is silent — a wrong key finds nothing rather than erroring.
Thank you for degrading rather than failing the batch on an unreadable brand;
that is the right behaviour and we should have done it on our side too.
Thank you for making a single unreadable brand degrade rather than fail the
batch. That is the right behaviour and we should have done it on our side too.
---
@@ -169,8 +182,8 @@ from — the exact inverse of `released_to`:
```
`null` for a file that went straight into a run without sitting in an inbox —
which, under `UPLOAD_AUTORUN=true`, is every file you send, because the id you
are handed is already the run.
which, under the current `UPLOAD_AUTORUN=true`, is every file you send, because
the id you are handed is already the run.
There is a test in our suite that stages two drops from different senders both
named `products.csv` and asserts they are distinguishable, so the collision you
@@ -193,43 +206,86 @@ set difference rather than a name-matching heuristic.
## 06 — Duplicate brands, duplicate products, stray names
Confirmed against the live database today: **55 brand tables, 1,414 products**
(you counted 1,614; the difference is a day of drift plus, we think, your count
including rejected rows — worth reconciling if it matters).
### The brand split — merged
`haldiram` (1 product) is gone; `haldirams` (2) is the survivor.
The split was not a typo. **Nothing in the system knew the two spellings were
one brand** — neither was in the alias map, so each resolved to itself and every
upload built whichever table its sheet happened to name. Merging the rows alone
would have fixed nothing: the next sheet spelling it without the "s" would
rebuild the table. The alias is in, so `Haldiram`, `haldiram`, `HALDIRAM` and
`Haldiram's` all now resolve to `haldirams`.
While merging we found the singular row was a corrupted duplicate of one already
in the plural table — same product, but with a stray `45` in the name, a raw
category id, and a bare-number size. We kept the clean row and carried across
the one thing the corrupted row had that it lacked: a newer price.
We also found, and corrected, something you could not have seen: **both
`haldirams` rows were carrying Lion Dates' FSSAI licence** (`10012042000244`,
the number on all 21 Lion Dates products) rather than Haldiram's own
(`10012011000140`). That is a regulatory identifier on the wrong manufacturer's
product, and it is fixed in the database and the seed file.
### The stray number — there are 18 of them, and we know what it is
You found `Lays Classic Salted 52g 150` and asked us to check for the pattern
elsewhere. **It affects 18 products across at least seven brands:**
```
haldiram 1 britannia 6
haldirams 2 parle 3 against hindustan_unilever 443
patanjali 3
Aashirvaad Shudh Chakki Atta 5kg 40 size_variants ['40'] price 299
Lays Classic Salted 52g 150 size_variants ['150'] price 21
Britannia Good Day Cashew 200g 60 size_variants ['60'] price 52
Coca-Cola 750ml 72 size_variants ['72'] price 42
Dove Cream Beauty Bar 100g 64 size_variants ['64'] price 76
Horlicks Classic Malt 500g 18 size_variants ['18'] price 289
... 12 more
```
So: the split brand is real, and Britannia/Parle/Patanjali do look like scrapes
that stopped part-way rather than genuinely small brands. We will re-run those
three.
The number is **not** a price — 40 against ₹299, 150 against ₹21. It is the
**case-pack count**: how many units come in a carton. A sheet's "Quantity"
column was mapped to the pack-size field, the bare number became the size, and
`_to_storage_row` then appended it to the product name.
### The trap, which is why we have not just fixed this
**The cause is already fixed**, on two layers, both verified today:
**De-duplicating the PepsiCo pairs means renaming a product, and `image_id` is
derived from the name.** Renaming `PepsiCo Kurkure Masala Munch 90g` to
`Kurkure Masala Munch 90g` does not merge the two rows — it mints a third id and
breaks any link pointing at either of the first two. You have just migrated onto
storing `image_id`. A well-meant cleanup on our side would re-break exactly what
you have finished repairing.
- The column mapper no longer maps a bare `Quantity` column to pack size. A
sheet with both `Pack Size` and `Quantity` now binds only `Pack Size`.
- `_sizes_for()` discards a unitless number and records why:
`ignored pack size '150': a number with no unit is a quantity, not a size`.
The same applies to stripping the stray `150` from `Lays Classic Salted 52g 150`.
So no new rows can acquire this. The 18 existing ones are legacy damage and we
will repair them — see the request below.
So before we touch it we would like to agree:
### The PepsiCo duplicate pairs — we need one decision from you first
Both pairs are still there (ids 661/730 and 662/731). We have deliberately not
touched them, because **de-duplicating means renaming, and `image_id` is derived
from the name.** Renaming `PepsiCo Kurkure Masala Munch 90g` to
`Kurkure Masala Munch 90g` does not merge the two rows — it mints a *third* id
and breaks any link pointing at either of the first two. You have just migrated
onto storing `image_id`; a well-meant cleanup on our side would re-break exactly
what you have finished repairing. The same applies to stripping the stray
numbers.
So, three things to agree before we act:
1. **Which convention wins** — brand prefix in the product name, or not? We have
no strong preference; we care only that it is one of them. Our lean is
*without* the prefix, since the brand is already a column.
2. **How the merge is communicated.** If we can give you the old-id →
new-id mapping for every row we touch, in advance, does that let you
re-point rather than clear? That is straightforward for us to produce.
no strong preference and care only that it is one of them. Our lean is
*without*, since the brand is already a column.
2. **Would an old-id → new-id mapping, delivered in advance for every row we
touch, let you re-point rather than clear?** That is straightforward for us
to produce and would cover both the 18 stray-number rows and the PepsiCo
pairs.
3. **Timing**, so it lands in one pass rather than trickling.
`haldiram` → `haldirams` is a three-row merge and much lower risk; we can do
that one immediately if you would rather not wait for the rest.
### Brand coverage
Confirmed from the live database: **55 tables, 1,414 products**. You counted
1,614 — worth reconciling, but the shape matches. And yes: **Britannia (6),
Parle (3) and Patanjali (3) are incomplete scrapes, not small brands.** We will
re-run those three.
---
@@ -241,7 +297,7 @@ because it was not only a coverage gap.
### What was actually happening
Produce rows were not rejected. They were **misfiled**, which is worse. The
brand fallback takes the first word of the name and then whole-word matches it
brand fallback takes the first word of the name and whole-word matches it
against our alias map:
```
@@ -249,12 +305,12 @@ Apple -> brand "Apple" -> junk table brand_apple
Tomato -> brand "Tomato" -> junk table brand_tomato
Bitter Gourd -> brand "Bitter" -> junk table brand_bitter
Curry Leaves -> brand "Curry" -> junk table brand_curry
Red Rose -> brand "Red" -> brand_brooke_bond <--
Red Rose -> brand "Red" -> brand_brooke_bond <--
```
That last one is not a typo. A rose was being written into the Brooke Bond tea
catalogue, and our enrichment then stamps that brand's real FSSAI licence number
onto the row. Your 139 hand-typed products were the visible symptom; this was
catalogue, and our enrichment then stamps that brand's real FSSAI licence onto
the row. Your 139 hand-typed products were the visible symptom; this was
underneath it.
### What now happens
@@ -265,28 +321,38 @@ them. Fruit, vegetables, greens, herbs, flowers, fish, eggs and loose dairy are
covered, alongside the pulses, grains, spices, oils and sugar that already were.
Five categories were added — Fruits & Vegetables, Fresh Herbs & Greens, Flowers,
Fish & Seafood, Eggs — with HSN codes and a 0% GST rate, since unprocessed
produce is nil-rated rather than reduced-rate.
Fish & Seafood, Eggs — with HSN codes at 0% GST, since unprocessed produce is
nil-rated rather than reduced-rate.
Merchant misspellings from your own data are handled: `Bitter guard`,
`Bottle ground`, `Ladies Finger` all resolve.
Misspellings from your own audit are handled: `Bitter guard`, `Bottle ground`
and `Ladies Finger` all resolve.
An uploaded produce row also now keeps only what the sheet actually said. Name,
weight and price are stored; HSN, SKU, barcode, FSSAI and description are left
null rather than invented, and **no pack sizes are generated** — which is cause
1 of your finding #01, kept out of this table from the start.
### The base list
**159 rows**, in the shape you asked for: name and category only, no brand, no
pack size, no price. Fruit (42), vegetables (53), greens and herbs (17), flowers
(14), fish and seafood (15), loose dairy (13), eggs (5). It includes the specific
items your audit listed — Jasmine, Lotus, Red Rose, Thulasi, Drumstick, Curry
Leaves, the four banana varieties, Tuna, Mackerel.
**159 rows**, in the shape you asked for: name and category, no brand, no pack
size, no price.
Each row carries a `search_query` embedding, so these are reachable through
semantic search and not just exact match. The list is hand-authored rather than
scraped, so it is not subject to any of #01.
```
Fruits & Vegetables 95 Fish & Seafood 15
Fresh Herbs & Greens 17 Dairy (loose) 13
Flowers 14 Eggs 5
```
**Images are not included yet.** You asked for name and image; we have shipped
the names. Sourcing 159 licensable produce photographs is a separate piece of
work and we did not want to hold the list for it — tell us if the list is not
useful to you without them and we will prioritise accordingly.
It includes the specific items your audit listed — Jasmine, Lotus, Red Rose,
Thulasi, Drumstick, Curry Leaves, the four banana varieties, Tuna, Mackerel.
Each row carries a search embedding, so these are reachable through semantic
search and not just exact match. The list is hand-authored rather than scraped,
so none of #01 applies to it.
**Images: 112 of the 159 rows (70%) currently have one** — all of the fruit,
vegetables, greens and herbs. Flowers, fish, loose dairy and eggs are still
name-only; we stopped the fetch part-way and will finish it. Tell us if the list
is more useful to you complete-but-later or partial-but-now.
### One limitation worth knowing
@@ -295,11 +361,11 @@ means "this is a brand". So place-qualified produce — `Salem Mango`,
`Mysore Banana`, `Jammu Apple`, all real strings from your Ragul Stores data —
still reads as branded, because `Mysore` is also a real brand (Mysore Sandal).
The workaround is already in the pipeline: **if the sheet has a Brand column and
leaves the cell empty, we believe it** and file the row under Own Products
There is a workaround already in the pipeline: **if the sheet has a Brand column
and leaves the cell empty, we believe it** and file the row under Own Products
regardless of the name. If your merchants' sheets carry an empty brand column,
those rows will land correctly. If they carry no brand column at all, the
name-based test is what applies.
those rows land correctly. If they carry no brand column at all, the name-based
test applies.
We would rather be conservative here. Collapsing a real regional brand into the
unbranded bucket is much harder to undo than a mango sitting in the wrong table.
@@ -311,13 +377,28 @@ unbranded bucket is much harder to undo than a mango sitting in the wrong table.
- The produce lexicon was run against **all 1,414 products in the live
catalogue** and against all **231 brand aliases**: zero reclassifications in
either. That check is now a test, so it runs on every change.
- It also surfaced a pre-existing bug we would not otherwise have found: our
pack-size stripper was eating the word after a number, so
`24 Mantra Organic Moong Dal 500g` lost its brand entirely and was being filed
as an unbranded commodity. **Every brand whose name starts with a digit hit
this.** "24 Mantra" appears on your finding-03 list, which is how we noticed.
Fixed.
- It surfaced a pre-existing bug we would not otherwise have found: our
pack-size stripper was eating the word *after* a number, so
`24 Mantra Organic Moong Dal 500g` lost its brand entirely and was filed as an
unbranded commodity. **Every brand whose name starts with a digit hit this.**
"24 Mantra" is on your finding-03 list, which is how we noticed. Fixed.
- All 159 seeded rows were checked to classify identically to how an uploaded
copy of the same name would, so a grocer typing "Tomato" lands on the seeded
row instead of creating a second one.
- Full suite: **1,054 tests passing.**
- Full suite: **1,108 tests passing.**
## Two things you did not ask about, but should know
**Barcodes.** We do not generate them; a product carries one only if the sheet
supplied it. 95 of our 1,414 products have one — and checking them against Open
Food Facts, **33 are attached to the wrong product** (`Lion Dates Powder` is
stored under a barcode Open Food Facts holds as a Dutch confection) and a
further 38 are not valid GTINs at all. If you join on barcode anywhere, treat
ours as unreliable until we have cleaned them. The lookup that produced them has
since been tightened.
**Nutrition provenance.** Rows in `nutrition_facts` sourced from Open Food Facts
were matched by *name*, with confidences as low as 0.32. We have built an exact
barcode-keyed lookup to replace that, but given the barcode quality above it can
currently upgrade only a handful of rows. Treat low-confidence nutrition as
indicative, not authoritative.