Reply to the Catalogue Drift Report · 31 Aug 2026
Every figure below was taken from the live deployment on 2 September, not read off source. Eight asks are fulfilled and serving. Three wait on one decision from you. Three are not done — and one answer we gave you last week was wrong.
Correction to our 01 September reply
We told you image_id was “a pure deterministic function of brand, product name and pack size.” That is true of the upload path and false of the scrape path — the one that broke your eleven links. Scraped ids carry a random uuid4 fragment. You migrated onto image_id on the strength of that answer, so please read 01b before anything else.
Two things that changed since we last wrote
The deploy has landed — from_drop is in the live openapi.json today — so the “please do not re-test” note in our last reply is out of date. Please do re-test. And 200 Hindustan Unilever products have been deleted since you took your measurement; that is finding 01 happening again, mid-conversation.
Your seven findings contain fourteen distinct asks. This is where each one stands.
| Ref | The ask | Status |
|---|---|---|
| 01a | Does a pack size that once existed survive a re-scrape? | Answered · no |
| 01b | Is image_id stable across re-scrapes? | Upload path only |
| 01c | superseded_by, a retired list, or a per-run changelog | Not built |
| 02 | Per-row rejection reason, with the sheet row number | Fulfilled |
| 03 | brand_key beside brand in the manifest | Fulfilled |
| 04 | from_drop on each file in a run | Fulfilled · live |
| 05 | source_row on each manifest entry | Fulfilled |
| 06a | Merge haldiram and haldirams | Fulfilled · live |
| 06b | De-duplicate the PepsiCo pairs, settle one prefix convention | Needs 06e |
| 06c | Strip the stray 150, check the pattern elsewhere | Cause fixed · 17 rows |
| 06d | Are Britannia, Parle and Patanjali complete? | Answered · no |
| 06e | Our question back: an old-id → new-id map before we rename | Awaiting you |
| 07 | A ~150-row loose-produce base list | 159 rows · live |
| — | Reconcile your 1,614 against our 1,414 | Exact |
You askedConfirm whether a pack size that once existed is meant to survive a re-scrape.
It is not, and that is unchanged. catalog_engine.py:999 still calls upsert_brand_products(brand, enhanced_products, cleanup=True), which deletes every row in the brand table whose image_id is absent from the batch being written. A SKU present on one run and absent on the next is removed, not retired.
The upload path is safe and always was: everything under POST /api/uploads/catalog uses cleanup=False, with a test holding it there. A sheet you send cannot delete a row it does not mention. The deletions come from brand scraping only.
image_id is deterministic on one path and random on the otherYou askedConfirm that image_id is stable across re-scrapes for a product whose name and pack size have not changed. We have now switched to storing it, and that switch only helps if the guarantee holds.
“Yes. It is a pure deterministic function of brand, product name and pack size, with no clock, counter or run id in it.” — our reply, 01 September
We answered from the wrong function. Two different paths mint image_id, and they do not behave alike:
build_image_id() is a pure function of brand + name + size. Verified today: two calls with the same input both return pepsico_cheetos_chips_100g. Storing it is safe.catalog_engine.py:695 mints ids with s3_service.generate_image_id(), which appends uuid4()[:8]. It reuses an existing id only when a row is found whose product_name is an exact string match — no size in the lookup, no normalisation, no fuzzy match.Both forms are visible side by side in the live catalogue right now:
scraped id 27 Cheetos Chips 250g image_id cheetos_chips_2d6bf74f
scraped id 6 Kurkure Menthol 50g image_id kurkure_menthol_49ef2d35
uploaded id 731 Kurkure Masala Munch 90g image_id pepsico_kurkure_masala_munch_90g
Those suffixes are random. So on a scrape, if the product name changes by a single character — exactly what a brand prefix appearing or disappearing does — the reuse lookup misses, a new random id is minted, and cleanup=True deletes the old row. That is the complete mechanism behind your eleven broken links, and image_id alone does not protect you from it.
What this means for your migration
Storing image_id instead of our row id is still the right move — it is strictly better than the row id, and it is stable for everything arriving through the upload path. But it is not yet the durable key we implied for scraped brands. Making it one means giving the scrape path the same deterministic build_image_id() the upload path uses. We have not done that, because it renames ids on the next scrape of every scraped brand — the same migration problem as 06b, needing the old-to-new map in 06e and your timing.
You askedAnything that lets us detect it — a superseded_by, a retired list, even a changelog per run — turns a silent break into something we can act on.
Not built. There is no retired_at, superseded_by or changelog anywhere in the codebase. Our questions from last week still stand, and the Hindustan Unilever deletion below makes them urgent:
retired_at timestamp plus exclusion from the default read work instead of the row being deleted? It preserves the image_id so your stored link resolves to something, and gives us somewhere to hang superseded_by and a per-run changelog.You askedA per-row reason on rejection, in the shape you already use for files — with row being the spreadsheet's own 1-based number, header included.
Delivered in that shape:
"rejections": [
{ "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
"reason": "title is too short to be a real product name" }
]
row is the 1-based sheet row with the header counted as row 1 — the row number is computed as position + 2, the same convention as the 422 responses, so it matches what the operator sees on screen. null only when the row cannot be located.slim=True strips only products.test_the_rejection_row_matches_the_offending_sheet_line. Documented in INGESTION_API.md with a field table.One correction to our own earlier account of this: the array itself was always in the response and only row is new. It was missing from our docs, which is why you could not find it. You were diffing 19 against 17 because we told you that was all you had.
brand_key is published beside the display nameYou askedInclude the catalogue key alongside the display name in the manifest — a brand_key field beside brand. Our normalisation is a guess that currently happens to be right.
Every products[] entry now carries it. It is produced by the same _sanitize_name() the storage layer uses to name the table, so it cannot drift from the key the catalogue is actually addressed by — and the test asserts exactly that identity rather than a hard-coded string:
assert product["brand_key"] == _sanitize_name(product["brand"])
Your normalisation is correct as far as we can tell. It stays a guess, though, and the failure mode is silent — a wrong key finds nothing rather than erroring. Use the published field.
from_drop — the one that could corrupt inventoryYou askedfrom_drop on each file in a run, carrying the drop id it was released from. Matching on it is exact, where matching on a filename is a coincidence we are relying on.
Each file in a run now carries from_drop, the exact inverse of released_to. It is stamped after staging, written to disk and re-read from the manifest — null only for a file that went straight into a run without sitting in an inbox.
Confirmed deployed. from_drop is in the live openapi.json served by mcp.nearle.ai.in today, which is also how we know the rest of this deploy has landed.
The collision you described is now a permanent regression guard: test_a_run_says_which_drop_each_file_came_from stages two drops from different senders both named a.csv and asserts they are distinguishable, and a second test proves the id survives a re-read of the run rather than existing only in the reply.
source_row traces every product to its sheet lineYou askedsource_row on each manifest entry — the 1-based sheet row that produced it, so we can say “row 14 became these three” and “rows 6 and 11 produced nothing”.
Both of those now work. source_row is the 1-based row with the header as row 1, and it is deliberately many-to-one: a pack-size cell reading 100g, 200g, 500g becomes three products that all report the same source_row. Rows that produced nothing are those absent from every entry — a set difference rather than a name-matching heuristic.
The map is built from the kept rows before dedupe and carried alongside the storage rows rather than inside them, so the number always describes the product that was actually written, not one that lost a dedupe. Tests cover the one-to-one case and the exploded case.
Live today: haldiram returns 0 products, haldirams returns 2. Merging the rows alone would have fixed nothing — the next sheet spelling it without the “s” would rebuild the table — so the alias went in too. Haldiram, haldiram, HALDIRAM and Haldiram's now all resolve to haldirams.
While merging we found something you could not have seen: both surviving rows were carrying Lion Dates' FSSAI licence rather than Haldiram's. That is a regulatory identifier on the wrong manufacturer's product. Fixed in the database and the seed file.
Confirmed live today, unchanged:
661 PepsiCo Lays Classic Salted 52g pepsico_pepsico_lays_classic_salted_52g
730 Lays Classic Salted 52g 150 pepsico_lays_classic_salted_52g_150
662 PepsiCo Kurkure Masala Munch 90g pepsico_pepsico_kurkure_masala_munch_90g
731 Kurkure Masala Munch 90g pepsico_kurkure_masala_munch_90g
We have not touched them, because de-duplicating means renaming, and image_id is derived from the name. Renaming PepsiCo Kurkure Masala Munch 90g does not merge the two rows — it mints a third id and breaks any link pointing at either of the first two. You have just finished migrating onto image_id; a well-meant cleanup on our side would re-break exactly what you repaired. Our lean on the convention is without the brand prefix, since brand is already its own column, but we will follow whichever you pick.
The number is not a price — 40 against ₹299, 150 against ₹21. It is the case-pack count: a sheet's “Quantity” column was being mapped to the pack-size field, the bare number became the size, and it was then appended to the product name.
The cause is fixed on two layers, both verified today. Quantity was removed from the pack-size keyword rule — only net qty / net quantity, the Indian labelling term for a real pack size, still match. A sheet with columns Product Name, Quantity, Pack Size, Brand now binds Pack Size and reports Quantity as unrecognised; previously a leading Quantity column also shut out the sheet's real Pack Size column, because mapping is first-wins by position. Independently, a unitless number is now discarded with a reason recorded: “ignored pack size '150': a number with no unit is a quantity, not a size.”
No new row can acquire this. The existing rows are still there — 17 of them, scanned across all 55 live brands today. We said 18 last week; the eighteenth was the corrupted Haldiram row removed in the merge.
brooke_bond 1 Red Label Tea 500g 30
fortune 2 Fortune Sunflower Oil 1L 48
hindustan_unilever 2135 Tata Salt Iodised 1kg 120
hindustan_unilever 2136 Bru Instant Coffee 100g 24
hindustan_unilever 2137 Horlicks Classic Malt 500g 18
hindustan_unilever 2138 Surf Excel Easy Wash 1kg 36
hindustan_unilever 2139 Vim Dishwash Bar 300g 90
hindustan_unilever 2140 Dove Cream Beauty Bar 100g 64
hindustan_unilever 2141 Clinic Plus Shampoo 175ml 40
india_gate 1 India Gate Basmati Rice 1kg 60
britannia 7 Britannia Good Day Cashew 200g 60
colgate_palmolive 483 Colgate Strong Teeth 200g 50
itc 5 Aashirvaad Shudh Chakki Atta 5kg 40
reckitt_benckiser 1 Harpic Power Plus 500ml 30
reckitt_benckiser 2 Dettol Original Soap 125g 80
coca_cola 1041 Coca-Cola 750ml 72
pepsico 730 Lays Classic Salted 52g 150
Repairing them is a rename, so it is blocked on 06e exactly as the PepsiCo pairs are. One thing unrelated but visible in that list: Tata Salt Iodised is filed under hindustan_unilever, and Tata Salt is not an HUL product. We are looking at it.
You read those right. Confirmed incomplete, not small brands — still 6, 3 and 3 products live today. The re-scrape has not been run yet; we would rather re-run them than have you build around the gap.
Three answers unblock 06b, 06c and the image_id determinism fix in 01b — all of which are renames, and all of which should land in one pass:
You askedA loose-produce base list — roughly 150 rows covering fruit, vegetables, greens, flowers, fish and milk. Name and image only; no brand, no pack size, no price.
Serving now under brand key own_products, in that shape:
Fruits & Vegetables 95 Flowers 14
Fresh Herbs & Greens 17 Dairy (loose) 13
Fish & Seafood 15 Eggs 5
159 rows
It includes the specific items your audit listed — Jasmine, Lotus, Red Rose, Thulasi, Drumstick, Curry Leaves, the four banana varieties, Tuna, Mackerel. Each row carries a search embedding, so these are reachable through semantic search and not just exact match, and every seeded row classifies identically to how an uploaded copy of the same name would — a grocer typing “Tomato” lands on the seeded row instead of creating a second one. The list is hand-authored rather than scraped, so none of finding 01 applies to it.
112 of the 159 (70%) carry an image — all of the fruit, vegetables, greens and herbs. Flowers, fish, loose dairy and eggs are still name-only; we stopped the fetch part-way. Tell us whether the list is more useful to you complete-but-later or partial-but-now.
The underlying bug was worse than a coverage gap
Produce rows were not rejected. They were misfiled. The brand fallback took the first word of the name and whole-word matched it against our alias map, so Apple became brand “Apple”, Curry Leaves became “Curry”, and Red Rose was being written into the Brooke Bond tea catalogue — where our enrichment then stamped that brand's real FSSAI licence onto it. Your 139 hand-typed products were the visible symptom; this was underneath. Loose goods are now recognised as commodities and filed under Own Products before brand inference can touch them.
Verified against your own audit strings today, including your merchants' misspellings:
Apple, Tomato, Curry Leaves, Red Rose, Thulasi, Drumstick, Tuna -> unbranded
Bitter guard, Bottle ground, Ladies Finger -> unbranded
Produce is also exempt from the invented-pack-size fallback that causes finding 01: an Own Products row with no weight column gets one Standard row, not three made-up ones. Name, weight and price are stored; HSN, SKU, barcode, FSSAI and description are left null rather than invented.
Salem Mango, Mysore Banana and Jammu Apple — all real strings from your Ragul Stores data — classify as branded, because Mysore is also a real brand. The classifier is deliberately conservative: collapsing a real regional brand into the unbranded bucket is much harder to undo than a mango in the wrong table. The workaround is already in the pipeline — if the sheet has a brand column and leaves the cell empty, we believe it and file the row under Own Products regardless of the name.Maceral does not resolve — your Ragul Stores spelling of mackerel. Mackerel does. Send us any other spellings your merchants actually use and we will add them.You counted 55 brands and 1,614 products. We measure 55 brands and 1,572 products today, of which 159 are the new produce rows — so 1,413 of the old kind. The 201-row gap resolves exactly:
your count 1,614
less Hindustan Unilever, 443 -> 243 today -200
less the Haldiram row removed in the merge -1
-------
1,413 our non-produce count
plus the produce base list +159
-------
1,572 live today
This is not a counting discrepancy
200 Hindustan Unilever products have been deleted since you measured — the largest brand in the catalogue, between your report and this reply. It is finding 01a happening again while we were writing about it, and it is the strongest argument we have for the retirement model in 01c. That is the answer we would like from you soonest.