Reply to the Catalogue Drift Report · 31 Aug 2026

Seven findings, fourteen asks, re‑measured

Every figure below was taken from the live deployment on 2 September, not read off source. Eight asks are fulfilled and serving. Three wait on one decision from you. Three are not done — and one answer we gave you last week was wrong.

Verified 2026-09-02 Live catalogue 55 brands / 1,572 products Backend suite 1,108 passing Deploy landed
8Fulfilled & live
3Blocked on your call
3Not done

Correction to our 01 September reply

We told you image_id was “a pure deterministic function of brand, product name and pack size.” That is true of the upload path and false of the scrape path — the one that broke your eleven links. Scraped ids carry a random uuid4 fragment. You migrated onto image_id on the strength of that answer, so please read 01b before anything else.

Two things that changed since we last wrote

The deploy has landed — from_drop is in the live openapi.json today — so the “please do not re-test” note in our last reply is out of date. Please do re-test. And 200 Hindustan Unilever products have been deleted since you took your measurement; that is finding 01 happening again, mid-conversation.

The ledger

Your seven findings contain fourteen distinct asks. This is where each one stands.

RefThe askStatus
01aDoes a pack size that once existed survive a re-scrape?Answered · no
01bIs image_id stable across re-scrapes?Upload path only
01csuperseded_by, a retired list, or a per-run changelogNot built
02Per-row rejection reason, with the sheet row numberFulfilled
03brand_key beside brand in the manifestFulfilled
04from_drop on each file in a runFulfilled · live
05source_row on each manifest entryFulfilled
06aMerge haldiram and haldiramsFulfilled · live
06bDe-duplicate the PepsiCo pairs, settle one prefix conventionNeeds 06e
06cStrip the stray 150, check the pattern elsewhereCause fixed · 17 rows
06dAre Britannia, Parle and Patanjali complete?Answered · no
06eOur question back: an old-id → new-id map before we renameAwaiting you
07A ~150-row loose-produce base list159 rows · live
—Reconcile your 1,614 against our 1,414Exact

01a  Pack sizes are still replaced, not added to

Root cause live

You askedConfirm whether a pack size that once existed is meant to survive a re-scrape.

It is not, and that is unchanged. catalog_engine.py:999 still calls upsert_brand_products(brand, enhanced_products, cleanup=True), which deletes every row in the brand table whose image_id is absent from the batch being written. A SKU present on one run and absent on the next is removed, not retired.

The upload path is safe and always was: everything under POST /api/uploads/catalog uses cleanup=False, with a test holding it there. A sheet you send cannot delete a row it does not mention. The deletions come from brand scraping only.

01b  image_id is deterministic on one path and random on the other

We answered wrong

You askedConfirm that image_id is stable across re-scrapes for a product whose name and pack size have not changed. We have now switched to storing it, and that switch only helps if the guarantee holds.

“Yes. It is a pure deterministic function of brand, product name and pack size, with no clock, counter or run id in it.” — our reply, 01 September

We answered from the wrong function. Two different paths mint image_id, and they do not behave alike:

Both forms are visible side by side in the live catalogue right now:

scraped   id  27   Cheetos Chips 250g          image_id  cheetos_chips_2d6bf74f
scraped   id   6   Kurkure Menthol 50g         image_id  kurkure_menthol_49ef2d35
uploaded  id 731   Kurkure Masala Munch 90g    image_id  pepsico_kurkure_masala_munch_90g

Those suffixes are random. So on a scrape, if the product name changes by a single character — exactly what a brand prefix appearing or disappearing does — the reuse lookup misses, a new random id is minted, and cleanup=True deletes the old row. That is the complete mechanism behind your eleven broken links, and image_id alone does not protect you from it.

What this means for your migration

Storing image_id instead of our row id is still the right move — it is strictly better than the row id, and it is stable for everything arriving through the upload path. But it is not yet the durable key we implied for scraped brands. Making it one means giving the scrape path the same deterministic build_image_id() the upload path uses. We have not done that, because it renames ids on the next scrape of every scraped brand — the same migration problem as 06b, needing the old-to-new map in 06e and your timing.

01c  Nothing tells you when a product is dropped or renamed

Not built

You askedAnything that lets us detect it — a superseded_by, a retired list, even a changelog per run — turns a silent break into something we can act on.

Not built. There is no retired_at, superseded_by or changelog anywhere in the codebase. Our questions from last week still stand, and the Hindustan Unilever deletion below makes them urgent:

  1. Would a retired_at timestamp plus exclusion from the default read work instead of the row being deleted? It preserves the image_id so your stored link resolves to something, and gives us somewhere to hang superseded_by and a per-run changelog.
  2. If so, should retired rows stay reachable through an explicit query, or vanish from the API entirely?

02  Rejections now name the row and the reason

Fulfilled

You askedA per-row reason on rejection, in the shape you already use for files — with row being the spreadsheet's own 1-based number, header included.

Delivered in that shape:

"rejections": [
  { "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
    "reason": "title is too short to be a real product name" }
]

One correction to our own earlier account of this: the array itself was always in the response and only row is new. It was missing from our docs, which is why you could not find it. You were diffing 19 against 17 because we told you that was all you had.

03  brand_key is published beside the display name

Fulfilled

You askedInclude the catalogue key alongside the display name in the manifest — a brand_key field beside brand. Our normalisation is a guess that currently happens to be right.

Every products[] entry now carries it. It is produced by the same _sanitize_name() the storage layer uses to name the table, so it cannot drift from the key the catalogue is actually addressed by — and the test asserts exactly that identity rather than a hard-coded string:

assert product["brand_key"] == _sanitize_name(product["brand"])

Your normalisation is correct as far as we can tell. It stays a guess, though, and the failure mode is silent — a wrong key finds nothing rather than erroring. Use the published field.

04  from_drop — the one that could corrupt inventory

Fulfilled & live

You askedfrom_drop on each file in a run, carrying the drop id it was released from. Matching on it is exact, where matching on a filename is a coincidence we are relying on.

Each file in a run now carries from_drop, the exact inverse of released_to. It is stamped after staging, written to disk and re-read from the manifest — null only for a file that went straight into a run without sitting in an inbox.

Confirmed deployed. from_drop is in the live openapi.json served by mcp.nearle.ai.in today, which is also how we know the rest of this deploy has landed.

The collision you described is now a permanent regression guard: test_a_run_says_which_drop_each_file_came_from stages two drops from different senders both named a.csv and asserts they are distinguishable, and a second test proves the id survives a re-read of the run rather than existing only in the reply.

05  source_row traces every product to its sheet line

Fulfilled

You askedsource_row on each manifest entry — the 1-based sheet row that produced it, so we can say “row 14 became these three” and “rows 6 and 11 produced nothing”.

Both of those now work. source_row is the 1-based row with the header as row 1, and it is deliberately many-to-one: a pack-size cell reading 100g, 200g, 500g becomes three products that all report the same source_row. Rows that produced nothing are those absent from every entry — a set difference rather than a name-matching heuristic.

The map is built from the kept rows before dedupe and carried alongside the storage rows rather than inside them, so the number always describes the product that was actually written, not one that lost a dedupe. Tests cover the one-to-one case and the exploded case.


06  Duplicates, stray values, and coverage

Mixed

06a  The Haldiram split is merged, and cannot recur

Live today: haldiram returns 0 products, haldirams returns 2. Merging the rows alone would have fixed nothing — the next sheet spelling it without the “s” would rebuild the table — so the alias went in too. Haldiram, haldiram, HALDIRAM and Haldiram's now all resolve to haldirams.

While merging we found something you could not have seen: both surviving rows were carrying Lion Dates' FSSAI licence rather than Haldiram's. That is a regulatory identifier on the wrong manufacturer's product. Fixed in the database and the seed file.

06b  The PepsiCo pairs are still there, deliberately

Confirmed live today, unchanged:

661  PepsiCo Lays Classic Salted 52g     pepsico_pepsico_lays_classic_salted_52g
730  Lays Classic Salted 52g 150         pepsico_lays_classic_salted_52g_150
662  PepsiCo Kurkure Masala Munch 90g    pepsico_pepsico_kurkure_masala_munch_90g
731  Kurkure Masala Munch 90g            pepsico_kurkure_masala_munch_90g

We have not touched them, because de-duplicating means renaming, and image_id is derived from the name. Renaming PepsiCo Kurkure Masala Munch 90g does not merge the two rows — it mints a third id and breaks any link pointing at either of the first two. You have just finished migrating onto image_id; a well-meant cleanup on our side would re-break exactly what you repaired. Our lean on the convention is without the brand prefix, since brand is already its own column, but we will follow whichever you pick.

06c  The stray number: cause fixed, 17 rows still carrying it

The number is not a price — 40 against ₹299, 150 against ₹21. It is the case-pack count: a sheet's “Quantity” column was being mapped to the pack-size field, the bare number became the size, and it was then appended to the product name.

The cause is fixed on two layers, both verified today. Quantity was removed from the pack-size keyword rule — only net qty / net quantity, the Indian labelling term for a real pack size, still match. A sheet with columns Product Name, Quantity, Pack Size, Brand now binds Pack Size and reports Quantity as unrecognised; previously a leading Quantity column also shut out the sheet's real Pack Size column, because mapping is first-wins by position. Independently, a unitless number is now discarded with a reason recorded: “ignored pack size '150': a number with no unit is a quantity, not a size.”

No new row can acquire this. The existing rows are still there — 17 of them, scanned across all 55 live brands today. We said 18 last week; the eighteenth was the corrupted Haldiram row removed in the merge.

brooke_bond           1  Red Label Tea 500g 30
fortune               2  Fortune Sunflower Oil 1L 48
hindustan_unilever 2135  Tata Salt Iodised 1kg 120
hindustan_unilever 2136  Bru Instant Coffee 100g 24
hindustan_unilever 2137  Horlicks Classic Malt 500g 18
hindustan_unilever 2138  Surf Excel Easy Wash 1kg 36
hindustan_unilever 2139  Vim Dishwash Bar 300g 90
hindustan_unilever 2140  Dove Cream Beauty Bar 100g 64
hindustan_unilever 2141  Clinic Plus Shampoo 175ml 40
india_gate            1  India Gate Basmati Rice 1kg 60
britannia             7  Britannia Good Day Cashew 200g 60
colgate_palmolive   483  Colgate Strong Teeth 200g 50
itc                   5  Aashirvaad Shudh Chakki Atta 5kg 40
reckitt_benckiser     1  Harpic Power Plus 500ml 30
reckitt_benckiser     2  Dettol Original Soap 125g 80
coca_cola          1041  Coca-Cola 750ml 72
pepsico             730  Lays Classic Salted 52g 150

Repairing them is a rename, so it is blocked on 06e exactly as the PepsiCo pairs are. One thing unrelated but visible in that list: Tata Salt Iodised is filed under hindustan_unilever, and Tata Salt is not an HUL product. We are looking at it.

06d  Britannia, Parle and Patanjali are incomplete scrapes

You read those right. Confirmed incomplete, not small brands — still 6, 3 and 3 products live today. The re-scrape has not been run yet; we would rather re-run them than have you build around the gap.

06e  What we need back from you

Three answers unblock 06b, 06c and the image_id determinism fix in 01b — all of which are renames, and all of which should land in one pass:

  1. Which convention wins? Brand prefix in the product name, or not. Either is fine; we care only that it is one of them.
  2. Would an old-id → new-id map, delivered in advance for every row we touch, let you re-point rather than clear? It is straightforward for us to produce, and it would cover the 17 stray-number rows, the PepsiCo pairs and the scraped-id migration together.
  3. Timing, so it lands in one pass rather than trickling.

07  Loose produce: 159 rows, live now

Fulfilled & live

You askedA loose-produce base list — roughly 150 rows covering fruit, vegetables, greens, flowers, fish and milk. Name and image only; no brand, no pack size, no price.

Serving now under brand key own_products, in that shape:

Fruits & Vegetables   95     Flowers          14
Fresh Herbs & Greens  17     Dairy (loose)    13
Fish & Seafood        15     Eggs              5
                                     159 rows

It includes the specific items your audit listed — Jasmine, Lotus, Red Rose, Thulasi, Drumstick, Curry Leaves, the four banana varieties, Tuna, Mackerel. Each row carries a search embedding, so these are reachable through semantic search and not just exact match, and every seeded row classifies identically to how an uploaded copy of the same name would — a grocer typing “Tomato” lands on the seeded row instead of creating a second one. The list is hand-authored rather than scraped, so none of finding 01 applies to it.

112 of the 159 (70%) carry an image — all of the fruit, vegetables, greens and herbs. Flowers, fish, loose dairy and eggs are still name-only; we stopped the fetch part-way. Tell us whether the list is more useful to you complete-but-later or partial-but-now.

The underlying bug was worse than a coverage gap

Produce rows were not rejected. They were misfiled. The brand fallback took the first word of the name and whole-word matched it against our alias map, so Apple became brand “Apple”, Curry Leaves became “Curry”, and Red Rose was being written into the Brooke Bond tea catalogue — where our enrichment then stamped that brand's real FSSAI licence onto it. Your 139 hand-typed products were the visible symptom; this was underneath. Loose goods are now recognised as commodities and filed under Own Products before brand inference can touch them.

Verified against your own audit strings today, including your merchants' misspellings:

Apple, Tomato, Curry Leaves, Red Rose, Thulasi, Drumstick, Tuna   -> unbranded
Bitter guard, Bottle ground, Ladies Finger                        -> unbranded

Produce is also exempt from the invented-pack-size fallback that causes finding 01: an Own Products row with no weight column gets one Standard row, not three made-up ones. Name, weight and price are stored; HSN, SKU, barcode, FSSAI and description are left null rather than invented.

Two limitations worth knowing


Your 1,614 against our count

You counted 55 brands and 1,614 products. We measure 55 brands and 1,572 products today, of which 159 are the new produce rows — so 1,413 of the old kind. The 201-row gap resolves exactly:

your count                                        1,614
  less Hindustan Unilever, 443 -> 243 today        -200
  less the Haldiram row removed in the merge         -1
                                                  -------
                                                   1,413   our non-produce count
  plus the produce base list                       +159
                                                  -------
                                                   1,572   live today

This is not a counting discrepancy

200 Hindustan Unilever products have been deleted since you measured — the largest brand in the catalogue, between your report and this reply. It is finding 01a happening again while we were writing about it, and it is the strongest argument we have for the retirement model in 01c. That is the answer we would like from you soonest.