diff --git a/docs/DRIFT_REPORT_VERIFICATION.html b/docs/DRIFT_REPORT_VERIFICATION.html
new file mode 100644
index 0000000..55efd87
--- /dev/null
+++ b/docs/DRIFT_REPORT_VERIFICATION.html
@@ -0,0 +1,644 @@
+
+
+
+
+
+ 8Fulfilled & live
+ 3Blocked on your call
+ 3Not done
+
+
+
+
Correction to our 01 September reply
+
We told you image_id was “a pure deterministic function of brand, product name and pack size.” That is true of the upload path and false of the scrape path — the one that broke your eleven links. Scraped ids carry a random uuid4 fragment. You migrated onto image_id on the strength of that answer, so please read 01b before anything else.
+
+
+
+
Two things that changed since we last wrote
+
The deploy has landed — from_drop is in the live openapi.json today — so the “please do not re-test” note in our last reply is out of date. Please do re-test. And 200 Hindustan Unilever products have been deleted since you took your measurement; that is finding 01 happening again, mid-conversation.
+
+
+
The ledger
+
Your seven findings contain fourteen distinct asks. This is where each one stands.
+
+
+
+
+
+
+
01a Pack sizes are still replaced, not added to
+ Root cause live
+
+
+
You askedConfirm whether a pack size that once existed is meant to survive a re-scrape.
+
+
It is not, and that is unchanged. catalog_engine.py:999 still calls upsert_brand_products(brand, enhanced_products, cleanup=True), which deletes every row in the brand table whose image_id is absent from the batch being written. A SKU present on one run and absent on the next is removed, not retired.
+
+
The upload path is safe and always was: everything under POST /api/uploads/catalog uses cleanup=False, with a test holding it there. A sheet you send cannot delete a row it does not mention. The deletions come from brand scraping only.
+
+
+
01b image_id is deterministic on one path and random on the other
+ We answered wrong
+
+
+
You askedConfirm that image_id is stable across re-scrapes for a product whose name and pack size have not changed. We have now switched to storing it, and that switch only helps if the guarantee holds.
+
+
“Yes. It is a pure deterministic function of brand, product name and pack size, with no clock, counter or run id in it.” — our reply, 01 September
+
+
We answered from the wrong function. Two different paths mint image_id, and they do not behave alike:
+
+
+ - Upload path —
build_image_id() is a pure function of brand + name + size. Verified today: two calls with the same input both return pepsico_cheetos_chips_100g. Storing it is safe.
+ - Scrape path —
catalog_engine.py:695 mints ids with s3_service.generate_image_id(), which appends uuid4()[:8]. It reuses an existing id only when a row is found whose product_name is an exact string match — no size in the lookup, no normalisation, no fuzzy match.
+
+
+
Both forms are visible side by side in the live catalogue right now:
+
+
scraped id 27 Cheetos Chips 250g image_id cheetos_chips_2d6bf74f
+scraped id 6 Kurkure Menthol 50g image_id kurkure_menthol_49ef2d35
+uploaded id 731 Kurkure Masala Munch 90g image_id pepsico_kurkure_masala_munch_90g
+
+
Those suffixes are random. So on a scrape, if the product name changes by a single character — exactly what a brand prefix appearing or disappearing does — the reuse lookup misses, a new random id is minted, and cleanup=True deletes the old row. That is the complete mechanism behind your eleven broken links, and image_id alone does not protect you from it.
+
+
+
What this means for your migration
+
Storing image_id instead of our row id is still the right move — it is strictly better than the row id, and it is stable for everything arriving through the upload path. But it is not yet the durable key we implied for scraped brands. Making it one means giving the scrape path the same deterministic build_image_id() the upload path uses. We have not done that, because it renames ids on the next scrape of every scraped brand — the same migration problem as 06b, needing the old-to-new map in 06e and your timing.
+
+
+
+
01c Nothing tells you when a product is dropped or renamed
+ Not built
+
+
+
You askedAnything that lets us detect it — a superseded_by, a retired list, even a changelog per run — turns a silent break into something we can act on.
+
+
Not built. There is no retired_at, superseded_by or changelog anywhere in the codebase. Our questions from last week still stand, and the Hindustan Unilever deletion below makes them urgent:
+
+ - Would a
retired_at timestamp plus exclusion from the default read work instead of the row being deleted? It preserves the image_id so your stored link resolves to something, and gives us somewhere to hang superseded_by and a per-run changelog.
+ - If so, should retired rows stay reachable through an explicit query, or vanish from the API entirely?
+
+
+
+
+
+
02 Rejections now name the row and the reason
+ Fulfilled
+
+
+
You askedA per-row reason on rejection, in the shape you already use for files — with row being the spreadsheet's own 1-based number, header included.
+
+
Delivered in that shape:
+
+
"rejections": [
+ { "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
+ "reason": "title is too short to be a real product name" }
+]
+
+
+ row is the 1-based sheet row with the header counted as row 1 — the row number is computed as position + 2, the same convention as the 422 responses, so it matches what the operator sees on screen. null only when the row cannot be located.
+ - Capped at 50 per file.
+ - Present on both the single-batch read and the list endpoint —
slim=True strips only products.
+ - Regression test:
test_the_rejection_row_matches_the_offending_sheet_line. Documented in INGESTION_API.md with a field table.
+
+
+
One correction to our own earlier account of this: the array itself was always in the response and only row is new. It was missing from our docs, which is why you could not find it. You were diffing 19 against 17 because we told you that was all you had.
+
+
+
03 brand_key is published beside the display name
+ Fulfilled
+
+
+
You askedInclude the catalogue key alongside the display name in the manifest — a brand_key field beside brand. Our normalisation is a guess that currently happens to be right.
+
+
Every products[] entry now carries it. It is produced by the same _sanitize_name() the storage layer uses to name the table, so it cannot drift from the key the catalogue is actually addressed by — and the test asserts exactly that identity rather than a hard-coded string:
+
+
assert product["brand_key"] == _sanitize_name(product["brand"])
+
+
Your normalisation is correct as far as we can tell. It stays a guess, though, and the failure mode is silent — a wrong key finds nothing rather than erroring. Use the published field.
+
+
+
04 from_drop — the one that could corrupt inventory
+ Fulfilled & live
+
+
+
You askedfrom_drop on each file in a run, carrying the drop id it was released from. Matching on it is exact, where matching on a filename is a coincidence we are relying on.
+
+
Each file in a run now carries from_drop, the exact inverse of released_to. It is stamped after staging, written to disk and re-read from the manifest — null only for a file that went straight into a run without sitting in an inbox.
+
+
Confirmed deployed. from_drop is in the live openapi.json served by mcp.nearle.ai.in today, which is also how we know the rest of this deploy has landed.
+
+
The collision you described is now a permanent regression guard: test_a_run_says_which_drop_each_file_came_from stages two drops from different senders both named a.csv and asserts they are distinguishable, and a second test proves the id survives a re-read of the run rather than existing only in the reply.
+
+
+
05 source_row traces every product to its sheet line
+ Fulfilled
+
+
+
You askedsource_row on each manifest entry — the 1-based sheet row that produced it, so we can say “row 14 became these three” and “rows 6 and 11 produced nothing”.
+
+
Both of those now work. source_row is the 1-based row with the header as row 1, and it is deliberately many-to-one: a pack-size cell reading 100g, 200g, 500g becomes three products that all report the same source_row. Rows that produced nothing are those absent from every entry — a set difference rather than a name-matching heuristic.
+
+
The map is built from the kept rows before dedupe and carried alongside the storage rows rather than inside them, so the number always describes the product that was actually written, not one that lost a dedupe. Tests cover the one-to-one case and the exploded case.
+
+
+
+
+
06 Duplicates, stray values, and coverage
+ Mixed
+
+
+
06a The Haldiram split is merged, and cannot recur
+
Live today: haldiram returns 0 products, haldirams returns 2. Merging the rows alone would have fixed nothing — the next sheet spelling it without the “s” would rebuild the table — so the alias went in too. Haldiram, haldiram, HALDIRAM and Haldiram's now all resolve to haldirams.
+
While merging we found something you could not have seen: both surviving rows were carrying Lion Dates' FSSAI licence rather than Haldiram's. That is a regulatory identifier on the wrong manufacturer's product. Fixed in the database and the seed file.
+
+
06b The PepsiCo pairs are still there, deliberately
+
Confirmed live today, unchanged:
+
661 PepsiCo Lays Classic Salted 52g pepsico_pepsico_lays_classic_salted_52g
+730 Lays Classic Salted 52g 150 pepsico_lays_classic_salted_52g_150
+662 PepsiCo Kurkure Masala Munch 90g pepsico_pepsico_kurkure_masala_munch_90g
+731 Kurkure Masala Munch 90g pepsico_kurkure_masala_munch_90g
+
We have not touched them, because de-duplicating means renaming, and image_id is derived from the name. Renaming PepsiCo Kurkure Masala Munch 90g does not merge the two rows — it mints a third id and breaks any link pointing at either of the first two. You have just finished migrating onto image_id; a well-meant cleanup on our side would re-break exactly what you repaired. Our lean on the convention is without the brand prefix, since brand is already its own column, but we will follow whichever you pick.
+
+
06c The stray number: cause fixed, 17 rows still carrying it
+
The number is not a price — 40 against ₹299, 150 against ₹21. It is the case-pack count: a sheet's “Quantity” column was being mapped to the pack-size field, the bare number became the size, and it was then appended to the product name.
+
The cause is fixed on two layers, both verified today. Quantity was removed from the pack-size keyword rule — only net qty / net quantity, the Indian labelling term for a real pack size, still match. A sheet with columns Product Name, Quantity, Pack Size, Brand now binds Pack Size and reports Quantity as unrecognised; previously a leading Quantity column also shut out the sheet's real Pack Size column, because mapping is first-wins by position. Independently, a unitless number is now discarded with a reason recorded: “ignored pack size '150': a number with no unit is a quantity, not a size.”
+
No new row can acquire this. The existing rows are still there — 17 of them, scanned across all 55 live brands today. We said 18 last week; the eighteenth was the corrupted Haldiram row removed in the merge.
+
brooke_bond 1 Red Label Tea 500g 30
+fortune 2 Fortune Sunflower Oil 1L 48
+hindustan_unilever 2135 Tata Salt Iodised 1kg 120
+hindustan_unilever 2136 Bru Instant Coffee 100g 24
+hindustan_unilever 2137 Horlicks Classic Malt 500g 18
+hindustan_unilever 2138 Surf Excel Easy Wash 1kg 36
+hindustan_unilever 2139 Vim Dishwash Bar 300g 90
+hindustan_unilever 2140 Dove Cream Beauty Bar 100g 64
+hindustan_unilever 2141 Clinic Plus Shampoo 175ml 40
+india_gate 1 India Gate Basmati Rice 1kg 60
+britannia 7 Britannia Good Day Cashew 200g 60
+colgate_palmolive 483 Colgate Strong Teeth 200g 50
+itc 5 Aashirvaad Shudh Chakki Atta 5kg 40
+reckitt_benckiser 1 Harpic Power Plus 500ml 30
+reckitt_benckiser 2 Dettol Original Soap 125g 80
+coca_cola 1041 Coca-Cola 750ml 72
+pepsico 730 Lays Classic Salted 52g 150
+
Repairing them is a rename, so it is blocked on 06e exactly as the PepsiCo pairs are. One thing unrelated but visible in that list: Tata Salt Iodised is filed under hindustan_unilever, and Tata Salt is not an HUL product. We are looking at it.
+
+
06d Britannia, Parle and Patanjali are incomplete scrapes
+
You read those right. Confirmed incomplete, not small brands — still 6, 3 and 3 products live today. The re-scrape has not been run yet; we would rather re-run them than have you build around the gap.
+
+
06e What we need back from you
+
Three answers unblock 06b, 06c and the image_id determinism fix in 01b — all of which are renames, and all of which should land in one pass:
+
+ - Which convention wins? Brand prefix in the product name, or not. Either is fine; we care only that it is one of them.
+ - Would an old-id → new-id map, delivered in advance for every row we touch, let you re-point rather than clear? It is straightforward for us to produce, and it would cover the 17 stray-number rows, the PepsiCo pairs and the scraped-id migration together.
+ - Timing, so it lands in one pass rather than trickling.
+
+
+
+
+
+
07 Loose produce: 159 rows, live now
+ Fulfilled & live
+
+
+
You askedA loose-produce base list — roughly 150 rows covering fruit, vegetables, greens, flowers, fish and milk. Name and image only; no brand, no pack size, no price.
+
+
Serving now under brand key own_products, in that shape:
+
+
Fruits & Vegetables 95 Flowers 14
+Fresh Herbs & Greens 17 Dairy (loose) 13
+Fish & Seafood 15 Eggs 5
+ 159 rows
+
+
It includes the specific items your audit listed — Jasmine, Lotus, Red Rose, Thulasi, Drumstick, Curry Leaves, the four banana varieties, Tuna, Mackerel. Each row carries a search embedding, so these are reachable through semantic search and not just exact match, and every seeded row classifies identically to how an uploaded copy of the same name would — a grocer typing “Tomato” lands on the seeded row instead of creating a second one. The list is hand-authored rather than scraped, so none of finding 01 applies to it.
+
+
112 of the 159 (70%) carry an image — all of the fruit, vegetables, greens and herbs. Flowers, fish, loose dairy and eggs are still name-only; we stopped the fetch part-way. Tell us whether the list is more useful to you complete-but-later or partial-but-now.
+
+
+
The underlying bug was worse than a coverage gap
+
Produce rows were not rejected. They were misfiled. The brand fallback took the first word of the name and whole-word matched it against our alias map, so Apple became brand “Apple”, Curry Leaves became “Curry”, and Red Rose was being written into the Brooke Bond tea catalogue — where our enrichment then stamped that brand's real FSSAI licence onto it. Your 139 hand-typed products were the visible symptom; this was underneath. Loose goods are now recognised as commodities and filed under Own Products before brand inference can touch them.
+
+
+
Verified against your own audit strings today, including your merchants' misspellings:
+
+
Apple, Tomato, Curry Leaves, Red Rose, Thulasi, Drumstick, Tuna -> unbranded
+Bitter guard, Bottle ground, Ladies Finger -> unbranded
+
+
Produce is also exempt from the invented-pack-size fallback that causes finding 01: an Own Products row with no weight column gets one Standard row, not three made-up ones. Name, weight and price are stored; HSN, SKU, barcode, FSSAI and description are left null rather than invented.
+
+
Two limitations worth knowing
+
+ - Place-qualified produce still reads as branded.
Salem Mango, Mysore Banana and Jammu Apple — all real strings from your Ragul Stores data — classify as branded, because Mysore is also a real brand. The classifier is deliberately conservative: collapsing a real regional brand into the unbranded bucket is much harder to undo than a mango in the wrong table. The workaround is already in the pipeline — if the sheet has a brand column and leaves the cell empty, we believe it and file the row under Own Products regardless of the name.
+ Maceral does not resolve — your Ragul Stores spelling of mackerel. Mackerel does. Send us any other spellings your merchants actually use and we will add them.
+
+
+
+
+
Your 1,614 against our count
+
You counted 55 brands and 1,614 products. We measure 55 brands and 1,572 products today, of which 159 are the new produce rows — so 1,413 of the old kind. The 201-row gap resolves exactly:
+
+
your count 1,614
+ less Hindustan Unilever, 443 -> 243 today -200
+ less the Haldiram row removed in the merge -1
+ -------
+ 1,413 our non-produce count
+ plus the produce base list +159
+ -------
+ 1,572 live today
+
+
+
This is not a counting discrepancy
+
200 Hindustan Unilever products have been deleted since you measured — the largest brand in the catalogue, between your report and this reply. It is finding 01a happening again while we were writing about it, and it is the strongest argument we have for the retirement model in 01c. That is the answer we would like from you soonest.
+
+
+
+
+
diff --git a/docs/DRIFT_REPORT_VERIFICATION.md b/docs/DRIFT_REPORT_VERIFICATION.md
new file mode 100644
index 0000000..a3e5da2
--- /dev/null
+++ b/docs/DRIFT_REPORT_VERIFICATION.md
@@ -0,0 +1,351 @@
+# Catalogue Drift Report — requirement-by-requirement verification
+
+Re-measured **2026-09-02** against the live deployment and the current source.
+This supersedes the figures in `DRIFT_REPORT_RESPONSE.md` (written 01 Sep), which
+was drafted before the deploy landed and contains **one claim that is wrong** —
+see the correction under 01(b).
+
+Fourteen distinct asks are contained in the seven findings. Eight are fulfilled,
+three are blocked on a decision from the integrator, three are not done.
+
+| # | The ask | Status |
+| --- | --- | --- |
+| 01a | Does a pack size that once existed survive a re-scrape? | **Answered: no.** Root cause still live |
+| 01b | Is `image_id` stable across re-scrapes? | **Answered: only on the upload path.** Previous answer was wrong |
+| 01c | `superseded_by` / retired list / per-run changelog | **Not built** |
+| 02 | Per-row rejection reason, with the sheet row number | **Fulfilled** |
+| 03 | `brand_key` beside `brand` in the manifest | **Fulfilled** |
+| 04 | `from_drop` on each file in a run | **Fulfilled, and live** |
+| 05 | `source_row` on each manifest entry | **Fulfilled** |
+| 06a | Merge `haldiram` and `haldirams` | **Fulfilled, and live** |
+| 06b | De-duplicate the PepsiCo pairs, settle one prefix convention | **Not done** — blocked on 06e |
+| 06c | Strip the stray `150`, check for the pattern elsewhere | **Cause fixed; 17 existing rows not yet repaired** |
+| 06d | Are Britannia, Parle, Patanjali complete? | **Answered: no, incomplete.** Re-scrape not yet run |
+| 06e | (our question back) old-id to new-id map before we rename anything | **Awaiting your answer** |
+| 07 | ~150-row loose-produce base list | **Fulfilled, and live** — 159 rows |
+| — | Reconcile your 1,614 against our 1,414 | **Reconciled exactly** — see below |
+
+Deployment state: the backend deploy **has** landed (`from_drop` is present in the
+live `openapi.json` at `mcp.nearle.ai.in`), and both database changes — the
+Haldiram merge and the 159 produce rows — are serving. The "not live yet, please
+do not re-test" caveat in the 01 Sep reply is out of date; **please do re-test.**
+
+---
+
+## 01a — Does a pack size survive a re-scrape?
+
+**No, and that is unchanged.** `catalog_engine.py:999` still calls
+`upsert_brand_products(brand, enhanced_products, cleanup=True)`, which deletes
+every row in the brand table whose `image_id` is absent from the batch being
+written. A SKU present on one run and absent on the next is removed, not retired.
+
+The upload path is still safe and always was: everything under
+`POST /api/uploads/catalog` uses `cleanup=False` (`store_catalog_pipeline.py:779`,
+with a test). A sheet you send cannot delete a row it does not mention.
+
+**This is still destroying rows.** See the reconciliation at the end — 200
+Hindustan Unilever products have disappeared since you took your measurement.
+
+## 01b — Is `image_id` stable across re-scrapes?
+
+**Correction to what we told you on 01 September.** We said:
+
+> "Yes. It is a pure deterministic function of brand, product name and pack
+> size, with no clock, counter or run id in it."
+
+That is true of the **upload** path and false of the **scrape** path, which is
+the path that broke your eleven links. We answered from the wrong function.
+
+- **Upload path** — `build_image_id()` (`store_catalog_pipeline.py:655`) is a
+ pure function of brand + name + size. Verified: two calls with the same input
+ both return `pepsico_cheetos_chips_100g`. Storing it is safe.
+- **Scrape path** — `catalog_engine.py:695` mints ids with
+ `s3_service.generate_image_id()`, which appends **`uuid4()[:8]`**
+ (`s3_service.py:59`). It is not deterministic. It reuses an existing id only
+ when `get_existing_product_image_id()` finds a row whose `product_name` is an
+ **exact string match** (`vector_store.py:489`) — no size in the lookup, no
+ normalisation, no fuzzy match.
+
+You can see both forms side by side in the live catalogue today:
+
+```
+scraped id 27 Cheetos Chips 250g image_id cheetos_chips_2d6bf74f
+scraped id 6 Kurkure Menthol 50g image_id kurkure_menthol_49ef2d35
+uploaded id 731 Kurkure Masala Munch 90g image_id pepsico_kurkure_masala_munch_90g
+```
+
+The `2d6bf74f` and `49ef2d35` are random. So on a scrape, if the product name
+changes by a single character — which is exactly what a brand prefix appearing
+or disappearing does — the reuse lookup misses, a **new random id** is minted,
+and `cleanup=True` deletes the old row. That is the full mechanism behind your
+eleven broken links, and `image_id` alone does not protect you from it.
+
+**What this means for your migration.** Storing `image_id` instead of our row id
+is still the right move — it is strictly better than the row id, and it is stable
+for everything that arrives through the upload path. But it is **not yet the
+durable key we implied** for scraped brands. Making it one means giving the
+scrape path the same deterministic `build_image_id()` the upload path uses. We
+have not done that, because it renames ids on the next scrape of every scraped
+brand, which is the same migration problem as 06b — it needs the old-to-new map
+in 06e, and it needs your timing.
+
+We are sorry for the incorrect answer; you told us you had switched on the
+strength of it.
+
+## 01c — `superseded_by`, a retired list, or a per-run changelog
+
+**Not built.** There is no `retired_at`, `superseded_by` or changelog anywhere in
+the codebase. Our questions from the 01 Sep reply still stand: would a
+`retired_at` timestamp with exclusion from the default read work instead of the
+row being deleted, and should retired rows stay reachable through an explicit
+query or vanish from the API entirely?
+
+---
+
+## 02 — Per-row rejection reasons — fulfilled
+
+`rejections[]` is populated at `store_catalog_pipeline.py:973` and carries
+exactly the shape you asked for:
+
+```jsonc
+"rejections": [
+ { "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
+ "reason": "title is too short to be a real product name" }
+]
+```
+
+- `row` is the 1-based sheet row with the header counted as row 1 —
+ `row_no = position + 2` (`store_catalog_pipeline.py:878`), the same convention
+ as the 422 responses, so it matches what the operator sees. `null` only when
+ the row cannot be located.
+- Capped at 50 per file (`as_dict`, line 188).
+- Present on **both** the single-batch read and the list endpoint: `slim=True`
+ strips only `products` (`batch_common.py:271`).
+- Regression test: `test_the_rejection_row_matches_the_offending_sheet_line`.
+- Documented in `INGESTION_API.md` with a field table.
+
+One correction to our own earlier account: the array itself was always in the
+response, and only `row` is new. It was missing from our docs, which is why you
+could not find it.
+
+## 03 — `brand_key` — fulfilled
+
+Every `products[]` entry carries `brand_key` beside `brand`
+(`store_catalog_pipeline.py:743`). It is produced by `_sanitize_name()`, the
+same function the storage layer uses to name the table, so it cannot drift from
+the key the catalogue is addressed by.
+
+Test `test_a_product_carries_the_key_the_catalogue_is_addressed_by` asserts
+`product["brand_key"] == _sanitize_name(product["brand"])`.
+
+## 04 — `from_drop` — fulfilled, and live
+
+`BatchFileOut.from_drop` (`batch_common.py:216`) is stamped by `_stamp_origins`
+at both `from-inbox` call sites (`batch_catalog.py:509,518`), written to disk,
+and re-read from the manifest. `null` for a file that went straight into a run
+without sitting in an inbox.
+
+**Confirmed deployed** — `from_drop` is in the live `openapi.json` served by
+`mcp.nearle.ai.in` today.
+
+Two regression tests cover the collision you described:
+`test_a_run_says_which_drop_each_file_came_from` stages two drops from different
+senders both named `a.csv` and asserts they are distinguishable, and
+`test_from_drop_survives_a_reread_of_the_run` proves it reaches disk.
+
+## 05 — `source_row` — fulfilled
+
+Every `products[]` entry carries `source_row`
+(`store_catalog_pipeline.py:751`), the 1-based sheet row with the header as row
+1. It is deliberately many-to-one: a pack-size cell reading `100g, 200g, 500g`
+becomes three products that all report the same `source_row`. Rows that produced
+nothing are those absent from every entry, so "rows 6 and 11 produced nothing"
+is now a set difference rather than a name-matching heuristic.
+
+The map is built from `kept` before dedupe and carried alongside the storage
+rows rather than inside them, so it describes the product actually written.
+Tests cover both the one-to-one and the exploded case.
+
+---
+
+## 06 — Duplicate brands, duplicate products, stray names
+
+### 06a — Haldiram merge: done, and live
+
+Live today: `haldiram` returns **0 products**, `haldirams` returns **2**. The
+alias is in the registry (`brand_registry.py:262-263`), so `Haldiram`,
+`haldiram`, `HALDIRAM` and `Haldiram's` all now resolve to `haldirams` — the
+merge cannot be undone by the next sheet that spells it without the "s".
+
+Both surviving rows also had their FSSAI licence corrected: they were carrying
+Lion Dates' number rather than Haldiram's.
+
+### 06b — The PepsiCo duplicate pairs: still there
+
+Confirmed live today, unchanged:
+
+```
+661 PepsiCo Lays Classic Salted 52g pepsico_pepsico_lays_classic_salted_52g
+730 Lays Classic Salted 52g 150 pepsico_lays_classic_salted_52g_150
+662 PepsiCo Kurkure Masala Munch 90g pepsico_pepsico_kurkure_masala_munch_90g
+731 Kurkure Masala Munch 90g pepsico_kurkure_masala_munch_90g
+```
+
+Deliberately untouched. De-duplicating means renaming, `image_id` is derived
+from the name, and renaming mints a *third* id rather than merging two — which
+would re-break the links you have just finished repairing. We need 06e first.
+
+Our lean on the convention is **without** the brand prefix, since brand is
+already its own column. We will follow whichever you prefer.
+
+### 06c — The stray number: cause fixed, rows not yet repaired
+
+**The cause is fixed, on two layers, both verified today:**
+
+- `Quantity` was removed from the `size_variants` keyword rule
+ (`user_products.py:176-182`); only `net qty` / `net quantity`, the Indian
+ labelling term for a real pack size, still match. Verified: a sheet with
+ columns `Product Name, Quantity, Pack Size, Brand` now binds `Pack Size` and
+ reports `Quantity` as unrecognised. Previously a leading `Quantity` column
+ also shut out the sheet's real `Pack Size` column, because mapping is
+ first-wins by position.
+- `_sizes_for()` discards a unitless number and records why
+ (`store_catalog_pipeline.py:428`): *"ignored pack size '150': a number with no
+ unit is a quantity, not a size"*.
+
+No new row can acquire this. **The existing rows are still there — 17 of them,
+scanned across all 55 live brands today** (we said 18 on 01 Sep; the 18th was
+the corrupted Haldiram row removed in the merge):
+
+```
+brooke_bond 1 Red Label Tea 500g 30
+fortune 2 Fortune Sunflower Oil 1L 48
+hindustan_unilever 2135 Tata Salt Iodised 1kg 120
+hindustan_unilever 2136 Bru Instant Coffee 100g 24
+hindustan_unilever 2137 Horlicks Classic Malt 500g 18
+hindustan_unilever 2138 Surf Excel Easy Wash 1kg 36
+hindustan_unilever 2139 Vim Dishwash Bar 300g 90
+hindustan_unilever 2140 Dove Cream Beauty Bar 100g 64
+hindustan_unilever 2141 Clinic Plus Shampoo 175ml 40
+india_gate 1 India Gate Basmati Rice 1kg 60
+britannia 7 Britannia Good Day Cashew 200g 60
+colgate_palmolive 483 Colgate Strong Teeth 200g 50
+itc 5 Aashirvaad Shudh Chakki Atta 5kg 40
+reckitt_benckiser 1 Harpic Power Plus 500ml 30
+reckitt_benckiser 2 Dettol Original Soap 125g 80
+coca_cola 1041 Coca-Cola 750ml 72
+pepsico 730 Lays Classic Salted 52g 150
+```
+
+Repairing them is a rename, so it is blocked on 06e exactly as 06b is.
+
+Unrelated but visible in that list: `Tata Salt Iodised` is filed under
+`hindustan_unilever`. Tata Salt is not an HUL product. We are looking at it.
+
+### 06d — Britannia, Parle, Patanjali
+
+**Confirmed incomplete scrapes, not small brands.** Still `6`, `3` and `3`
+products live today — the re-scrape has not been run yet.
+
+### 06e — What we need back from you
+
+Before we rename anything (06b, 06c, and the `image_id` determinism fix in 01b):
+
+1. Which prefixing convention wins — brand in the product name, or not?
+2. Would an **old-id to new-id map, delivered in advance for every row we touch**,
+ let you re-point rather than clear? It is straightforward for us to produce.
+3. Timing, so all of it lands in one pass rather than trickling.
+
+---
+
+## 07 — Loose produce — fulfilled, and live
+
+**159 rows, live now** under brand key `own_products`, in the shape you asked
+for: name and category, no brand, no pack size, no price.
+
+```
+Fruits & Vegetables 95 Flowers 14
+Fresh Herbs & Greens 17 Dairy (loose) 13
+Fish & Seafood 15 Eggs 5
+```
+
+Verified live: `getproducts?brand=own_products` returns them, e.g.
+`own_products_apple_standard`, with `highlights: ["Loose / unbranded", ...]`.
+**112 of the 159 (70%) carry an image** — all the fruit, vegetables, greens and
+herbs. Flowers, fish, loose dairy and eggs are still name-only. Tell us whether
+you want the list complete-but-later or partial-but-now.
+
+Each row carries a search embedding, so these reach semantic search and not just
+exact match, and every seeded row classifies identically to how an uploaded copy
+of the same name would — a grocer typing "Tomato" lands on the seeded row rather
+than creating a second one.
+
+**The underlying bug was worse than a coverage gap**, and is fixed: produce rows
+were not rejected, they were *misfiled*. The brand fallback took the first word
+of the name and matched it against the alias map, so `Apple` became brand
+"Apple", `Curry Leaves` became "Curry", and **`Red Rose` was being written into
+the Brooke Bond tea catalogue** and stamped with that brand's real FSSAI licence.
+Loose goods are now recognised as commodities and filed under `Own Products`
+before brand inference runs.
+
+Verified against your own audit strings today:
+
+```
+Apple, Tomato, Curry Leaves, Red Rose, Thulasi, Drumstick, Tuna -> unbranded
+Bitter guard, Bottle ground, Ladies Finger (your misspellings) -> unbranded
+```
+
+Produce is also exempt from the invented-pack-size fallback that causes 01a: an
+`Own Products` row with no weight gets one `Standard` row, not three made-up
+ones (`store_catalog_pipeline.py:460`).
+
+**Two limitations, unchanged and worth knowing:**
+
+- Place-qualified produce still reads as branded: `Salem Mango`, `Mysore Banana`
+ and `Jammu Apple` — all real strings from your Ragul Stores data — classify as
+ branded, because `Mysore` is also a real brand. The classifier is deliberately
+ conservative: collapsing a real regional brand into the unbranded bucket is
+ much harder to undo than a mango in the wrong table. The workaround is in the
+ pipeline already — **if the sheet has a brand column and leaves the cell
+ empty, we believe it** and file the row under Own Products regardless of name.
+- `Maceral` (your Ragul Stores spelling of mackerel) does not resolve. `Mackerel`
+ does. Send us any other spellings your merchants actually use and we will add
+ them.
+
+---
+
+## Reconciling your 1,614 against our count
+
+You counted 55 brands / 1,614 products. We now measure **55 brands / 1,572
+products**, of which 159 are the new produce rows — so **1,413 of the old kind**.
+
+Your number and ours differ by 201, and it resolves exactly:
+
+```
+your count 1,614
+ less Hindustan Unilever, 443 -> 243 today -200
+ less the Haldiram row removed in the merge -1
+ -------
+ 1,413 = our non-produce count
+ plus the produce base list +159
+ -------
+ 1,572 = live today
+```
+
+**200 Hindustan Unilever products have been deleted since you measured.** That is
+not a counting discrepancy — it is finding 01a happening again, to the largest
+brand in the catalogue, between your report and this reply. It is the strongest
+argument we have for the retirement model in 01c, and it is why we would like an
+answer on 01c sooner than on the rest.
+
+---
+
+## What we checked before writing this
+
+- Every figure re-measured today against the live catalogue API, not read off
+ source; the `image_id` behaviour read out of the two functions that mint it.
+- The ids you listed are still absent: pepsico `4, 19, 20, 25, 26`, dabur
+ `19, 20`, nestle `1`. Clearing those links rather than re-pointing them by
+ name remains the correct call.
+- Stray-number scan run across all 55 brands and all 1,572 live products.
+- Full backend suite: **1,108 tests passing.**
diff --git a/tests/test_review_inbox.py b/tests/test_review_inbox.py
index 6e376bf..b50fadc 100644
--- a/tests/test_review_inbox.py
+++ b/tests/test_review_inbox.py
@@ -11,11 +11,14 @@ says so, and every assertion below is ultimately about one of two failures -
* something running that nobody approved, and
* a file being lost, or run twice, on its way out of the inbox.
-The response SHAPES are asserted literally rather than loosely, because
-frontend/src/pages/InboxPanel.jsx was written against this contract before the
-backend existed. `submission_id`, `file_id`, `pending_count` and the `dismissed`
-count are read by name there; a rename that only this file catches is cheap, and
-one that nothing catches is a blank admin tab with no error.
+The response SHAPES are asserted literally rather than loosely, because the
+frontend was written against this contract before the backend existed.
+`submission_id`, `file_id`, `pending_count` and the `dismissed` count are read
+by name in frontend/src/pages/UploadResultsPanel.jsx - the approval gate lives
+inside that panel now, revealed only when `pending_count` is above zero, since
+UPLOAD_AUTORUN=true leaves the inbox permanently empty. A rename that only this
+file catches is cheap, and one that nothing catches is a blank admin tab with no
+error.
"""
from __future__ import annotations