New updates on DB and JSON

This commit is contained in:
sriram
2026-09-01 13:55:15 +05:30
parent 183b65b3bd
commit 6c7a886659
20 changed files with 68753 additions and 135 deletions

View File

@@ -0,0 +1,323 @@
# Response to the Catalogue Drift Report
Reply to the seven findings dated 31 August 2026 (drop
`8e1448e176d843d08183d387ad724f95` → run `0ed4c77b0ca14e03b1aaf6b5b77d1994`).
Every figure below was measured against the same live deployment, not read off
source. Where we disagree with a finding, the evidence is included so you can
check it rather than take our word for it.
**Summary:** four items are fixed and ship in the next backend deploy. One
(#02) was already in the API and we had failed to document it — that is our
fault and the docs are now corrected. #01 is diagnosed, and the cause is not
what either of us assumed. #06 is confirmed but carries a trap that means we
should agree an approach before touching it.
| # | Finding | Status |
| --- | --- | --- |
| 01 | Pack sizes replaced between scrapes | **Diagnosed** — three causes, not one. Fix needs your input |
| 02 | `rejected` is a bare count | **Already shipped, now documented** — plus `row` added |
| 03 | Manifest brands are not catalogue keys | **Fixed** — `brand_key` published |
| 04 | Run files carry no `from_drop` | **Fixed** |
| 05 | Manifest carries no `source_row` | **Fixed** |
| 06 | Duplicate brands and products | **Confirmed.** Read the trap below before we act |
| 07 | No loose-produce coverage | **Fixed** — 159-row base list, and the upload path now handles produce |
---
## 01 — Pack sizes and names are replaced between scrapes
You asked us to confirm whether a pack size that once existed is meant to
survive a re-scrape. **It is not, today** — but that is only the last of three
causes, and fixing it alone would not have saved your links.
### Cause 1: pack sizes are invented when a scrape does not supply them
`_sizes_for()` falls back to `default_size_variants(category, name)` when a
product declares no size. That fallback is **keyed on the resolved category**,
and the resolved category is not stable between runs. Measured today:
```
category "Snacks" -> ['55g', '150g', '200g']
category "" (unresolved) -> ['100g', '250g', '500g']
category "Breakfast Cereal" -> ['250g', '500g', '1kg']
```
Now compare your table. Cheetos id 25 is **55g** — the *Snacks* set. Ids 26 and
27 are **100g** and **250g** — the *unresolved* set. The same product was
ingested once with its category resolved and once without, and produced two
disjoint sets of pack sizes. Your Cheerios ids (100g, 250g, 500g) sit across the
Breakfast Cereal set and the unresolved set the same way.
So the pack sizes were never scraped facts that changed. Some of them were
generated, and the generator's input moved.
### Cause 2: the product name gains or loses a brand prefix
Your own #06 has the evidence: `Hot Heads 30g` became `Nestle Hot Heads`, and
`PepsiCo Kurkure Masala Munch 90g` coexists with `Kurkure Masala Munch 90g`.
`image_id` is derived from the name, so a prefix appearing or disappearing moves
the id even when the product is identical.
### Cause 3: the write then deletes whatever is not in the new set
The brand-scrape path calls `upsert_brand_products(..., cleanup=True)`, which
deletes every row in the table whose `image_id` is absent from the batch being
written. That is what turns causes 1 and 2 from "duplicate rows" into "the row
you stored is gone".
**Any one of these is survivable. Together they guarantee broken links on every
re-scrape**, which matches your finding that zero of eleven could be repaired.
### Not the upload path
Worth stating plainly, because it affects how much you need to worry: the
**upload** path — everything reached through `POST /api/uploads/catalog` — uses
`cleanup=False` and has always done so. A sheet you send can never delete a row
it does not mention. The deletions came from brand scraping only.
### What we need from you
The real fix is to stop causes 1 and 2 (do not invent sizes for a product
already in the catalogue; settle the naming convention), and to soft-retire
rather than delete for cause 3. That third part changes how the production
catalogue is written and we would rather agree it with you than spring it:
- Would a `retired_at` timestamp plus exclusion from the default read work for
you, instead of the row being deleted? That preserves the `image_id` so your
stored link resolves to something, and lets us give you the `superseded_by`
and per-run changelog you asked for.
- If so, do you want retired rows visible through an explicit query, or gone
from the API entirely?
### Your two direct questions
**Is `image_id` stable across re-scrapes for a product whose name and pack size
have not changed?** Yes. It is a pure deterministic function of brand, product
name and pack size, with no clock, counter or run id in it. Verified:
```
build_image_id('pepsico', 'Cheetos Chips', '100g') -> pepsico_cheetos_chips_100g
```
Storing it rather than our row id is the right call and we have documented the
guarantee so it does not quietly change. **The caveat is #06:** the guarantee is
only as good as the stability of the name, and inconsistent brand prefixing
breaks exactly that.
**Do you want to know when a product is dropped or renamed?** Yes, and we agree
it should exist. It falls out of the retirement model above rather than being a
separate feature, which is why we would like to settle that first.
---
## 02 — `rejected` is a count with no reason
**This is our documentation failure, not a missing feature.** `rejections[]` has
been in every response — single-batch read and list endpoint both, since
`to_out(slim=True)` strips only `products` — carrying `product_name`, `size` and
`reason` per refused row. It was absent from `INGESTION_API.md`, which documents
`"rejected": 0` and never mentions the array, so there was no way for you to
know it was there. Sorry — that is a straightforwardly bad docs bug.
The genuine gap was the row number, which is now added:
```jsonc
"rejections": [
{ "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
"reason": "title is too short to be a real product name; image_urls: no images were found for this product" }
]
```
`row` is the 1-based sheet row with the header counted as row 1 — the same
convention as the `422` responses, so it matches what the operator sees on
screen. `null` only when the row cannot be located. Capped at 50 per file.
---
## 03 — Manifest brands are not catalogue keys
Fixed. Every entry in `products[]` now carries `brand_key` beside `brand`:
```jsonc
{ "brand": "24 Mantra", "brand_key": "24_mantra", ... }
```
This is generated by the same function the storage layer uses to name the table,
so it cannot drift from the key the catalogue is actually addressed by. Your
normalisation is correct as far as we can tell, but it is a guess, and the
failure mode is silent — a wrong key finds nothing rather than erroring.
Thank you for degrading rather than failing the batch on an unreadable brand;
that is the right behaviour and we should have done it on our side too.
---
## 04 — Run files carry no `from_drop`
Fixed, and we agree with your assessment that this was the one item that could
corrupt a merchant's inventory rather than merely inconvenience you.
Each file in a run now carries `from_drop`, the id of the drop it was released
from — the exact inverse of `released_to`:
```jsonc
"files": [
{ "index": 0, "filename": "products.csv", "from_drop": "8e1448e1...", ... },
{ "index": 1, "filename": "products.csv", "from_drop": "a91c02f4...", ... }
]
```
`null` for a file that went straight into a run without sitting in an inbox —
which, under `UPLOAD_AUTORUN=true`, is every file you send, because the id you
are handed is already the run.
There is a test in our suite that stages two drops from different senders both
named `products.csv` and asserts they are distinguishable, so the collision you
described is now a permanent regression guard rather than a hope.
---
## 05 — Manifest cannot be traced back to the spreadsheet row
Fixed. Every entry in `products[]` carries `source_row`, the 1-based sheet row
with the header as row 1.
It is deliberately **many-to-one**: a pack-size cell reading `100g, 200g, 500g`
becomes three products that all report the same `source_row`, which is what lets
you say "row 14 of your sheet became these three". Rows that produced nothing
are those absent from every entry — "rows 6 and 11 produced nothing" is now a
set difference rather than a name-matching heuristic.
---
## 06 — Duplicate brands, duplicate products, stray names
Confirmed against the live database today: **55 brand tables, 1,414 products**
(you counted 1,614; the difference is a day of drift plus, we think, your count
including rejected rows — worth reconciling if it matters).
```
haldiram 1 britannia 6
haldirams 2 parle 3 against hindustan_unilever 443
patanjali 3
```
So: the split brand is real, and Britannia/Parle/Patanjali do look like scrapes
that stopped part-way rather than genuinely small brands. We will re-run those
three.
### The trap, which is why we have not just fixed this
**De-duplicating the PepsiCo pairs means renaming a product, and `image_id` is
derived from the name.** Renaming `PepsiCo Kurkure Masala Munch 90g` to
`Kurkure Masala Munch 90g` does not merge the two rows — it mints a third id and
breaks any link pointing at either of the first two. You have just migrated onto
storing `image_id`. A well-meant cleanup on our side would re-break exactly what
you have finished repairing.
The same applies to stripping the stray `150` from `Lays Classic Salted 52g 150`.
So before we touch it we would like to agree:
1. **Which convention wins** — brand prefix in the product name, or not? We have
no strong preference; we care only that it is one of them. Our lean is
*without* the prefix, since the brand is already a column.
2. **How the merge is communicated.** If we can give you the old-id →
new-id mapping for every row we touch, in advance, does that let you
re-point rather than clear? That is straightforward for us to produce.
3. **Timing**, so it lands in one pass rather than trickling.
`haldiram` → `haldirams` is a three-row merge and much lower risk; we can do
that one immediately if you would rather not wait for the rest.
---
## 07 — Loose produce has no coverage
Fixed, and this turned out to be the most valuable finding in your report,
because it was not only a coverage gap.
### What was actually happening
Produce rows were not rejected. They were **misfiled**, which is worse. The
brand fallback takes the first word of the name and then whole-word matches it
against our alias map:
```
Apple -> brand "Apple" -> junk table brand_apple
Tomato -> brand "Tomato" -> junk table brand_tomato
Bitter Gourd -> brand "Bitter" -> junk table brand_bitter
Curry Leaves -> brand "Curry" -> junk table brand_curry
Red Rose -> brand "Red" -> brand_brooke_bond <--
```
That last one is not a typo. A rose was being written into the Brooke Bond tea
catalogue, and our enrichment then stamps that brand's real FSSAI licence number
onto the row. Your 139 hand-typed products were the visible symptom; this was
underneath it.
### What now happens
Loose goods are recognised as commodities and filed under a single house brand,
`Own Products` (table `brand_own_products`), before brand inference can touch
them. Fruit, vegetables, greens, herbs, flowers, fish, eggs and loose dairy are
covered, alongside the pulses, grains, spices, oils and sugar that already were.
Five categories were added — Fruits & Vegetables, Fresh Herbs & Greens, Flowers,
Fish & Seafood, Eggs — with HSN codes and a 0% GST rate, since unprocessed
produce is nil-rated rather than reduced-rate.
Merchant misspellings from your own data are handled: `Bitter guard`,
`Bottle ground`, `Ladies Finger` all resolve.
### The base list
**159 rows**, in the shape you asked for: name and category only, no brand, no
pack size, no price. Fruit (42), vegetables (53), greens and herbs (17), flowers
(14), fish and seafood (15), loose dairy (13), eggs (5). It includes the specific
items your audit listed — Jasmine, Lotus, Red Rose, Thulasi, Drumstick, Curry
Leaves, the four banana varieties, Tuna, Mackerel.
Each row carries a `search_query` embedding, so these are reachable through
semantic search and not just exact match. The list is hand-authored rather than
scraped, so it is not subject to any of #01.
**Images are not included yet.** You asked for name and image; we have shipped
the names. Sourcing 159 licensable produce photographs is a separate piece of
work and we did not want to hold the list for it — tell us if the list is not
useful to you without them and we will prioritise accordingly.
### One limitation worth knowing
The classifier is deliberately conservative: a single word it does not recognise
means "this is a brand". So place-qualified produce — `Salem Mango`,
`Mysore Banana`, `Jammu Apple`, all real strings from your Ragul Stores data —
still reads as branded, because `Mysore` is also a real brand (Mysore Sandal).
The workaround is already in the pipeline: **if the sheet has a Brand column and
leaves the cell empty, we believe it** and file the row under Own Products
regardless of the name. If your merchants' sheets carry an empty brand column,
those rows will land correctly. If they carry no brand column at all, the
name-based test is what applies.
We would rather be conservative here. Collapsing a real regional brand into the
unbranded bucket is much harder to undo than a mango sitting in the wrong table.
---
## What we verified before sending this
- The produce lexicon was run against **all 1,414 products in the live
catalogue** and against all **231 brand aliases**: zero reclassifications in
either. That check is now a test, so it runs on every change.
- It also surfaced a pre-existing bug we would not otherwise have found: our
pack-size stripper was eating the word after a number, so
`24 Mantra Organic Moong Dal 500g` lost its brand entirely and was being filed
as an unbranded commodity. **Every brand whose name starts with a digit hit
this.** "24 Mantra" appears on your finding-03 list, which is how we noticed.
Fixed.
- All 159 seeded rows were checked to classify identically to how an uploaded
copy of the same name would, so a grocer typing "Tomato" lands on the seeded
row instead of creating a second one.
- Full suite: **1,054 tests passing.**

View File

@@ -176,7 +176,7 @@ The bad file is kept as a failed member rather than dropped, so a sender who sub
"brands": [],
"files": [
{ "index": 0, "filename": "catalog.csv", "status": "queued",
"total_stages": 11, "rows_total": 1, "size_bytes": 57 },
"from_drop": null, "total_stages": 11, "rows_total": 1, "size_bytes": 57 },
{ "index": 1, "filename": "notes.txt", "status": "failed",
"detail": "The file has no data rows." }
],
@@ -184,6 +184,22 @@ The bad file is kept as a failed member rather than dropped, so a sender who sub
}
```
#### `from_drop` — which file in this run is yours
**Match on this, never on `filename`.** An admin can assemble one run from several
drops, so a run's `files` may contain sheets you did not send — and two senders can
both upload `products.csv`. Matching on the name is a coincidence; matching on
`from_drop` is exact.
| Value | Meaning |
| --- | --- |
| the drop id you were given | this file is the one you sent in that drop |
| `null` | the file went straight into a run and never sat in an inbox — under `UPLOAD_AUTORUN=true` that is every file, and the run id you hold is already the only id involved |
It is the exact inverse of `released_to`, which points from your drop to the run
that took it. Both are needed: `released_to` answers *where did my drop go*,
`from_drop` answers *whose file is this*.
`use_llm` and `fetch_images` are reported, never accepted. They decide how much
outbound work a run commits the host to, and this endpoint's caller is anonymous, so
they come from settings — sending them in the request has no effect.
@@ -333,19 +349,112 @@ catalogue:
"rows_total": 2, "inserted": 1, "backfilled": 0, "skipped_existing": 1,
"rejected": 0, "brands": ["amul"],
"products": [
{ "image_id": "amul_amul_butter_100g", "brand": "amul",
"product_name": "Amul Butter 100g",
{ "image_id": "amul_amul_butter_100g", "brand": "amul", "brand_key": "amul",
"product_name": "Amul Butter 100g", "source_row": 2,
"product_sku": "ACME-BUT-100", "sku_source": "sheet",
"disposition": "inserted" },
{ "image_id": "amul_amul_ghee_1l", "brand": "amul",
"product_name": "Amul Ghee 1L",
{ "image_id": "amul_amul_ghee_1l", "brand": "amul", "brand_key": "amul",
"product_name": "Amul Ghee 1L", "source_row": 3,
"product_sku": "AMUL-GHE-1-001", "sku_source": "Internal",
"disposition": "unchanged" }
],
"products_truncated": false
"products_truncated": false,
"rejections": [
{ "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
"reason": "title is too short to be a real product name; image_urls: no images were found for this product" }
]
}
```
#### `image_id` — the join key, and what it is stable against
`image_id` is a pure deterministic function of **brand, product name and pack size**.
Re-sending an unchanged sheet produces byte-identical ids, which is what makes the
pipeline idempotent, and it is the column the catalogue deduplicates on. Store it
rather than a row id.
What it is *not* stable against is any change to those three inputs. A pack size
moving from `100g` to `250g` is a different SKU at a different price and is correctly
a different id; so is a product name gaining or losing a brand prefix
(`Hot Heads` vs `Nestle Hot Heads`). If a name is rewritten upstream, the id moves
with it.
#### `brand_key` — the key the catalogue is addressed by
`brand` is the display name; `brand_key` is the identifier the catalogue is keyed
on, and the two are not the same string:
| `brand` | `brand_key` |
| --- | --- |
| `24 Mantra` | `24_mantra` |
| `Paper Boat` | `paper_boat` |
| `coca-cola` | `coca_cola` |
| `Own Products` | `own_products` |
Use `brand_key` rather than normalising `brand` yourself. The rule (lower-case,
non-alphanumerics to underscores) is stable, but deriving it is a guess and the
failure is silent — a wrong key finds nothing rather than erroring.
#### Unbranded rows — what `Own Products` means, and what it does not fill in
A row with no brand — loose fruit, vegetables, greens, flowers, fish, or staples
like dal and sugar — is filed under the display brand **`Own Products`**
(`brand_own_products`) rather than having a brand guessed from its first word.
**For these rows we store what your sheet said and nothing more.** A shopkeeper
bills from this record, so a tax code or article number we invented would be our
guess wearing your letterhead:
| Field | For an unbranded row |
| --- | --- |
| `product_name`, `size_variants` | from your sheet |
| `selling_price`, `final_selling_price` | from your sheet |
| `price_range` | a ±8% band around your price. Null if you sent no price |
| `category` | derived from the product name — deterministic, not guessed |
| `image_url`, `image_urls` | searched for, as with any other row |
| `hsn_code`, `product_sku`, `barcode`, `fssai_license`, `description` | **null**, unless your sheet supplied them |
Anything you *do* send is kept: a sheet with its own HSN, SKU or description
column keeps all three. The rule is "we do not invent", not "we discard".
Two consequences worth planning for:
- **No pack-size explosion.** A branded row with no size gets a plausible set
(100g/250g/500g); an unbranded one does not, because a shop sells apples by
whatever the customer asks for. A produce row with no weight column produces
exactly one entry, with size `Standard`.
- **`validation_status` is often `needs_review`.** For these rows that reflects
a missing image, not a suspect product — the absence of price or SKU is no
longer counted against them. `needs_review` rows are stored like any other;
only `rejected` rows are dropped.
#### `source_row` — which line of the sheet produced this
The 1-based row number as the sender sees it on screen, **header counted as row 1**,
so the first data row is `2`. The same convention the `422` responses use.
This is **many-to-one**: a pack-size cell reading `100g, 200g, 500g` legitimately
becomes three products, and all three carry the same `source_row`. Rows that produced
nothing are the ones absent from every entry — which is how you tell a shopkeeper
"rows 6 and 11 produced nothing".
#### `rejections[]` — which rows were refused, and why
`rejected` is a count; `rejections` is the explanation, and it has always been sent.
One entry per refused row:
| Field | Meaning |
| --- | --- |
| `row` | The 1-based sheet row, same convention as `source_row`. `null` if the row could not be located |
| `product_name` | The name as the sheet gave it |
| `size` | The pack the refusal applies to, since one row can yield several |
| `reason` | Every validation issue, joined with `; ` |
Capped at 50 entries per file. Present on both the single-batch read and the list
endpoint — only `products` is dropped from list responses.
#### `disposition`
| `disposition` | What happened |
| --- | --- |
| `inserted` | New product, created by this run |