New updates on DB and JSON

This commit is contained in:
sriram
2026-09-01 13:55:15 +05:30
parent 183b65b3bd
commit 6c7a886659
20 changed files with 68753 additions and 135 deletions

View File

@@ -176,7 +176,7 @@ The bad file is kept as a failed member rather than dropped, so a sender who sub
"brands": [],
"files": [
{ "index": 0, "filename": "catalog.csv", "status": "queued",
"total_stages": 11, "rows_total": 1, "size_bytes": 57 },
"from_drop": null, "total_stages": 11, "rows_total": 1, "size_bytes": 57 },
{ "index": 1, "filename": "notes.txt", "status": "failed",
"detail": "The file has no data rows." }
],
@@ -184,6 +184,22 @@ The bad file is kept as a failed member rather than dropped, so a sender who sub
}
```
#### `from_drop` — which file in this run is yours
**Match on this, never on `filename`.** An admin can assemble one run from several
drops, so a run's `files` may contain sheets you did not send — and two senders can
both upload `products.csv`. Matching on the name is a coincidence; matching on
`from_drop` is exact.
| Value | Meaning |
| --- | --- |
| the drop id you were given | this file is the one you sent in that drop |
| `null` | the file went straight into a run and never sat in an inbox — under `UPLOAD_AUTORUN=true` that is every file, and the run id you hold is already the only id involved |
It is the exact inverse of `released_to`, which points from your drop to the run
that took it. Both are needed: `released_to` answers *where did my drop go*,
`from_drop` answers *whose file is this*.
`use_llm` and `fetch_images` are reported, never accepted. They decide how much
outbound work a run commits the host to, and this endpoint's caller is anonymous, so
they come from settings — sending them in the request has no effect.
@@ -333,19 +349,112 @@ catalogue:
"rows_total": 2, "inserted": 1, "backfilled": 0, "skipped_existing": 1,
"rejected": 0, "brands": ["amul"],
"products": [
{ "image_id": "amul_amul_butter_100g", "brand": "amul",
"product_name": "Amul Butter 100g",
{ "image_id": "amul_amul_butter_100g", "brand": "amul", "brand_key": "amul",
"product_name": "Amul Butter 100g", "source_row": 2,
"product_sku": "ACME-BUT-100", "sku_source": "sheet",
"disposition": "inserted" },
{ "image_id": "amul_amul_ghee_1l", "brand": "amul",
"product_name": "Amul Ghee 1L",
{ "image_id": "amul_amul_ghee_1l", "brand": "amul", "brand_key": "amul",
"product_name": "Amul Ghee 1L", "source_row": 3,
"product_sku": "AMUL-GHE-1-001", "sku_source": "Internal",
"disposition": "unchanged" }
],
"products_truncated": false
"products_truncated": false,
"rejections": [
{ "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
"reason": "title is too short to be a real product name; image_urls: no images were found for this product" }
]
}
```
#### `image_id` — the join key, and what it is stable against
`image_id` is a pure deterministic function of **brand, product name and pack size**.
Re-sending an unchanged sheet produces byte-identical ids, which is what makes the
pipeline idempotent, and it is the column the catalogue deduplicates on. Store it
rather than a row id.
What it is *not* stable against is any change to those three inputs. A pack size
moving from `100g` to `250g` is a different SKU at a different price and is correctly
a different id; so is a product name gaining or losing a brand prefix
(`Hot Heads` vs `Nestle Hot Heads`). If a name is rewritten upstream, the id moves
with it.
#### `brand_key` — the key the catalogue is addressed by
`brand` is the display name; `brand_key` is the identifier the catalogue is keyed
on, and the two are not the same string:
| `brand` | `brand_key` |
| --- | --- |
| `24 Mantra` | `24_mantra` |
| `Paper Boat` | `paper_boat` |
| `coca-cola` | `coca_cola` |
| `Own Products` | `own_products` |
Use `brand_key` rather than normalising `brand` yourself. The rule (lower-case,
non-alphanumerics to underscores) is stable, but deriving it is a guess and the
failure is silent — a wrong key finds nothing rather than erroring.
#### Unbranded rows — what `Own Products` means, and what it does not fill in
A row with no brand — loose fruit, vegetables, greens, flowers, fish, or staples
like dal and sugar — is filed under the display brand **`Own Products`**
(`brand_own_products`) rather than having a brand guessed from its first word.
**For these rows we store what your sheet said and nothing more.** A shopkeeper
bills from this record, so a tax code or article number we invented would be our
guess wearing your letterhead:
| Field | For an unbranded row |
| --- | --- |
| `product_name`, `size_variants` | from your sheet |
| `selling_price`, `final_selling_price` | from your sheet |
| `price_range` | a ±8% band around your price. Null if you sent no price |
| `category` | derived from the product name — deterministic, not guessed |
| `image_url`, `image_urls` | searched for, as with any other row |
| `hsn_code`, `product_sku`, `barcode`, `fssai_license`, `description` | **null**, unless your sheet supplied them |
Anything you *do* send is kept: a sheet with its own HSN, SKU or description
column keeps all three. The rule is "we do not invent", not "we discard".
Two consequences worth planning for:
- **No pack-size explosion.** A branded row with no size gets a plausible set
(100g/250g/500g); an unbranded one does not, because a shop sells apples by
whatever the customer asks for. A produce row with no weight column produces
exactly one entry, with size `Standard`.
- **`validation_status` is often `needs_review`.** For these rows that reflects
a missing image, not a suspect product — the absence of price or SKU is no
longer counted against them. `needs_review` rows are stored like any other;
only `rejected` rows are dropped.
#### `source_row` — which line of the sheet produced this
The 1-based row number as the sender sees it on screen, **header counted as row 1**,
so the first data row is `2`. The same convention the `422` responses use.
This is **many-to-one**: a pack-size cell reading `100g, 200g, 500g` legitimately
becomes three products, and all three carry the same `source_row`. Rows that produced
nothing are the ones absent from every entry — which is how you tell a shopkeeper
"rows 6 and 11 produced nothing".
#### `rejections[]` — which rows were refused, and why
`rejected` is a count; `rejections` is the explanation, and it has always been sent.
One entry per refused row:
| Field | Meaning |
| --- | --- |
| `row` | The 1-based sheet row, same convention as `source_row`. `null` if the row could not be located |
| `product_name` | The name as the sheet gave it |
| `size` | The pack the refusal applies to, since one row can yield several |
| `reason` | Every validation issue, joined with `; ` |
Capped at 50 entries per file. Present on both the single-batch read and the list
endpoint — only `products` is dropped from list responses.
#### `disposition`
| `disposition` | What happened |
| --- | --- |
| `inserted` | New product, created by this run |