Feature Updation

This commit is contained in:
sriram
2026-09-02 12:25:36 +05:30
parent 1de518feed
commit 5399fea4cc
3 changed files with 1003 additions and 5 deletions

View File

@@ -0,0 +1,644 @@
<title>Catalogue Drift Verification</title>
<link rel="preconnect" href="https://fonts.googleapis.com">
<link rel="preconnect" href="https://fonts.gstatic.com" crossorigin>
<link rel="stylesheet" href="https://fonts.googleapis.com/css2?family=Newsreader:ital,opsz,wght@0,6..72,400;0,6..72,500;0,6..72,600;1,6..72,400&family=Source+Sans+3:ital,wght@0,400;0,600;0,700;1,400&family=JetBrains+Mono:wght@400;700&display=swap">
<style>
:root {
--ground: #F6F7F9;
--surface: #FFFFFF;
--surface-alt: #EFF1F5;
--ink: #191D25;
--ink-soft: #3D4552;
--muted: #616B7B;
--rule: #DCE0E8;
--rule-strong: #C3C9D4;
--accent: #1F4E79;
--accent-soft: #E7EEF6;
--ok: #2C6E49;
--ok-bg: #E4F0E8;
--warn: #8A6212;
--warn-bg: #F7EDD8;
--bad: #9C2B2B;
--bad-bg: #F7E4E4;
--code-bg: #F1F3F7;
color-scheme: light;
}
@media (prefers-color-scheme: dark) {
:root:not([data-theme="light"]) {
--ground: #11141A;
--surface: #181C23;
--surface-alt: #1F242D;
--ink: #E5E9F0;
--ink-soft: #C2C9D4;
--muted: #929BA9;
--rule: #2B313B;
--rule-strong: #3B434F;
--accent: #86B4E0;
--accent-soft: #1B2836;
--ok: #7FC79A;
--ok-bg: #16281E;
--warn: #DCB463;
--warn-bg: #2B2313;
--bad: #E39191;
--bad-bg: #2E1919;
--code-bg: #1C2129;
color-scheme: dark;
}
}
:root[data-theme="dark"] {
--ground: #11141A;
--surface: #181C23;
--surface-alt: #1F242D;
--ink: #E5E9F0;
--ink-soft: #C2C9D4;
--muted: #929BA9;
--rule: #2B313B;
--rule-strong: #3B434F;
--accent: #86B4E0;
--accent-soft: #1B2836;
--ok: #7FC79A;
--ok-bg: #16281E;
--warn: #DCB463;
--warn-bg: #2B2313;
--bad: #E39191;
--bad-bg: #2E1919;
--code-bg: #1C2129;
color-scheme: dark;
}
* { box-sizing: border-box; }
body {
background: var(--ground);
color: var(--ink);
font-family: "Source Sans 3", ui-sans-serif, system-ui, -apple-system, "Segoe UI", sans-serif;
font-size: 16.5px;
line-height: 1.62;
-webkit-font-smoothing: antialiased;
}
.wrap {
max-width: 890px;
margin: 0 auto;
padding: 56px 28px 110px;
}
/* ---------- masthead ---------- */
.masthead {
border-bottom: 2px solid var(--ink);
padding-bottom: 22px;
margin-bottom: 34px;
}
.eyebrow {
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 11.5px;
letter-spacing: 0.13em;
text-transform: uppercase;
color: var(--muted);
margin: 0 0 14px;
}
h1 {
font-family: "Newsreader", ui-serif, Georgia, serif;
font-weight: 500;
font-size: clamp(2.1rem, 5.2vw, 3.05rem);
line-height: 1.08;
letter-spacing: -0.015em;
text-wrap: balance;
margin: 0 0 16px;
}
.standfirst {
font-family: "Newsreader", ui-serif, Georgia, serif;
font-size: 1.18rem;
line-height: 1.52;
color: var(--ink-soft);
margin: 0;
max-width: 62ch;
}
.meta-strip {
display: flex;
flex-wrap: wrap;
gap: 10px 26px;
margin-top: 22px;
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 12px;
color: var(--muted);
}
.meta-strip b { color: var(--ink-soft); font-weight: 700; }
/* ---------- verdict ---------- */
.verdict {
display: grid;
grid-template-columns: repeat(3, 1fr);
gap: 1px;
background: var(--rule);
border: 1px solid var(--rule);
margin: 0 0 40px;
}
.verdict div {
background: var(--surface);
padding: 18px 20px;
}
.verdict .n {
font-family: "Newsreader", ui-serif, Georgia, serif;
font-size: 2.3rem;
line-height: 1;
font-variant-numeric: tabular-nums;
display: block;
margin-bottom: 6px;
}
.verdict .l {
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 11px;
letter-spacing: 0.09em;
text-transform: uppercase;
color: var(--muted);
}
.v-ok .n { color: var(--ok); }
.v-wait .n { color: var(--warn); }
.v-no .n { color: var(--bad); }
@media (max-width: 620px) { .verdict { grid-template-columns: 1fr; } }
/* ---------- ledger table ---------- */
.scroller { overflow-x: auto; margin: 0 0 14px; }
table {
border-collapse: collapse;
width: 100%;
font-size: 14.5px;
min-width: 560px;
}
thead th {
text-align: left;
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 10.5px;
letter-spacing: 0.1em;
text-transform: uppercase;
color: var(--muted);
font-weight: 400;
padding: 0 12px 8px;
border-bottom: 1px solid var(--rule-strong);
}
tbody td {
padding: 10px 12px;
border-bottom: 1px solid var(--rule);
vertical-align: top;
}
tbody tr:hover td { background: var(--surface-alt); }
td.ref {
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 12.5px;
color: var(--muted);
white-space: nowrap;
width: 1%;
}
td.st { width: 1%; white-space: nowrap; }
.chip {
display: inline-block;
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 10.5px;
letter-spacing: 0.06em;
text-transform: uppercase;
padding: 3px 8px 2px;
border-radius: 2px;
white-space: nowrap;
}
.c-ok { background: var(--ok-bg); color: var(--ok); }
.c-wait { background: var(--warn-bg); color: var(--warn); }
.c-no { background: var(--bad-bg); color: var(--bad); }
/* ---------- sections ---------- */
h2 {
font-family: "Newsreader", ui-serif, Georgia, serif;
font-weight: 500;
font-size: 1.72rem;
letter-spacing: -0.01em;
line-height: 1.2;
text-wrap: balance;
margin: 0 0 6px;
}
h3 {
font-family: "Source Sans 3", sans-serif;
font-weight: 700;
font-size: 1.02rem;
letter-spacing: 0.005em;
margin: 30px 0 8px;
}
.rule-major {
border: 0;
border-top: 2px solid var(--ink);
margin: 56px 0 30px;
}
.sec-head {
display: flex;
align-items: baseline;
justify-content: space-between;
gap: 18px;
flex-wrap: wrap;
margin-bottom: 20px;
}
.ask {
border-left: 3px solid var(--rule-strong);
padding: 2px 0 2px 16px;
margin: 0 0 20px;
color: var(--muted);
font-family: "Newsreader", ui-serif, Georgia, serif;
font-style: italic;
font-size: 1.04rem;
}
.ask span {
display: block;
font-family: "JetBrains Mono", ui-monospace, monospace;
font-style: normal;
font-size: 10.5px;
letter-spacing: 0.1em;
text-transform: uppercase;
margin-bottom: 4px;
}
p { margin: 0 0 15px; max-width: 68ch; }
ul { margin: 0 0 15px; padding-left: 20px; max-width: 68ch; }
li { margin-bottom: 7px; }
strong { font-weight: 700; }
a { color: var(--accent); }
code {
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 0.855em;
background: var(--code-bg);
padding: 1px 5px;
border-radius: 2px;
}
pre {
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 12.5px;
line-height: 1.62;
background: var(--code-bg);
border: 1px solid var(--rule);
border-left: 3px solid var(--rule-strong);
padding: 15px 18px;
overflow-x: auto;
margin: 0 0 20px;
color: var(--ink-soft);
tab-size: 2;
}
pre code { background: none; padding: 0; font-size: inherit; }
pre b { color: var(--ink); font-weight: 700; }
/* ---------- callouts ---------- */
.callout {
background: var(--surface);
border: 1px solid var(--rule);
border-top: 3px solid var(--accent);
padding: 20px 22px 6px;
margin: 0 0 24px;
}
.callout.is-bad { border-top-color: var(--bad); }
.callout.is-warn { border-top-color: var(--warn); }
.callout .kicker {
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 10.5px;
letter-spacing: 0.11em;
text-transform: uppercase;
color: var(--muted);
margin: 0 0 9px;
}
blockquote {
margin: 0 0 15px;
padding-left: 16px;
border-left: 2px solid var(--rule-strong);
font-family: "Newsreader", ui-serif, Georgia, serif;
font-style: italic;
color: var(--ink-soft);
max-width: 64ch;
}
ol.decisions { counter-reset: d; list-style: none; padding: 0; margin: 0 0 15px; max-width: 68ch; }
ol.decisions li {
counter-increment: d;
position: relative;
padding-left: 34px;
margin-bottom: 12px;
}
ol.decisions li::before {
content: counter(d);
position: absolute;
left: 0; top: 1px;
font-family: "JetBrains Mono", ui-monospace, monospace;
font-size: 11px;
color: var(--accent);
border: 1px solid var(--rule-strong);
width: 22px; height: 22px;
display: grid; place-items: center;
border-radius: 50%;
}
footer {
margin-top: 60px;
padding-top: 20px;
border-top: 1px solid var(--rule);
font-size: 14px;
color: var(--muted);
}
footer ul { padding-left: 18px; }
:focus-visible { outline: 2px solid var(--accent); outline-offset: 2px; }
@media (prefers-reduced-motion: reduce) { * { transition: none !important; animation: none !important; } }
</style>
<div class="wrap">
<header class="masthead">
<p class="eyebrow">Reply to the Catalogue Drift Report &middot; 31 Aug 2026</p>
<h1>Seven findings, fourteen asks, re&#8209;measured</h1>
<p class="standfirst">Every figure below was taken from the live deployment on 2 September, not read off source. Eight asks are fulfilled and serving. Three wait on one decision from you. Three are not done — and one answer we gave you last week was wrong.</p>
<div class="meta-strip">
<span><b>Verified</b> 2026-09-02</span>
<span><b>Live catalogue</b> 55 brands / 1,572 products</span>
<span><b>Backend suite</b> 1,108 passing</span>
<span><b>Deploy</b> landed</span>
</div>
</header>
<section class="verdict">
<div class="v-ok"><span class="n">8</span><span class="l">Fulfilled &amp; live</span></div>
<div class="v-wait"><span class="n">3</span><span class="l">Blocked on your call</span></div>
<div class="v-no"><span class="n">3</span><span class="l">Not done</span></div>
</section>
<div class="callout is-bad">
<p class="kicker">Correction to our 01 September reply</p>
<p>We told you <code>image_id</code> was “a pure deterministic function of brand, product name and pack size.” That is true of the <strong>upload</strong> path and false of the <strong>scrape</strong> path — the one that broke your eleven links. Scraped ids carry a random <code>uuid4</code> fragment. You migrated onto <code>image_id</code> on the strength of that answer, so please read 01b before anything else.</p>
</div>
<div class="callout is-warn">
<p class="kicker">Two things that changed since we last wrote</p>
<p>The deploy <strong>has</strong> landed — <code>from_drop</code> is in the live <code>openapi.json</code> today — so the “please do not re-test” note in our last reply is out of date. <strong>Please do re-test.</strong> And 200 Hindustan Unilever products have been deleted since you took your measurement; that is finding 01 happening again, mid-conversation.</p>
</div>
<h2 style="margin-top:44px">The ledger</h2>
<p>Your seven findings contain fourteen distinct asks. This is where each one stands.</p>
<div class="scroller">
<table>
<thead>
<tr><th>Ref</th><th>The ask</th><th>Status</th></tr>
</thead>
<tbody>
<tr><td class="ref">01a</td><td>Does a pack size that once existed survive a re-scrape?</td><td class="st"><span class="chip c-no">Answered · no</span></td></tr>
<tr><td class="ref">01b</td><td>Is <code>image_id</code> stable across re-scrapes?</td><td class="st"><span class="chip c-no">Upload path only</span></td></tr>
<tr><td class="ref">01c</td><td><code>superseded_by</code>, a retired list, or a per-run changelog</td><td class="st"><span class="chip c-no">Not built</span></td></tr>
<tr><td class="ref">02</td><td>Per-row rejection reason, with the sheet row number</td><td class="st"><span class="chip c-ok">Fulfilled</span></td></tr>
<tr><td class="ref">03</td><td><code>brand_key</code> beside <code>brand</code> in the manifest</td><td class="st"><span class="chip c-ok">Fulfilled</span></td></tr>
<tr><td class="ref">04</td><td><code>from_drop</code> on each file in a run</td><td class="st"><span class="chip c-ok">Fulfilled · live</span></td></tr>
<tr><td class="ref">05</td><td><code>source_row</code> on each manifest entry</td><td class="st"><span class="chip c-ok">Fulfilled</span></td></tr>
<tr><td class="ref">06a</td><td>Merge <code>haldiram</code> and <code>haldirams</code></td><td class="st"><span class="chip c-ok">Fulfilled · live</span></td></tr>
<tr><td class="ref">06b</td><td>De-duplicate the PepsiCo pairs, settle one prefix convention</td><td class="st"><span class="chip c-wait">Needs 06e</span></td></tr>
<tr><td class="ref">06c</td><td>Strip the stray <code>150</code>, check the pattern elsewhere</td><td class="st"><span class="chip c-wait">Cause fixed · 17 rows</span></td></tr>
<tr><td class="ref">06d</td><td>Are Britannia, Parle and Patanjali complete?</td><td class="st"><span class="chip c-no">Answered · no</span></td></tr>
<tr><td class="ref">06e</td><td>Our question back: an old-id → new-id map before we rename</td><td class="st"><span class="chip c-wait">Awaiting you</span></td></tr>
<tr><td class="ref">07</td><td>A ~150-row loose-produce base list</td><td class="st"><span class="chip c-ok">159 rows · live</span></td></tr>
<tr><td class="ref">—</td><td>Reconcile your 1,614 against our 1,414</td><td class="st"><span class="chip c-ok">Exact</span></td></tr>
</tbody>
</table>
</div>
<hr class="rule-major">
<div class="sec-head">
<h2>01a &nbsp;Pack sizes are still replaced, not added to</h2>
<span class="chip c-no">Root cause live</span>
</div>
<p class="ask"><span>You asked</span>Confirm whether a pack size that once existed is meant to survive a re-scrape.</p>
<p><strong>It is not, and that is unchanged.</strong> <code>catalog_engine.py:999</code> still calls <code>upsert_brand_products(brand, enhanced_products, cleanup=True)</code>, which deletes every row in the brand table whose <code>image_id</code> is absent from the batch being written. A SKU present on one run and absent on the next is removed, not retired.</p>
<p>The upload path is safe and always was: everything under <code>POST /api/uploads/catalog</code> uses <code>cleanup=False</code>, with a test holding it there. A sheet you send cannot delete a row it does not mention. The deletions come from brand scraping only.</p>
<div class="sec-head" style="margin-top:44px">
<h2>01b &nbsp;<code>image_id</code> is deterministic on one path and random on the other</h2>
<span class="chip c-no">We answered wrong</span>
</div>
<p class="ask"><span>You asked</span>Confirm that <code>image_id</code> is stable across re-scrapes for a product whose name and pack size have not changed. We have now switched to storing it, and that switch only helps if the guarantee holds.</p>
<blockquote>“Yes. It is a pure deterministic function of brand, product name and pack size, with no clock, counter or run id in it.” — our reply, 01 September</blockquote>
<p>We answered from the wrong function. Two different paths mint <code>image_id</code>, and they do not behave alike:</p>
<ul>
<li><strong>Upload path</strong> — <code>build_image_id()</code> is a pure function of brand + name + size. Verified today: two calls with the same input both return <code>pepsico_cheetos_chips_100g</code>. Storing it is safe.</li>
<li><strong>Scrape path</strong> — <code>catalog_engine.py:695</code> mints ids with <code>s3_service.generate_image_id()</code>, which appends <strong><code>uuid4()[:8]</code></strong>. It reuses an existing id only when a row is found whose <code>product_name</code> is an <strong>exact string match</strong> — no size in the lookup, no normalisation, no fuzzy match.</li>
</ul>
<p>Both forms are visible side by side in the live catalogue right now:</p>
<pre><code>scraped id 27 Cheetos Chips 250g image_id cheetos_chips_<b>2d6bf74f</b>
scraped id 6 Kurkure Menthol 50g image_id kurkure_menthol_<b>49ef2d35</b>
uploaded id 731 Kurkure Masala Munch 90g image_id pepsico_kurkure_masala_munch_90g</code></pre>
<p>Those suffixes are random. So on a scrape, if the product name changes by a single character — exactly what a brand prefix appearing or disappearing does — the reuse lookup misses, a new random id is minted, and <code>cleanup=True</code> deletes the old row. That is the complete mechanism behind your eleven broken links, and <code>image_id</code> alone does not protect you from it.</p>
<div class="callout">
<p class="kicker">What this means for your migration</p>
<p>Storing <code>image_id</code> instead of our row id is still the right move — it is strictly better than the row id, and it is stable for everything arriving through the upload path. But it is <strong>not yet the durable key we implied</strong> for scraped brands. Making it one means giving the scrape path the same deterministic <code>build_image_id()</code> the upload path uses. We have not done that, because it renames ids on the next scrape of every scraped brand — the same migration problem as 06b, needing the old-to-new map in 06e and your timing.</p>
</div>
<div class="sec-head" style="margin-top:44px">
<h2>01c &nbsp;Nothing tells you when a product is dropped or renamed</h2>
<span class="chip c-no">Not built</span>
</div>
<p class="ask"><span>You asked</span>Anything that lets us detect it — a <code>superseded_by</code>, a retired list, even a changelog per run — turns a silent break into something we can act on.</p>
<p><strong>Not built.</strong> There is no <code>retired_at</code>, <code>superseded_by</code> or changelog anywhere in the codebase. Our questions from last week still stand, and the Hindustan Unilever deletion below makes them urgent:</p>
<ol class="decisions">
<li>Would a <code>retired_at</code> timestamp plus exclusion from the default read work instead of the row being deleted? It preserves the <code>image_id</code> so your stored link resolves to <em>something</em>, and gives us somewhere to hang <code>superseded_by</code> and a per-run changelog.</li>
<li>If so, should retired rows stay reachable through an explicit query, or vanish from the API entirely?</li>
</ol>
<hr class="rule-major">
<div class="sec-head">
<h2>02 &nbsp;Rejections now name the row and the reason</h2>
<span class="chip c-ok">Fulfilled</span>
</div>
<p class="ask"><span>You asked</span>A per-row reason on rejection, in the shape you already use for files — with <code>row</code> being the spreadsheet's own 1-based number, header included.</p>
<p>Delivered in that shape:</p>
<pre><code>"rejections": [
{ "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
"reason": "title is too short to be a real product name" }
]</code></pre>
<ul>
<li><code>row</code> is the 1-based sheet row with the header counted as row 1 — the row number is computed as <code>position + 2</code>, the same convention as the 422 responses, so it matches what the operator sees on screen. <code>null</code> only when the row cannot be located.</li>
<li>Capped at 50 per file.</li>
<li>Present on <strong>both</strong> the single-batch read and the list endpoint — <code>slim=True</code> strips only <code>products</code>.</li>
<li>Regression test: <code>test_the_rejection_row_matches_the_offending_sheet_line</code>. Documented in <code>INGESTION_API.md</code> with a field table.</li>
</ul>
<p>One correction to our own earlier account of this: the array itself was always in the response and only <code>row</code> is new. It was missing from our docs, which is why you could not find it. You were diffing 19 against 17 because we told you that was all you had.</p>
<div class="sec-head" style="margin-top:44px">
<h2>03 &nbsp;<code>brand_key</code> is published beside the display name</h2>
<span class="chip c-ok">Fulfilled</span>
</div>
<p class="ask"><span>You asked</span>Include the catalogue key alongside the display name in the manifest — a <code>brand_key</code> field beside <code>brand</code>. Our normalisation is a guess that currently happens to be right.</p>
<p>Every <code>products[]</code> entry now carries it. It is produced by the same <code>_sanitize_name()</code> the storage layer uses to name the table, so it cannot drift from the key the catalogue is actually addressed by — and the test asserts exactly that identity rather than a hard-coded string:</p>
<pre><code>assert product["brand_key"] == _sanitize_name(product["brand"])</code></pre>
<p>Your normalisation is correct as far as we can tell. It stays a guess, though, and the failure mode is silent — a wrong key finds nothing rather than erroring. Use the published field.</p>
<div class="sec-head" style="margin-top:44px">
<h2>04 &nbsp;<code>from_drop</code> — the one that could corrupt inventory</h2>
<span class="chip c-ok">Fulfilled &amp; live</span>
</div>
<p class="ask"><span>You asked</span><code>from_drop</code> on each file in a run, carrying the drop id it was released from. Matching on it is exact, where matching on a filename is a coincidence we are relying on.</p>
<p>Each file in a run now carries <code>from_drop</code>, the exact inverse of <code>released_to</code>. It is stamped after staging, written to disk and re-read from the manifest — <code>null</code> only for a file that went straight into a run without sitting in an inbox.</p>
<p><strong>Confirmed deployed.</strong> <code>from_drop</code> is in the live <code>openapi.json</code> served by <code>mcp.nearle.ai.in</code> today, which is also how we know the rest of this deploy has landed.</p>
<p>The collision you described is now a permanent regression guard: <code>test_a_run_says_which_drop_each_file_came_from</code> stages two drops from different senders both named <code>a.csv</code> and asserts they are distinguishable, and a second test proves the id survives a re-read of the run rather than existing only in the reply.</p>
<div class="sec-head" style="margin-top:44px">
<h2>05 &nbsp;<code>source_row</code> traces every product to its sheet line</h2>
<span class="chip c-ok">Fulfilled</span>
</div>
<p class="ask"><span>You asked</span><code>source_row</code> on each manifest entry — the 1-based sheet row that produced it, so we can say “row 14 became these three” and “rows 6 and 11 produced nothing”.</p>
<p>Both of those now work. <code>source_row</code> is the 1-based row with the header as row 1, and it is deliberately <strong>many-to-one</strong>: a pack-size cell reading <code>100g, 200g, 500g</code> becomes three products that all report the same <code>source_row</code>. Rows that produced nothing are those absent from every entry — a set difference rather than a name-matching heuristic.</p>
<p>The map is built from the kept rows before dedupe and carried alongside the storage rows rather than inside them, so the number always describes the product that was actually written, not one that lost a dedupe. Tests cover the one-to-one case and the exploded case.</p>
<hr class="rule-major">
<div class="sec-head">
<h2>06 &nbsp;Duplicates, stray values, and coverage</h2>
<span class="chip c-wait">Mixed</span>
</div>
<h3>06a &nbsp;The Haldiram split is merged, and cannot recur</h3>
<p>Live today: <code>haldiram</code> returns <strong>0 products</strong>, <code>haldirams</code> returns <strong>2</strong>. Merging the rows alone would have fixed nothing — the next sheet spelling it without the “s” would rebuild the table — so the alias went in too. <code>Haldiram</code>, <code>haldiram</code>, <code>HALDIRAM</code> and <code>Haldiram's</code> now all resolve to <code>haldirams</code>.</p>
<p>While merging we found something you could not have seen: both surviving rows were carrying <strong>Lion Dates' FSSAI licence</strong> rather than Haldiram's. That is a regulatory identifier on the wrong manufacturer's product. Fixed in the database and the seed file.</p>
<h3>06b &nbsp;The PepsiCo pairs are still there, deliberately</h3>
<p>Confirmed live today, unchanged:</p>
<pre><code>661 PepsiCo Lays Classic Salted 52g pepsico_pepsico_lays_classic_salted_52g
730 Lays Classic Salted 52g 150 pepsico_lays_classic_salted_52g_150
662 PepsiCo Kurkure Masala Munch 90g pepsico_pepsico_kurkure_masala_munch_90g
731 Kurkure Masala Munch 90g pepsico_kurkure_masala_munch_90g</code></pre>
<p>We have not touched them, because <strong>de-duplicating means renaming, and <code>image_id</code> is derived from the name.</strong> Renaming <code>PepsiCo Kurkure Masala Munch 90g</code> does not merge the two rows — it mints a <em>third</em> id and breaks any link pointing at either of the first two. You have just finished migrating onto <code>image_id</code>; a well-meant cleanup on our side would re-break exactly what you repaired. Our lean on the convention is <strong>without</strong> the brand prefix, since brand is already its own column, but we will follow whichever you pick.</p>
<h3>06c &nbsp;The stray number: cause fixed, 17 rows still carrying it</h3>
<p>The number is not a price — 40 against ₹299, 150 against ₹21. It is the <strong>case-pack count</strong>: a sheet's “Quantity” column was being mapped to the pack-size field, the bare number became the size, and it was then appended to the product name.</p>
<p><strong>The cause is fixed on two layers, both verified today.</strong> <code>Quantity</code> was removed from the pack-size keyword rule — only <code>net qty</code> / <code>net quantity</code>, the Indian labelling term for a real pack size, still match. A sheet with columns <code>Product Name, Quantity, Pack Size, Brand</code> now binds <code>Pack Size</code> and reports <code>Quantity</code> as unrecognised; previously a leading <code>Quantity</code> column also shut out the sheet's real <code>Pack Size</code> column, because mapping is first-wins by position. Independently, a unitless number is now discarded with a reason recorded: <em>“ignored pack size '150': a number with no unit is a quantity, not a size.”</em></p>
<p>No new row can acquire this. <strong>The existing rows are still there — 17 of them, scanned across all 55 live brands today.</strong> We said 18 last week; the eighteenth was the corrupted Haldiram row removed in the merge.</p>
<pre><code>brooke_bond 1 Red Label Tea 500g <b>30</b>
fortune 2 Fortune Sunflower Oil 1L <b>48</b>
hindustan_unilever 2135 Tata Salt Iodised 1kg <b>120</b>
hindustan_unilever 2136 Bru Instant Coffee 100g <b>24</b>
hindustan_unilever 2137 Horlicks Classic Malt 500g <b>18</b>
hindustan_unilever 2138 Surf Excel Easy Wash 1kg <b>36</b>
hindustan_unilever 2139 Vim Dishwash Bar 300g <b>90</b>
hindustan_unilever 2140 Dove Cream Beauty Bar 100g <b>64</b>
hindustan_unilever 2141 Clinic Plus Shampoo 175ml <b>40</b>
india_gate 1 India Gate Basmati Rice 1kg <b>60</b>
britannia 7 Britannia Good Day Cashew 200g <b>60</b>
colgate_palmolive 483 Colgate Strong Teeth 200g <b>50</b>
itc 5 Aashirvaad Shudh Chakki Atta 5kg <b>40</b>
reckitt_benckiser 1 Harpic Power Plus 500ml <b>30</b>
reckitt_benckiser 2 Dettol Original Soap 125g <b>80</b>
coca_cola 1041 Coca-Cola 750ml <b>72</b>
pepsico 730 Lays Classic Salted 52g <b>150</b></code></pre>
<p>Repairing them is a rename, so it is blocked on 06e exactly as the PepsiCo pairs are. One thing unrelated but visible in that list: <code>Tata Salt Iodised</code> is filed under <code>hindustan_unilever</code>, and Tata Salt is not an HUL product. We are looking at it.</p>
<h3>06d &nbsp;Britannia, Parle and Patanjali are incomplete scrapes</h3>
<p>You read those right. <strong>Confirmed incomplete, not small brands</strong> — still <code>6</code>, <code>3</code> and <code>3</code> products live today. The re-scrape has not been run yet; we would rather re-run them than have you build around the gap.</p>
<h3>06e &nbsp;What we need back from you</h3>
<p>Three answers unblock 06b, 06c and the <code>image_id</code> determinism fix in 01b — all of which are renames, and all of which should land in one pass:</p>
<ol class="decisions">
<li><strong>Which convention wins?</strong> Brand prefix in the product name, or not. Either is fine; we care only that it is one of them.</li>
<li><strong>Would an old-id → new-id map, delivered in advance for every row we touch, let you re-point rather than clear?</strong> It is straightforward for us to produce, and it would cover the 17 stray-number rows, the PepsiCo pairs and the scraped-id migration together.</li>
<li><strong>Timing</strong>, so it lands in one pass rather than trickling.</li>
</ol>
<hr class="rule-major">
<div class="sec-head">
<h2>07 &nbsp;Loose produce: 159 rows, live now</h2>
<span class="chip c-ok">Fulfilled &amp; live</span>
</div>
<p class="ask"><span>You asked</span>A loose-produce base list — roughly 150 rows covering fruit, vegetables, greens, flowers, fish and milk. Name and image only; no brand, no pack size, no price.</p>
<p>Serving now under brand key <code>own_products</code>, in that shape:</p>
<pre><code>Fruits &amp; Vegetables 95 Flowers 14
Fresh Herbs &amp; Greens 17 Dairy (loose) 13
Fish &amp; Seafood 15 Eggs 5
<b>159 rows</b></code></pre>
<p>It includes the specific items your audit listed — Jasmine, Lotus, Red Rose, Thulasi, Drumstick, Curry Leaves, the four banana varieties, Tuna, Mackerel. Each row carries a search embedding, so these are reachable through semantic search and not just exact match, and every seeded row classifies identically to how an uploaded copy of the same name would — a grocer typing “Tomato” lands on the seeded row instead of creating a second one. The list is hand-authored rather than scraped, so none of finding 01 applies to it.</p>
<p><strong>112 of the 159 (70%) carry an image</strong> — all of the fruit, vegetables, greens and herbs. Flowers, fish, loose dairy and eggs are still name-only; we stopped the fetch part-way. Tell us whether the list is more useful to you complete-but-later or partial-but-now.</p>
<div class="callout is-warn">
<p class="kicker">The underlying bug was worse than a coverage gap</p>
<p>Produce rows were not rejected. They were <strong>misfiled</strong>. The brand fallback took the first word of the name and whole-word matched it against our alias map, so <code>Apple</code> became brand “Apple”, <code>Curry Leaves</code> became “Curry”, and <strong><code>Red Rose</code> was being written into the Brooke Bond tea catalogue</strong> — where our enrichment then stamped that brand's real FSSAI licence onto it. Your 139 hand-typed products were the visible symptom; this was underneath. Loose goods are now recognised as commodities and filed under <code>Own Products</code> before brand inference can touch them.</p>
</div>
<p>Verified against your own audit strings today, including your merchants' misspellings:</p>
<pre><code>Apple, Tomato, Curry Leaves, Red Rose, Thulasi, Drumstick, Tuna -> unbranded
Bitter guard, Bottle ground, Ladies Finger -> unbranded</code></pre>
<p>Produce is also exempt from the invented-pack-size fallback that causes finding 01: an <code>Own Products</code> row with no weight column gets one <code>Standard</code> row, not three made-up ones. Name, weight and price are stored; HSN, SKU, barcode, FSSAI and description are left null rather than invented.</p>
<h3>Two limitations worth knowing</h3>
<ul>
<li><strong>Place-qualified produce still reads as branded.</strong> <code>Salem Mango</code>, <code>Mysore Banana</code> and <code>Jammu Apple</code> — all real strings from your Ragul Stores data — classify as branded, because <code>Mysore</code> is also a real brand. The classifier is deliberately conservative: collapsing a real regional brand into the unbranded bucket is much harder to undo than a mango in the wrong table. The workaround is already in the pipeline — <strong>if the sheet has a brand column and leaves the cell empty, we believe it</strong> and file the row under Own Products regardless of the name.</li>
<li><strong><code>Maceral</code> does not resolve</strong> — your Ragul Stores spelling of mackerel. <code>Mackerel</code> does. Send us any other spellings your merchants actually use and we will add them.</li>
</ul>
<hr class="rule-major">
<h2>Your 1,614 against our count</h2>
<p>You counted 55 brands and 1,614 products. We measure <strong>55 brands and 1,572 products</strong> today, of which 159 are the new produce rows — so 1,413 of the old kind. The 201-row gap resolves exactly:</p>
<pre><code>your count 1,614
less Hindustan Unilever, 443 -> 243 today <b>-200</b>
less the Haldiram row removed in the merge <b>-1</b>
-------
1,413 our non-produce count
plus the produce base list <b>+159</b>
-------
<b>1,572</b> live today</code></pre>
<div class="callout is-bad">
<p class="kicker">This is not a counting discrepancy</p>
<p><strong>200 Hindustan Unilever products have been deleted since you measured</strong> — the largest brand in the catalogue, between your report and this reply. It is finding 01a happening again while we were writing about it, and it is the strongest argument we have for the retirement model in 01c. That is the answer we would like from you soonest.</p>
</div>
<footer>
<p><strong>What we checked before sending this</strong></p>
<ul>
<li>Every figure re-measured against the live catalogue API on 2 September, not read off source; the <code>image_id</code> behaviour read out of both functions that mint it.</li>
<li>The ids you listed are still absent — pepsico <code>4, 19, 20, 25, 26</code>, dabur <code>19, 20</code>, nestle <code>1</code>. Clearing those links rather than re-pointing them by name remains the correct call.</li>
<li>Stray-number scan run across all 55 brands and all 1,572 live products.</li>
<li>Full backend suite: 1,108 tests passing.</li>
</ul>
</footer>
</div>

View File

@@ -0,0 +1,351 @@
# Catalogue Drift Report — requirement-by-requirement verification
Re-measured **2026-09-02** against the live deployment and the current source.
This supersedes the figures in `DRIFT_REPORT_RESPONSE.md` (written 01 Sep), which
was drafted before the deploy landed and contains **one claim that is wrong** —
see the correction under 01(b).
Fourteen distinct asks are contained in the seven findings. Eight are fulfilled,
three are blocked on a decision from the integrator, three are not done.
| # | The ask | Status |
| --- | --- | --- |
| 01a | Does a pack size that once existed survive a re-scrape? | **Answered: no.** Root cause still live |
| 01b | Is `image_id` stable across re-scrapes? | **Answered: only on the upload path.** Previous answer was wrong |
| 01c | `superseded_by` / retired list / per-run changelog | **Not built** |
| 02 | Per-row rejection reason, with the sheet row number | **Fulfilled** |
| 03 | `brand_key` beside `brand` in the manifest | **Fulfilled** |
| 04 | `from_drop` on each file in a run | **Fulfilled, and live** |
| 05 | `source_row` on each manifest entry | **Fulfilled** |
| 06a | Merge `haldiram` and `haldirams` | **Fulfilled, and live** |
| 06b | De-duplicate the PepsiCo pairs, settle one prefix convention | **Not done** — blocked on 06e |
| 06c | Strip the stray `150`, check for the pattern elsewhere | **Cause fixed; 17 existing rows not yet repaired** |
| 06d | Are Britannia, Parle, Patanjali complete? | **Answered: no, incomplete.** Re-scrape not yet run |
| 06e | (our question back) old-id to new-id map before we rename anything | **Awaiting your answer** |
| 07 | ~150-row loose-produce base list | **Fulfilled, and live** — 159 rows |
| — | Reconcile your 1,614 against our 1,414 | **Reconciled exactly** — see below |
Deployment state: the backend deploy **has** landed (`from_drop` is present in the
live `openapi.json` at `mcp.nearle.ai.in`), and both database changes — the
Haldiram merge and the 159 produce rows — are serving. The "not live yet, please
do not re-test" caveat in the 01 Sep reply is out of date; **please do re-test.**
---
## 01a — Does a pack size survive a re-scrape?
**No, and that is unchanged.** `catalog_engine.py:999` still calls
`upsert_brand_products(brand, enhanced_products, cleanup=True)`, which deletes
every row in the brand table whose `image_id` is absent from the batch being
written. A SKU present on one run and absent on the next is removed, not retired.
The upload path is still safe and always was: everything under
`POST /api/uploads/catalog` uses `cleanup=False` (`store_catalog_pipeline.py:779`,
with a test). A sheet you send cannot delete a row it does not mention.
**This is still destroying rows.** See the reconciliation at the end — 200
Hindustan Unilever products have disappeared since you took your measurement.
## 01b — Is `image_id` stable across re-scrapes?
**Correction to what we told you on 01 September.** We said:
> "Yes. It is a pure deterministic function of brand, product name and pack
> size, with no clock, counter or run id in it."
That is true of the **upload** path and false of the **scrape** path, which is
the path that broke your eleven links. We answered from the wrong function.
- **Upload path** — `build_image_id()` (`store_catalog_pipeline.py:655`) is a
pure function of brand + name + size. Verified: two calls with the same input
both return `pepsico_cheetos_chips_100g`. Storing it is safe.
- **Scrape path** — `catalog_engine.py:695` mints ids with
`s3_service.generate_image_id()`, which appends **`uuid4()[:8]`**
(`s3_service.py:59`). It is not deterministic. It reuses an existing id only
when `get_existing_product_image_id()` finds a row whose `product_name` is an
**exact string match** (`vector_store.py:489`) — no size in the lookup, no
normalisation, no fuzzy match.
You can see both forms side by side in the live catalogue today:
```
scraped id 27 Cheetos Chips 250g image_id cheetos_chips_2d6bf74f
scraped id 6 Kurkure Menthol 50g image_id kurkure_menthol_49ef2d35
uploaded id 731 Kurkure Masala Munch 90g image_id pepsico_kurkure_masala_munch_90g
```
The `2d6bf74f` and `49ef2d35` are random. So on a scrape, if the product name
changes by a single character — which is exactly what a brand prefix appearing
or disappearing does — the reuse lookup misses, a **new random id** is minted,
and `cleanup=True` deletes the old row. That is the full mechanism behind your
eleven broken links, and `image_id` alone does not protect you from it.
**What this means for your migration.** Storing `image_id` instead of our row id
is still the right move — it is strictly better than the row id, and it is stable
for everything that arrives through the upload path. But it is **not yet the
durable key we implied** for scraped brands. Making it one means giving the
scrape path the same deterministic `build_image_id()` the upload path uses. We
have not done that, because it renames ids on the next scrape of every scraped
brand, which is the same migration problem as 06b — it needs the old-to-new map
in 06e, and it needs your timing.
We are sorry for the incorrect answer; you told us you had switched on the
strength of it.
## 01c — `superseded_by`, a retired list, or a per-run changelog
**Not built.** There is no `retired_at`, `superseded_by` or changelog anywhere in
the codebase. Our questions from the 01 Sep reply still stand: would a
`retired_at` timestamp with exclusion from the default read work instead of the
row being deleted, and should retired rows stay reachable through an explicit
query or vanish from the API entirely?
---
## 02 — Per-row rejection reasons — fulfilled
`rejections[]` is populated at `store_catalog_pipeline.py:973` and carries
exactly the shape you asked for:
```jsonc
"rejections": [
{ "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
"reason": "title is too short to be a real product name" }
]
```
- `row` is the 1-based sheet row with the header counted as row 1 —
`row_no = position + 2` (`store_catalog_pipeline.py:878`), the same convention
as the 422 responses, so it matches what the operator sees. `null` only when
the row cannot be located.
- Capped at 50 per file (`as_dict`, line 188).
- Present on **both** the single-batch read and the list endpoint: `slim=True`
strips only `products` (`batch_common.py:271`).
- Regression test: `test_the_rejection_row_matches_the_offending_sheet_line`.
- Documented in `INGESTION_API.md` with a field table.
One correction to our own earlier account: the array itself was always in the
response, and only `row` is new. It was missing from our docs, which is why you
could not find it.
## 03 — `brand_key` — fulfilled
Every `products[]` entry carries `brand_key` beside `brand`
(`store_catalog_pipeline.py:743`). It is produced by `_sanitize_name()`, the
same function the storage layer uses to name the table, so it cannot drift from
the key the catalogue is addressed by.
Test `test_a_product_carries_the_key_the_catalogue_is_addressed_by` asserts
`product["brand_key"] == _sanitize_name(product["brand"])`.
## 04 — `from_drop` — fulfilled, and live
`BatchFileOut.from_drop` (`batch_common.py:216`) is stamped by `_stamp_origins`
at both `from-inbox` call sites (`batch_catalog.py:509,518`), written to disk,
and re-read from the manifest. `null` for a file that went straight into a run
without sitting in an inbox.
**Confirmed deployed** — `from_drop` is in the live `openapi.json` served by
`mcp.nearle.ai.in` today.
Two regression tests cover the collision you described:
`test_a_run_says_which_drop_each_file_came_from` stages two drops from different
senders both named `a.csv` and asserts they are distinguishable, and
`test_from_drop_survives_a_reread_of_the_run` proves it reaches disk.
## 05 — `source_row` — fulfilled
Every `products[]` entry carries `source_row`
(`store_catalog_pipeline.py:751`), the 1-based sheet row with the header as row
1. It is deliberately many-to-one: a pack-size cell reading `100g, 200g, 500g`
becomes three products that all report the same `source_row`. Rows that produced
nothing are those absent from every entry, so "rows 6 and 11 produced nothing"
is now a set difference rather than a name-matching heuristic.
The map is built from `kept` before dedupe and carried alongside the storage
rows rather than inside them, so it describes the product actually written.
Tests cover both the one-to-one and the exploded case.
---
## 06 — Duplicate brands, duplicate products, stray names
### 06a — Haldiram merge: done, and live
Live today: `haldiram` returns **0 products**, `haldirams` returns **2**. The
alias is in the registry (`brand_registry.py:262-263`), so `Haldiram`,
`haldiram`, `HALDIRAM` and `Haldiram's` all now resolve to `haldirams` — the
merge cannot be undone by the next sheet that spells it without the "s".
Both surviving rows also had their FSSAI licence corrected: they were carrying
Lion Dates' number rather than Haldiram's.
### 06b — The PepsiCo duplicate pairs: still there
Confirmed live today, unchanged:
```
661 PepsiCo Lays Classic Salted 52g pepsico_pepsico_lays_classic_salted_52g
730 Lays Classic Salted 52g 150 pepsico_lays_classic_salted_52g_150
662 PepsiCo Kurkure Masala Munch 90g pepsico_pepsico_kurkure_masala_munch_90g
731 Kurkure Masala Munch 90g pepsico_kurkure_masala_munch_90g
```
Deliberately untouched. De-duplicating means renaming, `image_id` is derived
from the name, and renaming mints a *third* id rather than merging two — which
would re-break the links you have just finished repairing. We need 06e first.
Our lean on the convention is **without** the brand prefix, since brand is
already its own column. We will follow whichever you prefer.
### 06c — The stray number: cause fixed, rows not yet repaired
**The cause is fixed, on two layers, both verified today:**
- `Quantity` was removed from the `size_variants` keyword rule
(`user_products.py:176-182`); only `net qty` / `net quantity`, the Indian
labelling term for a real pack size, still match. Verified: a sheet with
columns `Product Name, Quantity, Pack Size, Brand` now binds `Pack Size` and
reports `Quantity` as unrecognised. Previously a leading `Quantity` column
also shut out the sheet's real `Pack Size` column, because mapping is
first-wins by position.
- `_sizes_for()` discards a unitless number and records why
(`store_catalog_pipeline.py:428`): *"ignored pack size '150': a number with no
unit is a quantity, not a size"*.
No new row can acquire this. **The existing rows are still there — 17 of them,
scanned across all 55 live brands today** (we said 18 on 01 Sep; the 18th was
the corrupted Haldiram row removed in the merge):
```
brooke_bond 1 Red Label Tea 500g 30
fortune 2 Fortune Sunflower Oil 1L 48
hindustan_unilever 2135 Tata Salt Iodised 1kg 120
hindustan_unilever 2136 Bru Instant Coffee 100g 24
hindustan_unilever 2137 Horlicks Classic Malt 500g 18
hindustan_unilever 2138 Surf Excel Easy Wash 1kg 36
hindustan_unilever 2139 Vim Dishwash Bar 300g 90
hindustan_unilever 2140 Dove Cream Beauty Bar 100g 64
hindustan_unilever 2141 Clinic Plus Shampoo 175ml 40
india_gate 1 India Gate Basmati Rice 1kg 60
britannia 7 Britannia Good Day Cashew 200g 60
colgate_palmolive 483 Colgate Strong Teeth 200g 50
itc 5 Aashirvaad Shudh Chakki Atta 5kg 40
reckitt_benckiser 1 Harpic Power Plus 500ml 30
reckitt_benckiser 2 Dettol Original Soap 125g 80
coca_cola 1041 Coca-Cola 750ml 72
pepsico 730 Lays Classic Salted 52g 150
```
Repairing them is a rename, so it is blocked on 06e exactly as 06b is.
Unrelated but visible in that list: `Tata Salt Iodised` is filed under
`hindustan_unilever`. Tata Salt is not an HUL product. We are looking at it.
### 06d — Britannia, Parle, Patanjali
**Confirmed incomplete scrapes, not small brands.** Still `6`, `3` and `3`
products live today — the re-scrape has not been run yet.
### 06e — What we need back from you
Before we rename anything (06b, 06c, and the `image_id` determinism fix in 01b):
1. Which prefixing convention wins — brand in the product name, or not?
2. Would an **old-id to new-id map, delivered in advance for every row we touch**,
let you re-point rather than clear? It is straightforward for us to produce.
3. Timing, so all of it lands in one pass rather than trickling.
---
## 07 — Loose produce — fulfilled, and live
**159 rows, live now** under brand key `own_products`, in the shape you asked
for: name and category, no brand, no pack size, no price.
```
Fruits & Vegetables 95 Flowers 14
Fresh Herbs & Greens 17 Dairy (loose) 13
Fish & Seafood 15 Eggs 5
```
Verified live: `getproducts?brand=own_products` returns them, e.g.
`own_products_apple_standard`, with `highlights: ["Loose / unbranded", ...]`.
**112 of the 159 (70%) carry an image** — all the fruit, vegetables, greens and
herbs. Flowers, fish, loose dairy and eggs are still name-only. Tell us whether
you want the list complete-but-later or partial-but-now.
Each row carries a search embedding, so these reach semantic search and not just
exact match, and every seeded row classifies identically to how an uploaded copy
of the same name would — a grocer typing "Tomato" lands on the seeded row rather
than creating a second one.
**The underlying bug was worse than a coverage gap**, and is fixed: produce rows
were not rejected, they were *misfiled*. The brand fallback took the first word
of the name and matched it against the alias map, so `Apple` became brand
"Apple", `Curry Leaves` became "Curry", and **`Red Rose` was being written into
the Brooke Bond tea catalogue** and stamped with that brand's real FSSAI licence.
Loose goods are now recognised as commodities and filed under `Own Products`
before brand inference runs.
Verified against your own audit strings today:
```
Apple, Tomato, Curry Leaves, Red Rose, Thulasi, Drumstick, Tuna -> unbranded
Bitter guard, Bottle ground, Ladies Finger (your misspellings) -> unbranded
```
Produce is also exempt from the invented-pack-size fallback that causes 01a: an
`Own Products` row with no weight gets one `Standard` row, not three made-up
ones (`store_catalog_pipeline.py:460`).
**Two limitations, unchanged and worth knowing:**
- Place-qualified produce still reads as branded: `Salem Mango`, `Mysore Banana`
and `Jammu Apple` — all real strings from your Ragul Stores data — classify as
branded, because `Mysore` is also a real brand. The classifier is deliberately
conservative: collapsing a real regional brand into the unbranded bucket is
much harder to undo than a mango in the wrong table. The workaround is in the
pipeline already — **if the sheet has a brand column and leaves the cell
empty, we believe it** and file the row under Own Products regardless of name.
- `Maceral` (your Ragul Stores spelling of mackerel) does not resolve. `Mackerel`
does. Send us any other spellings your merchants actually use and we will add
them.
---
## Reconciling your 1,614 against our count
You counted 55 brands / 1,614 products. We now measure **55 brands / 1,572
products**, of which 159 are the new produce rows — so **1,413 of the old kind**.
Your number and ours differ by 201, and it resolves exactly:
```
your count 1,614
less Hindustan Unilever, 443 -> 243 today -200
less the Haldiram row removed in the merge -1
-------
1,413 = our non-produce count
plus the produce base list +159
-------
1,572 = live today
```
**200 Hindustan Unilever products have been deleted since you measured.** That is
not a counting discrepancy — it is finding 01a happening again, to the largest
brand in the catalogue, between your report and this reply. It is the strongest
argument we have for the retirement model in 01c, and it is why we would like an
answer on 01c sooner than on the rest.
---
## What we checked before writing this
- Every figure re-measured today against the live catalogue API, not read off
source; the `image_id` behaviour read out of the two functions that mint it.
- The ids you listed are still absent: pepsico `4, 19, 20, 25, 26`, dabur
`19, 20`, nestle `1`. Clearing those links rather than re-pointing them by
name remains the correct call.
- Stray-number scan run across all 55 brands and all 1,572 live products.
- Full backend suite: **1,108 tests passing.**

View File

@@ -11,11 +11,14 @@ says so, and every assertion below is ultimately about one of two failures -
* something running that nobody approved, and
* a file being lost, or run twice, on its way out of the inbox.
The response SHAPES are asserted literally rather than loosely, because
frontend/src/pages/InboxPanel.jsx was written against this contract before the
backend existed. `submission_id`, `file_id`, `pending_count` and the `dismissed`
count are read by name there; a rename that only this file catches is cheap, and
one that nothing catches is a blank admin tab with no error.
The response SHAPES are asserted literally rather than loosely, because the
frontend was written against this contract before the backend existed.
`submission_id`, `file_id`, `pending_count` and the `dismissed` count are read
by name in frontend/src/pages/UploadResultsPanel.jsx - the approval gate lives
inside that panel now, revealed only when `pending_count` is above zero, since
UPLOAD_AUTORUN=true leaves the inbox permanently empty. A rename that only this
file catches is cheap, and one that nothing catches is a blank admin tab with no
error.
"""
from __future__ import annotations