Files
catalogue_backend/docs/INGESTION_API.md
2026-09-08 15:18:29 +05:30

46 KiB
Raw Blame History

Sending spreadsheets to the catalogue

Base: https://mcp.nearle.ai.in · Formats: .xlsx .xls .csv .tsv

Four endpoints accept a spreadsheet. One is open — no credential needed — and stages files for review; three are operator imports that require a credential and write directly to the store, sales and nutrition tables.

If you only need to send sheets and watch them run, §1 is the whole story and you need no credential. §2 covers the admin side — the review inbox, starting a batch, and picking who executes it.

Deployment status. The endpoints, the open drop and the stage timeline in Stage-by-stage progress are live on https://mcp.nearle.ai.in — verified 2026-08-31 by probing the deployment, not by reading the source.

UPLOAD_AUTORUN is the exception: committed, not yet deployed. Until it ships, an upload still lands in the review inbox and answers pending. Both modes are documented below and both are supported; if you are integrating right now, write the client so it does not care which is running — §8 Checking what is deployed shows how to tell in one curl.


Read this first

These endpoints used to fill a missing column with a plausible constant. A sheet whose headers didn't match still returned 200 OK — having written rows attributed to brand amul in store store_mumbai_1 at MRP 100. Nutrition rows were worse: absent nutrients became invented numbers and an absent allergens column became allergens = ['None'], stored as data_status: "verified".

The rule now: nothing is invented. A missing value is stored as NULL. A row missing the fields that identify it is skipped and reported back with its row number. An upload where nothing could be imported is a 422, not a 200 with rows_imported: 0.

A blank allergens cell records "we were not told" — which is not the claim "contains no allergens". Leave the column out rather than writing None.

Arithmetic is not invention. line_total is still derived from quantity × unit_price when your sheet omits it, because both were supplied. That distinction is the whole design: compute from given values, never substitute for absent ones.


Authentication

The catalogue drop (§1) needs no credential. Anyone who can reach the host can send it a spreadsheet — and with UPLOAD_AUTORUN on, that spreadsheet runs through the pipeline immediately and its products land in the live catalogue. There is no undo; ingestion is an upsert.

That is a deliberate trade, not an oversight. The requirement was uploads that run without manual intervention, and a review queue that needs an admin to press a button is not that. What bounds the endpoint is throughput rather than identity: the limits in Limits, and a single worker thread behind a queue of BATCH_QUEUE_MAX. A sender can occupy the ingestion worker; they cannot multiply it.

With UPLOAD_AUTORUN=false the older behaviour returns — files wait in an admin review inbox and the cost of an unwanted drop is disk until somebody declines it.

The orchestration routes (§2) and the three operator imports (§3) write straight to the database and stay credentialed. Two credential types are accepted there:

X-API-Key: <your secret>          # machine clients
Authorization: Bearer <token>     # from POST /api/auth/login, 12h lifetime

Send one or the other. Sending both is not rejected — the bearer token is tried first and the API key ignored — but relying on that is a way to spend an hour debugging the wrong credential.

Endpoint Permission Who holds it
POST /api/uploads/catalog none anyone
GET /api/uploads/catalog/{batch_id} none — the id is the credential anyone holding that id
GET /api/uploads/catalog (list) upload_catalog uploader, admin
/api/admin/catalog-batch/* (§2) admin admin only
/api/upload/stores upload_store_inventory user, admin
/api/upload/analytics manage_analytics admin
/api/upload/nutrition manage_nutrition admin

admin is a superuser: Principal.has_permission grants it every permission rather than requiring each to be listed. On the credentialed endpoints, no credential → 401 and a valid credential without the permission → 403.

A credential that is present but wrong is a 401 everywhere, including on the open drop. Anonymous is allowed; mistyped is not — a typo must not silently downgrade to an anonymous submission the sender then cannot find.


1. The catalogue drop

This is the one to give a colleague. No credential, up to 20 sheets per request, answers immediately with a batch_id.

What happens on arrival depends on one setting, UPLOAD_AUTORUN:

true (the intended production mode) false
On arrival The batch is queued and the 11 stages run Files wait in the admin review inbox
Batch status queued, then running pending
The id you get back is the run — poll it and watch stages[] is a drop id; the run gets a different one, reached via the file's released_to
Anything to do No An admin must tick the files and press Start

Write your client so it does not care which is running. Both answer 202 with an id that GET /api/uploads/catalog/{id} understands; only the number of hops differs. Poll the id you were given, and follow released_to if a file ever grows one.

POST /api/uploads/catalog → 202

curl -X POST https://mcp.nearle.ai.in/api/uploads/catalog \
  -F 'files=@store-catalog.xlsx' \
  -F 'files=@second-store.xlsx' \
  -F 'sender=priya'
Form field Purpose
files required The spreadsheets. Repeat the field for more than one; up to 20.
sender optional A label for the inbox, so the admin can see who sent what. Free text, trimmed to 60 chars. Defaults to anonymous.

use_llm and fetch_images are not accepted here. They commit the host to outbound work, and this endpoint's caller is anonymous, so the choice is not theirs to make. Passing them is inert — the response reports what was actually used.

Both are on for an auto-started run: image search fills the product images, and the LLM fills blank descriptions wherever an Ollama server is reachable. There is none in production today, so in practice descriptions arrive exactly as your sheet wrote them and the run is otherwise unaffected.

Which columns are read

Headers match case-insensitively by keyword, so Product Name, ITEM and variant all resolve to the same field. A product-name column is the only hard requirement — a sheet without one is rejected during the request.

Field Headers that match
product_name required product · item · variant · name
brand optional brand · manufacturer · company
category optional category · segment
size_variants optional size · pack · weight · volume · net qty · quantity
final_selling_price optional final price · selling price · price · mrp · rate · cost
barcode optional barcode · bar code · gtin · ean · upc
fssai_license optional any header containing fssai
hsn_code optional any header containing hsn
image_url optional image · photo · picture · url · link
description optional description · desc · detail
title optional title
product_sku optional sku
price_range optional price range · range
providers optional provider · platform · marketplace · available at
highlights optional highlight · feature · benefit
nutrients optional nutrient · nutrition

Rules are evaluated in the order above, so a header matching two of them takes the earlier field. Where two columns map to the same field, the leftmost wins and the other is reported as ignored. Unrecognised columns are listed back, not an error.

Response — 202, one good file and one bad

The bad file is kept as a failed member rather than dropped, so a sender who submitted two files and sees one knows what happened to the other.

{
  "batch_id": "49a82536866a483a9189954d3c749243",
  "status": "queued",                    // "pending" when UPLOAD_AUTORUN=false
  "detail": null,
  "submitted_by": "priya",
  "created_at": 1756370000.0,
  "updated_at": 1756370000.0,
  "files_total": 2,
  "files_done": 0,
  "files_failed": 1,
  "current_file": null,
  "use_llm": true,                       // UPLOAD_AUTORUN_USE_LLM
  "fetch_images": true,                  // UPLOAD_AUTORUN_FETCH_IMAGES
  "runner": "inprocess",
  "totals": { "rows_total": 0, "products_built": 0, "inserted": 0,
              "backfilled": 0, "skipped_existing": 0, "rejected": 0 },
  "brands": [],
  "files": [
    { "index": 0, "filename": "catalog.csv", "status": "queued",
      "from_drop": null, "total_stages": 11, "rows_total": 1, "size_bytes": 57 },
    { "index": 1, "filename": "notes.txt", "status": "failed",
      "detail": "The file has no data rows." }
  ],
  "message": "1 file(s) accepted and queued for ingestion. 1 could not be read - see 'files' for the reason on each, and resend those."
}

from_drop — which file in this run is yours

Match on this, never on filename. An admin can assemble one run from several drops, so a run's files may contain sheets you did not send — and two senders can both upload products.csv. Matching on the name is a coincidence; matching on from_drop is exact.

Value Meaning
the drop id you were given this file is the one you sent in that drop
null the file went straight into a run and never sat in an inbox — under UPLOAD_AUTORUN=true that is every file, and the run id you hold is already the only id involved

It is the exact inverse of released_to, which points from your drop to the run that took it. Both are needed: released_to answers where did my drop go, from_drop answers whose file is this.

use_llm and fetch_images are reported, never accepted. They decide how much outbound work a run commits the host to, and this endpoint's caller is anonymous, so they come from settings — sending them in the request has no effect.

There are two levels of status, and they are not the same question. status: "queued" on the batch means it is behind the worker; on a file it means intact and not started yet. On the review-inbox path a pending batch containing queued files means something different again: those files are waiting behind a decision, not the worker.

Polling

GET /api/uploads/catalog/{batch_id} — the same body, live. No credential: the batch_id is a uuid4 handed only to whoever sent the drop, so holding it is the proof of having sent it. An id nobody issued is a 404.

Under autorun that id is already the run, so the rest of this subsection does not apply — go straight to Stage-by-stage progress. The release hop below is the UPLOAD_AUTORUN=false path, and is kept because that mode is still supported and a client that handles both needs no branch.

Your drop id stays valid for the whole lifecycle. Poll it and read the per-file status:

File status released_to Meaning
queued null Still in the inbox; nobody has looked yet
released the run id Accepted and started. Follow that id for progress and results
dismissed null An admin declined this file
// GET /api/uploads/catalog/{drop_id}, after an admin pressed Start
{ "status": "retired",
  "files": [ { "filename": "catalog.csv",
               "status": "released",
               "released_to": "8dcef8a2ad94..." } ] }

Then GET /api/uploads/catalog/{released_to} — also with no credential — for the run itself. A run an admin assembled from several drops lists every file in it, so you may see filenames batched alongside your own.

GET /api/uploads/catalog?limit=20 — your recent submissions, newest first, under {"batches": [...]}. This one does need a credential: letting an anonymous caller enumerate every sender's drops is a different thing from letting one check the id they hold. Note that production currently has no API keys configured at all (auth.api_keys_count is 0), so today this list is reachable only with an admin bearer token. Ask for a key to be issued if you need it; the two endpoints above need none.

Poll anonymously, even if you hold a credential

Ownership is matched on the sender name, exactly. A run an admin assembled from several drops carries a joined list — "alice, bob" — so a caller presenting a credential as alice gets a 404 on a run containing their own file, while the identical request without the credential returns 200. That is not a bug you can work around from the outside; it is the ownership filter doing what it says. Hold the id, send no credential.

Stage-by-stage progress

The poll body carries enough to draw the whole pipeline, not just a percentage.

On the batch:

Field Meaning
stage_names All eleven names, in order — see The eleven stages
runner "inprocess" or "dagster" — who is executing it (§2)

stage_names is served rather than left for you to hardcode, deliberately: draw the pipeline from it and your client cannot drift out of step when a stage is added or renamed. Read it once per poll; it is eleven short strings.

On each file:

Field Meaning
stage_index / stage_name Where this file is right now (0 = not started)
total_stages 11, so a progress bar needs no constant
rows_done / rows_total Progress within the current stage
stages[] The timeline — one entry per stage this file has entered

Each stages[] entry is {index, name, rows_done, rows_total, started_at, finished_at}. finished_at: null is the stage running now. The array persists after the run ends, so a finished file can still show its full timeline — the scalars above only ever describe the present moment, which is why they are not enough on their own.

// GET /api/uploads/catalog/{released_to} — mid-run
{
  "batch_id": "8dcef8a2ad94...",
  "status": "running",
  "runner": "inprocess",
  "stage_names": ["Brand Resolution & FSSAI Licence Mapping", "Row Intake & Normalisation", "..."],
  "files": [
    {
      "filename": "catalog.csv",
      "status": "running",
      "stage_index": 6,
      "stage_name": "Image Search & Contamination Filtering",
      "total_stages": 11,
      "rows_done": 120,
      "rows_total": 400,
      "stages": [
        { "index": 1, "name": "Brand Resolution & FSSAI Licence Mapping",
          "rows_done": 400, "rows_total": 400,
          "started_at": 1756612800.1, "finished_at": 1756612801.4 },
        { "index": 6, "name": "Image Search & Contamination Filtering",
          "rows_done": 120, "rows_total": 400,
          "started_at": 1756612809.7, "finished_at": null }
      ]
    }
  ]
}

Timestamps are epoch seconds as floats. Stage 6 (image search) and stage 8 (barcode enrichment) reach the network and are by far the slowest; a file sitting on either for minutes is normal, which is what rows_done is for.

Poll every few seconds. Stop when status leaves queued/running — see the table below.

Status values

Status Meaning
pending In the review inbox. Nothing has run. UPLOAD_AUTORUN=false only
retired Every file in this drop has been released or dismissed. Read the per-file status
queued Waiting for the worker. Under autorun this is the first status an upload gets
running In the pipeline now
done Every file completed
partial Some landed, some failed. Deliberately not done — four of five succeeding must not read as flat success
failed No file completed
interrupted A restart cut a run short. An admin resumes it; it never auto-restarts
cancelled Stopped by an admin before the remaining files began

What a finished run tells you

Each file in a completed run carries a result. Alongside the counts it lists every row that resolved, with the identifiers needed to reconcile the sheet against the catalogue:

"result": {
  "rows_total": 2, "inserted": 1, "backfilled": 0, "skipped_existing": 1,
  "rejected": 0, "brands": ["amul"],
  "products": [
    { "image_id": "amul_amul_butter_100g", "brand": "amul", "brand_key": "amul",
      "product_name": "Amul Butter 100g", "source_row": 2,
      "product_sku": "ACME-BUT-100", "sku_source": "sheet",
      "disposition": "inserted" },
    { "image_id": "amul_amul_ghee_1l", "brand": "amul", "brand_key": "amul",
      "product_name": "Amul Ghee 1L", "source_row": 3,
      "product_sku": "AMUL-GHE-1-001", "sku_source": "Internal",
      "disposition": "unchanged" }
  ],
  "products_truncated": false,
  "rejections": [
    { "row": 7, "product_name": "Kurkure Menthol", "size": "10g",
      "reason": "title is too short to be a real product name; image_urls: no images were found for this product" }
  ]
}

image_id — the join key, and what it is stable against

image_id is a pure deterministic function of brand, product name and pack size. Re-sending an unchanged sheet produces byte-identical ids, which is what makes the pipeline idempotent, and it is the column the catalogue deduplicates on. Store it rather than a row id.

What it is not stable against is any change to those three inputs. A pack size moving from 100g to 250g is a different SKU at a different price and is correctly a different id; so is a product name gaining or losing a brand prefix (Hot Heads vs Nestle Hot Heads). If a name is rewritten upstream, the id moves with it.

brand_key — the key the catalogue is addressed by

brand is the display name; brand_key is the identifier the catalogue is keyed on, and the two are not the same string:

brand brand_key
24 Mantra 24_mantra
Paper Boat paper_boat
coca-cola coca_cola
Own Products own_products

Use brand_key rather than normalising brand yourself. The rule (lower-case, non-alphanumerics to underscores) is stable, but deriving it is a guess and the failure is silent — a wrong key finds nothing rather than erroring.

Unbranded rows — what Own Products means, and what it does not fill in

A row with no brand — loose fruit, vegetables, greens, flowers, fish, or staples like dal and sugar — is filed under the display brand Own Products (brand_own_products) rather than having a brand guessed from its first word.

For these rows we store what your sheet said and nothing more. A shopkeeper bills from this record, so a tax code or article number we invented would be our guess wearing your letterhead:

Field For an unbranded row
product_name, size_variants from your sheet
selling_price, final_selling_price from your sheet
price_range a ±8% band around your price. Null if you sent no price
category derived from the product name — deterministic, not guessed
image_url, image_urls searched for, as with any other row
hsn_code, product_sku, barcode, fssai_license, description null, unless your sheet supplied them

Anything you do send is kept: a sheet with its own HSN, SKU or description column keeps all three. The rule is "we do not invent", not "we discard".

Two consequences worth planning for:

  • No pack-size explosion. A branded row with no size gets a plausible set (100g/250g/500g); an unbranded one does not, because a shop sells apples by whatever the customer asks for. A produce row with no weight column produces exactly one entry, with size Standard.
  • validation_status is often needs_review. For these rows that reflects a missing image, not a suspect product — the absence of price or SKU is no longer counted against them. needs_review rows are stored like any other; only rejected rows are dropped.

source_row — which line of the sheet produced this

The 1-based row number as the sender sees it on screen, header counted as row 1, so the first data row is 2. The same convention the 422 responses use.

This is many-to-one: a pack-size cell reading 100g, 200g, 500g legitimately becomes three products, and all three carry the same source_row. Rows that produced nothing are the ones absent from every entry — which is how you tell a shopkeeper "rows 6 and 11 produced nothing".

rejections[] — which rows were refused, and why

rejected is a count; rejections is the explanation, and it has always been sent. One entry per refused row:

Field Meaning
row The 1-based sheet row, same convention as source_row. null if the row could not be located
product_name The name as the sheet gave it
size The pack the refusal applies to, since one row can yield several
reason Every validation issue, joined with ;

Capped at 50 entries per file. Present on both the single-batch read and the list endpoint — only products is dropped from list responses.

disposition

disposition What happened
inserted New product, created by this run
backfilled Existed already; this sheet filled columns it had left empty
unchanged Existed already and was complete. Nothing was written

Join on image_id. It is the key the catalogue deduplicates every product by. Do not match on product_name — a name differing by one character is a different image_id and therefore a different product, so name matching silently creates duplicates instead of updating.

unchanged rows are listed too. Re-sending a sheet is the normal case and writes nothing; without them you would get an empty list back for a completely successful upload. products_truncated is true when a file resolved more than 5,000 rows (MAX_REPORTED_PRODUCTS) — the counts stay accurate, only the listing is cut.

products is returned on the single-batch read only. The list endpoints omit it, since twenty runs of several thousand rows each is not a list payload.

Which SKU wins

Your sheet's sku column Result sku_source
Has a value Preserved verbatim. The resolver is never called sheet
Blank A deterministic internal SKU is minted, e.g. AMUL-GHE-1-001 Internal
Blank, and marketplace lookup enabled A real marketplace product id the marketplace name

The third row does not occur in production: ENABLE_SKU_WEB_LOOKUP is false there, so a blank cell always yields an Internal SKU.

If your sheet carries its own sku_source column, that is preserved too and not overwritten with sheet.

A product's SKU is stable once stored. Re-sending the same product does not renumber it — a backfill keeps the stored value rather than the one that run minted. But that stability is per image_id: change the product name and you get a new image_id, a new product, and a new SKU. Another reason to join on image_id.

The eleven stages

  1. Brand Resolution & FSSAI Licence Mapping
  2. Row Intake & Normalisation
  3. Title & Category Consistency Guard
  4. Pack-Size Variant Explosion & Unit Safety
  5. Pricing Band Estimation
  6. Image Search & Contamination Filtering
  7. Marketplace & Internal SKU Resolution
  8. Barcode Retrieval & Enrichment
  9. HSN / GST Tax Enrichment
  10. Deterministic Product Validation Gate
  11. Vector Embedding & Storage

Stages 8 and 9 each run several enrichment steps internally; the stage count and numbering are unchanged.

What gets filled for a branded row, and how far to trust it

The rules above for Own Products still hold — an unbranded row is stored as you sent it. A branded row is gap-filled, and every filled value records how it was arrived at in a field_sources map on the row, so nothing has to be taken on trust:

method Means Example
sourced a real value from a real external source nutrients measured per 100 g, from Open Food Facts
derived computed from another field on the same row ean13 from barcode; tax_amount from your price
estimated a category-level or brand-level inference the keyword nutrient list before a lookup succeeds
not_applicable cannot exist for this product, and should not nutrition on a shampoo; an FSSAI food licence on a detergent
unknown we have not got it and have no source for it an FSSAI licence for a brand not in the registry

Two consequences worth reading twice:

  • not_applicable is not a gap. Roughly 30% of a general FMCG catalog is soap, shampoo and detergent. Those rows will never have nutrients or a health score. Coverage percentages are reported against the applicable rows, so they describe work remaining rather than work impossible.
  • barcode_verified: false means the pack is not confirmed. A barcode found by name match against a brand's catalogue is a real GS1 code for that brand, written to every size variant of the title. Good enough for catalog matching, dedup and nutrition lookups; not good enough for logistics, invoicing, or anything a scanner drives. Filter on it before relying on a barcode commercially.

Nothing is ever invented. There is no LLM anywhere in the nutrient, barcode, FSSAI or tax path, and a value that cannot be sourced or safely derived is left empty and labelled rather than guessed.

Barcodes arrive after the upload finishes

Barcode and nutrition enrichment need the network, so they do not hold up your response. Ingestion returns as soon as the rows are stored; a background job then fetches each brand's catalogue and fills barcodes, nutrition and health scores over the following minutes. Its id is on the batch as nutrition_job_id and its progress is at GET /api/admin/nutrition-intelligence/jobs/{job_id}.

So a row read immediately after a 200 may have no barcode yet and have one a few minutes later. That is expected, and re-reading is the only action needed.

Limits

Limit Value Env var Exceeded
Files per request 20 BATCH_MAX_FILES 413
Bytes per file 10 MB — 413
Rows per file 2,000 — file marked failed, others accepted
Bytes per request 50 MB BATCH_MAX_TOTAL_BYTES 413
Rows per request 20,000 BATCH_MAX_TOTAL_ROWS 413
Batches queued behind the running one 4 BATCH_QUEUE_MAX 429 — autorun only
Files awaiting review 200 INBOX_MAX_PENDING_FILES 429 — UPLOAD_AUTORUN=false only
Bytes awaiting review 200 MB INBOX_MAX_PENDING_BYTES 429 — UPLOAD_AUTORUN=false only

Under autorun, expect a 429 occasionally and retry. One batch runs at a time and only four may wait behind it, so a handful of uploads in quick succession will hit the ceiling. The files are staged when this happens and the message names the batch id, so an admin can resume that one instead of you resending — but a client that treats 429 as a hard failure will report a working system as broken. Back off and retry.

The two inbox ceilings apply only when UPLOAD_AUTORUN=false. They count what is still awaiting review, so starting or dismissing a drop frees its share immediately, and a 429 from them stores nothing — resend once an admin has cleared space.

Drops nobody acts on are deleted after BATCH_RETENTION_DAYS (7).


2. Orchestration (admin)

§1 is the sender's half: drop a file, poll an id. This is the other half — seeing what is waiting, deciding what runs, and choosing who runs it. Every route here is admin-only.

Getting a token

curl -s -X POST https://mcp.nearle.ai.in/api/auth/login \
  -H 'Content-Type: application/json' \
  -d '{"username":"admin","password":"<password>"}'
# -> {"access_token":"eyJ...","token_type":"bearer", ...}

Then Authorization: Bearer <access_token> on everything below. The token lasts 12 hours.

Before you hand this to anyone: there is exactly one admin account. The second interactive account is disabled by design (its password hash is deliberately unset), so there is no way to issue a scoped orchestration login. Sharing this credential shares full administrative access to the whole application — catalogue writes, cancel, resume, the training endpoints — not just the routes in this section. If that is not what you want, keep the orchestration in-house and give the sender only §1, which needs no credential at all.

The routes

Method Path Purpose
GET /api/admin/catalog-batch/inbox Everything awaiting review, grouped by sender
POST /api/admin/catalog-batch/from-inbox Tick files and run them as one batch
GET /api/admin/catalog-batch/batches?limit=20 Recent batches, slim (no per-product manifest)
GET /api/admin/catalog-batch/batches/{batch_id} One batch in full, with the stage timeline
POST /api/admin/catalog-batch/batches/{batch_id}/resume Restart a queued or interrupted batch
POST /api/admin/catalog-batch/batches/{batch_id}/cancel Stop before the remaining files begin
POST /api/admin/catalog-batch/inbox/dismiss Decline files without running them
POST /api/admin/catalog-batch/preview Parse sheets and report what was read — stores nothing
POST /api/admin/catalog-batch/ingest Upload and run directly, skipping the inbox

The two batches reads return the same shape as §1's poll, so Stage-by-stage progress applies unchanged. Use preview before ingest when a sheet's headers are in doubt: it answers the column-mapping question without writing anything.

Starting a batch

# 1. See what is waiting
curl -s https://mcp.nearle.ai.in/api/admin/catalog-batch/inbox \
  -H "Authorization: Bearer $TOKEN"
{ "pending_count": 2,
  "submissions": [
    { "submission_id": "49a82536866a...", "submitted_by": "priya",
      "created_at": 1756370000.0,
      "files": [ { "file_id": "49a82536866a...:0", "filename": "catalog.csv",
                   "rows_total": 412, "size_bytes": 57000 } ] }
  ] }
# 2. Start the ones you want, as one batch
curl -X POST https://mcp.nearle.ai.in/api/admin/catalog-batch/from-inbox \
  -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
  -d '{"file_ids":["49a82536866a...:0"],
       "runner":"inprocess","use_llm":false,"fetch_images":true}'
# -> 202, and a NEW batch_id. Poll that one for the stages.

file_ids are compound — "{batch_id}:{index}" — and you take them verbatim from file_id in the inbox response. Never assemble one by hand; a bare index is not unique across two submissions.

A partial start is a normal outcome, not an error. Malformed or stale ids are dropped silently rather than failing the request, because the admin UI polls every five seconds and a tick can legitimately refer to a file another tab started a moment ago. So check the returned batch's files for what was actually taken. If nothing was, you get a 409 telling you to refresh.

The files you started leave the inbox, and the original drop's per-file released_to is set to the new batch id — which is how the sender in §1 gets from the id they hold to the run that carries their results.

Choosing a runner

runner Who executes it
"inprocess" (default) This container's worker thread. One batch at a time behind a queue.
"dagster" Nobody, unless a Dagster instance is running to claim it.

⚠️ runner: "dagster" does not work in production, and fails silently. Dagster is a local development orchestrator: it is absent from requirements.txt, from the production compose file and from the deployed image, which never even copies the orchestration/ directory. A batch staged for it is deliberately handed to no one — it waits for dagster dev to pick it up, which in production is never. You get a batch parked at queued with the detail "Waiting for the Dagster orchestrator to pick this batch up", indefinitely, which reads exactly like a hang.

Against https://mcp.nearle.ai.in, always send "inprocess" — or omit the field, which means the same thing. To rescue one already stuck, POST /api/admin/catalog-batch/batches/{batch_id}/resume hands it to the in-process worker.

use_llm and fetch_images are chosen here rather than by the sender, because they commit the host to outbound work. fetch_images: true makes stage 6 substantially slower.


3. Operator imports

Three endpoints that write straight to the store, sales and nutrition tables — no pipeline, no review, no polling. Synchronous, with a per-row report. All three share one response shape. Form field name is file (singular).

Each also answers at an /upload alias — /api/upload/stores/upload, /api/upload/analytics/upload, /api/upload/nutrition/upload — carrying the same guard. Prefer the short form.

// success
{
  "status": "success",
  "filename": "sheet.csv",
  "rows_total": 1,
  "rows_imported": 1,
  "rows_skipped": 0,
  "errors": [],
  "message": "Successfully imported 1 store inventory items."
}

// partial — row 3 had no brand
{
  "status": "partial",
  "rows_total": 2,
  "rows_imported": 1,
  "rows_skipped": 1,
  "errors": [
    { "row": 3, "error": "missing required field(s): brand" }
  ],
  "message": "Imported 1 store inventory items; skipped 1 row(s) that were missing required fields - see 'errors'."
}

row is the number your spreadsheet shows — 1-based, counting the header — so row 3 is the second data row. At most 50 errors are returned; the rest are still skipped.

POST /api/upload/stores

Column When absent
store_id required Row skipped and reported
brand required Row skipped and reported (lowercased on write)
product_name required Row skipped and reported
image_id optional Derived from brand + product name, so a re-upload lands on the same key
store_name optional Falls back to the store_id, underscores replaced and title-cased
category, city optional Stored NULL
available_stock, reserved_stock, reorder_level, safety_stock optional Stored 0 — the schema's own default; columns are NOT NULL
mrp, cost_price, selling_price all three Prices written only when all three present. Inventory still imports; counted in prices_skipped

Extra response fields: stores_affected, prices_written, prices_skipped.

POST /api/upload/analytics

This endpoint never worked before. Its insert named a total_price column that doesn't exist on order_items (the column is line_total) and omitted the NOT NULL store_id. Every call raised and returned 500 Database import failed. If you have an integration that gave up on it, retry.

Column When absent
store_id, brand, customer_id, quantity, unit_price required Row skipped and reported
image_id or product_name one of Row skipped. image_id derived from the name when only the name is given
order_id optional Generated. Each row then becomes its own order
total_price optional Derived as quantity × unit_price. Accepted under this name; stored in line_total
order_date optional Import time, counted in dates_defaulted_to_now so you can see how much of the file isn't really dated
payment_method, delivery_status optional Stored NULL

Extra response fields: total_revenue, dates_defaulted_to_now.

customer_id is required rather than defaulted: it used to fall back to a shared cust_imported, which merges every buyer in a file into one customer and corrupts exactly the per-customer models this table feeds. Rows sharing an order_id become one order, whose order_value is recomputed as the sum of its lines.

Import each file once. order_items has a surrogate key and nothing to conflict on, so re-uploading appends the lines again.

POST /api/upload/nutrition

Only brand, plus one of image_id / product_name, is required. Every nutrient column is optional and stays NULL when omitted. Recognised: calories_kcal, protein_g, carbohydrates_g, total_sugar_g, dietary_fiber_g, total_fat_g, sodium_mg, calcium_mg, iron_mg, vitamin_c_mg, health_score, diet_tags, allergens (plus category).

What you send decides the row's data_status, which every downstream reader must check before treating a NULL as zero:

Status Awarded when
verified All seven core nutrients present — calories, protein, carbohydrates, sugar, fibre, fat, sodium
partial Some of the seven present
unavailable None present. Row still created, carrying only its identity
// two of seven nutrients
{ "rows_imported": 1,
  "data_status_counts": { "verified": 0, "partial": 1, "unavailable": 0 } }

// full core set
{ "rows_imported": 1,
  "data_status_counts": { "verified": 1, "partial": 0, "unavailable": 0 } }

Never write "None" in the allergens column. An omitted allergens column stores NULL with allergen_source: "unavailable". A literal None stores the string as an allergen name. The two must stay distinguishable: one says nobody told us, the other claims the product is allergen-free — and that claim is served by the public /api/nutrition endpoints.

Upserts use COALESCE throughout, so a later, thinner sheet can never blank a value an earlier trusted source established, and a row already verified is never downgraded by a partial upload.


4. Errors

Code Cause What to do
400 File unparseable, no data rows, or (catalogue drop) no product-name column in any file Read detail; it names the headers it found
401 Credential present but invalid; or absent on an endpoint that needs one Check the table in Authentication
403 Valid credential, wrong permission Check the role table
413 Over a file-count, byte or row ceiling Split the drop
422 Operator imports only: not one row could be imported detail.errors lists row numbers and what each was missing
429 Catalogue drop, autorun: four batches already queued Normal under load. Back off and retry; detail names the staged batch id
429 Catalogue drop, UPLOAD_AUTORUN=false: the review inbox is full Nothing was stored. Ask an admin to clear it, then resend
// 400 — wrong columns
{ "detail": "None of the uploaded files could be ingested. wrong.csv: No product name column was found. Headers read: colour, size. Recognised fields: size_variants." }

// 422 — nothing importable
{ "detail": {
    "message": "No store inventory items could be imported from 'sheet.csv'. Check the column headers against the sample template (GET /api/upload/template/...).",
    "rows_total": 1,
    "errors": [ { "row": 2, "error": "missing required field(s): store_id, brand" } ] } }

Note the shape difference: on 400, 413 and 429, detail is a string; on 422 it is an object. A client that renders detail directly will print [object Object] for a 422 unless it handles both.

Sample sheets with the correct headers: GET /api/upload/template/stores, /analytics, /nutrition — no credential needed.


5. If you already integrated

Was Is now
POST /api/uploads/catalog needed an X-API-Key No credential. Send the file; drop the header
Files waited for review (28 Aug – 31 Aug 2026) They run on arrival again. Expect queued, not pending, and no released_to hop — the id you get back is the run. UPLOAD_AUTORUN=false restores the review inbox
The drop id 404'd once an admin started it It stays valid. The file reads released and carries released_to
The result was counts only It also lists products with image_id / product_sku / disposition
Progress was one stage_index scalar Each file also carries a stages[] timeline, and the batch carries stage_names and runner — see Stage-by-stage progress
?use_llm / ?fetch_images on the drop Still ignored. Under autorun they come from UPLOAD_AUTORUN_USE_LLM / UPLOAD_AUTORUN_FETCH_IMAGES; the response reports what was used
429 meant the review inbox was full Under autorun it means the worker queue is full again — retryable, and the files were staged
200 with rows_imported: 0 422 with per-row reasons. Handle as a client error, not a server one
Missing customer_id became cust_imported Row is skipped. Supply a real customer id
Missing cost_price invented as mrp × 0.7 Prices skipped for that row; inventory still lands. Send all three
Missing nutrients and allergens got plausible defaults Stored NULL, and the row is partial/unavailable rather than verified

6. API keys are now optional

Nobody needs a key to send catalogue spreadsheets. Issue one only for a machine client that wants the credentialed reads (GET /api/uploads/catalog) or the operator imports.

python scripts/make_auth_secrets.py --api-key catalog-drop:uploader

Put the resulting name:role:secret triple in the Dokploy Environment tab, not in .env.production, and restart the service. Two reasons, both recorded in .env.production's own comments:

  1. .env.production is committed. A per-consumer key is the one credential that gets issued and revoked often, and it does not belong in git.
  2. backend/Dockerfile does COPY .env.production .env, so a value there is baked at build time — issuing or revoking would mean rebuilding an image that installs CPU torch, a build that has already failed once on disk space. settings.py calls load_dotenv() without override=True, so the process environment wins.

Container environment is fixed at creation, so the service must be recreated, not merely restarted. /api/health will then report api_keys_source: "process-env".

Constraints enforced at boot, before any request is served:

  • Format name:role:secret, comma-separated between entries.
  • role is admin, user or uploader. Keys never expire — treat one as a long-lived secret and rotate it deliberately. Two entries can be live at once, which is how you rotate without a cutover window.
  • The secret must be at least 32 characters, because /api/health publishes a digest.
  • Name the key for its function, not the person holding it. /api/health is public and reports {name, role, fingerprint} for every configured key.

7. What an open drop costs

With UPLOAD_AUTORUN=true, it costs CPU and it reaches the catalogue. This is the part to be clear-eyed about: the endpoint takes no credential, so anyone who can reach the host can cause products to be written to the live catalogue, and an ingest is an upsert with no undo. That was chosen knowingly — the requirement was uploads that run without manual intervention — but it should never be a surprise to whoever operates this next.

What still bounds it is throughput, not identity:

  • the per-request ceilings in Limits — 20 files, 50 MB, 20,000 rows;
  • one worker thread running a single batch at a time, with BATCH_QUEUE_MAX waiting behind it and a 429 past that. A sender can occupy the ingestion worker — that is what it is for — but cannot multiply it, and cannot touch the request path the healthcheck reads;
  • BATCH_RETENTION_DAYS, which reclaims staged bytes either way.

Nothing here bounds who. If the host is reachable from the open internet, put an IP allow-list on this route at the proxy — with autorun on, that is no longer defence in depth, it is the only control over who may write.

With UPLOAD_AUTORUN=false the cost is disk and nothing else until somebody looks: the endpoint queues nothing, so it cannot occupy the worker or reach the catalogue on its own. INBOX_MAX_PENDING_FILES and INBOX_MAX_PENDING_BYTES bound what unreviewed submissions occupy, past either it answers 429 and stores nothing, and dismissing a drop deletes its bytes immediately. The worst a stranger can then do is fill an inbox an admin has to decline.


8. Checking what is deployed

GET /api/health is public and answers 200 even when dependencies are down.

curl -s https://mcp.nearle.ai.in/api/health
  • POST /api/uploads/catalog with no credential and no file returns 400/422 → the open drop is live. This is what production answers today; a 401 would mean the old, credentialed build had been rolled back.
  • GET /api/uploads/catalog/<32 random hex chars> returns 404, not 401 → the anonymous read is live and the id really is the credential.
  • Which mode is running. There is no settings endpoint, so send one small sheet and read the batch status it comes back with:
    printf 'Product Name\nAmul Butter 100g\n' > /tmp/probe.csv
    curl -s -X POST https://mcp.nearle.ai.in/api/uploads/catalog \
      -F 'files=@/tmp/probe.csv' -F 'sender=deployment-probe' \
      | python -c "import json,sys; print(json.load(sys.stdin)['status'])"
    # queued  -> UPLOAD_AUTORUN is on; that row is being ingested now
    # pending -> the review inbox is on; nothing runs until an admin starts it
    
    Note what the first answer means: the probe row is ingested. Use a product name you are willing to see in the catalogue, or run this against a staging host.
  • The served schema carries the stage timeline → stages[] and stage_names are there, so a client can render the eleven stages:
    curl -s https://mcp.nearle.ai.in/openapi.json \
      | python -c "import json,sys; s=json.load(sys.stdin)['components']['schemas']; \
    

print('stage_names' in s['BatchOut']['properties'], 'stages' in s['BatchFileOut']['properties'])"

-> True True

- **`auth.api_keys_count` / `api_keys` / `api_keys_source` present** → the build
includes the auth diagnostics. Absent → the deployment predates them.
- **`auth.api_keys[].fingerprint`** answers *"is my key on this deployment?"* without
anyone sending the secret — the question a `401` cannot answer, since an undeployed
key and a wrong key fail identically. Compare against
`scripts/make_auth_secrets.py --fingerprint`.

`status: "degraded"` with `ollama: false` is expected in production: `USE_OLLAMA` is
`false` there.