# Sending spreadsheets to the catalogue **Base:** `https://mcp.nearle.ai.in` · **Formats:** `.xlsx` `.xls` `.csv` `.tsv` Four endpoints accept a spreadsheet. One is **open — no credential needed** — and stages files for review; three are operator imports that require a credential and write directly to the store, sales and nutrition tables. If you only need to **send sheets and watch them run**, §1 is the whole story and you need no credential. §2 covers the admin side — the review inbox, starting a batch, and picking who executes it. > **Deployment status.** The endpoints, the open drop and the stage timeline in > [Stage-by-stage progress](#stage-by-stage-progress) are **live on > `https://mcp.nearle.ai.in`** — verified 2026-08-31 by probing the deployment, > not by reading the source. > > **`UPLOAD_AUTORUN` is the exception: committed, not yet deployed.** Until it > ships, an upload still lands in the review inbox and answers `pending`. Both > modes are documented below and both are supported; if you are integrating > right now, write the client so it does not care which is running — > [§8 Checking what is deployed](#8-checking-what-is-deployed) shows how to tell > in one curl. --- ## Read this first These endpoints used to fill a missing column with a plausible constant. A sheet whose headers didn't match still returned `200 OK` — having written rows attributed to brand `amul` in store `store_mumbai_1` at MRP 100. Nutrition rows were worse: absent nutrients became invented numbers and an absent allergens column became `allergens = ['None']`, stored as `data_status: "verified"`. The rule now: **nothing is invented.** A missing value is stored as NULL. A row missing the fields that identify it is skipped and reported back with its row number. An upload where nothing could be imported is a `422`, not a `200` with `rows_imported: 0`. A blank allergens cell records *"we were not told"* — which is not the claim *"contains no allergens"*. Leave the column out rather than writing `None`. **Arithmetic is not invention.** `line_total` is still derived from `quantity × unit_price` when your sheet omits it, because both were supplied. That distinction is the whole design: compute from given values, never substitute for absent ones. --- ## Authentication **The catalogue drop (§1) needs no credential.** Anyone who can reach the host can send it a spreadsheet — and with `UPLOAD_AUTORUN` on, that spreadsheet runs through the pipeline immediately and its products land in the live catalogue. There is no undo; ingestion is an upsert. That is a deliberate trade, not an oversight. The requirement was uploads that run without manual intervention, and a review queue that needs an admin to press a button is not that. What bounds the endpoint is throughput rather than identity: the limits in [Limits](#limits), and a single worker thread behind a queue of `BATCH_QUEUE_MAX`. A sender can occupy the ingestion worker; they cannot multiply it. With `UPLOAD_AUTORUN=false` the older behaviour returns — files wait in an admin review inbox and the cost of an unwanted drop is disk until somebody declines it. The orchestration routes (§2) and the three operator imports (§3) write straight to the database and stay credentialed. Two credential types are accepted there: ``` X-API-Key: # machine clients Authorization: Bearer # from POST /api/auth/login, 12h lifetime ``` Send one or the other. Sending both is not rejected — the bearer token is tried first and the API key ignored — but relying on that is a way to spend an hour debugging the wrong credential. | Endpoint | Permission | Who holds it | | --- | --- | --- | | `POST /api/uploads/catalog` | **none** | anyone | | `GET /api/uploads/catalog/{batch_id}` | **none** — the id is the credential | anyone holding that id | | `GET /api/uploads/catalog` (list) | `upload_catalog` | `uploader`, `admin` | | `/api/admin/catalog-batch/*` (§2) | `admin` | `admin` only | | `/api/upload/stores` | `upload_store_inventory` | `user`, `admin` | | `/api/upload/analytics` | `manage_analytics` | `admin` | | `/api/upload/nutrition` | `manage_nutrition` | `admin` | `admin` is a superuser: `Principal.has_permission` grants it every permission rather than requiring each to be listed. On the credentialed endpoints, no credential → `401` and a valid credential without the permission → `403`. A credential that is **present but wrong** is a `401` everywhere, including on the open drop. Anonymous is allowed; mistyped is not — a typo must not silently downgrade to an anonymous submission the sender then cannot find. --- ## 1. The catalogue drop This is the one to give a colleague. No credential, up to 20 sheets per request, answers immediately with a `batch_id`. **What happens on arrival depends on one setting, `UPLOAD_AUTORUN`:** | | `true` (the intended production mode) | `false` | | --- | --- | --- | | On arrival | The batch is queued and the 11 stages run | Files wait in the admin review inbox | | Batch `status` | `queued`, then `running` | `pending` | | The id you get back | **is the run** — poll it and watch `stages[]` | is a *drop* id; the run gets a different one, reached via the file's `released_to` | | Anything to do | No | An admin must tick the files and press Start | **Write your client so it does not care which is running.** Both answer `202` with an id that `GET /api/uploads/catalog/{id}` understands; only the number of hops differs. Poll the id you were given, and follow `released_to` if a file ever grows one. ### `POST /api/uploads/catalog` → `202` ```bash curl -X POST https://mcp.nearle.ai.in/api/uploads/catalog \ -F 'files=@store-catalog.xlsx' \ -F 'files=@second-store.xlsx' \ -F 'sender=priya' ``` | Form field | | Purpose | | --- | --- | --- | | `files` | **required** | The spreadsheets. Repeat the field for more than one; up to 20. | | `sender` | optional | A label for the inbox, so the admin can see who sent what. Free text, trimmed to 60 chars. Defaults to `anonymous`. | `use_llm` and `fetch_images` are **no longer accepted here.** They decide how a run behaves and commit the host to outbound work, so the choice belongs to the admin pressing Start — not to the sender. Passing them is inert. ### Which columns are read Headers match case-insensitively by keyword, so `Product Name`, `ITEM` and `variant` all resolve to the same field. **A product-name column is the only hard requirement** — a sheet without one is rejected during the request. | Field | | Headers that match | | --- | --- | --- | | `product_name` | **required** | `product` · `item` · `variant` · `name` | | `brand` | optional | `brand` · `manufacturer` · `company` | | `category` | optional | `category` · `segment` | | `size_variants` | optional | `size` · `pack` · `weight` · `volume` · `net qty` · `quantity` | | `final_selling_price` | optional | `final price` · `selling price` · `price` · `mrp` · `rate` · `cost` | | `barcode` | optional | `barcode` · `bar code` · `gtin` · `ean` · `upc` | | `fssai_license` | optional | any header containing `fssai` | | `hsn_code` | optional | any header containing `hsn` | | `image_url` | optional | `image` · `photo` · `picture` · `url` · `link` | | `description` | optional | `description` · `desc` · `detail` | | `title` | optional | `title` | | `product_sku` | optional | `sku` | | `price_range` | optional | `price range` · `range` | | `providers` | optional | `provider` · `platform` · `marketplace` · `available at` | | `highlights` | optional | `highlight` · `feature` · `benefit` | | `nutrients` | optional | `nutrient` · `nutrition` | Rules are evaluated in the order above, so a header matching two of them takes the earlier field. Where two *columns* map to the same field, the **leftmost wins** and the other is reported as ignored. Unrecognised columns are listed back, not an error. ### Response — `202`, one good file and one bad The bad file is kept as a failed member rather than dropped, so a sender who submitted two files and sees one knows what happened to the other. ```jsonc { "batch_id": "49a82536866a483a9189954d3c749243", "status": "queued", // "pending" when UPLOAD_AUTORUN=false "detail": null, "submitted_by": "priya", "created_at": 1756370000.0, "updated_at": 1756370000.0, "files_total": 2, "files_done": 0, "files_failed": 1, "current_file": null, "use_llm": false, // UPLOAD_AUTORUN_USE_LLM "fetch_images": true, // UPLOAD_AUTORUN_FETCH_IMAGES "runner": "inprocess", "totals": { "rows_total": 0, "products_built": 0, "inserted": 0, "backfilled": 0, "skipped_existing": 0, "rejected": 0 }, "brands": [], "files": [ { "index": 0, "filename": "catalog.csv", "status": "queued", "total_stages": 11, "rows_total": 1, "size_bytes": 57 }, { "index": 1, "filename": "notes.txt", "status": "failed", "detail": "The file has no data rows." } ], "message": "1 file(s) accepted and queued for ingestion. 1 could not be read - see 'files' for the reason on each, and resend those." } ``` `use_llm` and `fetch_images` are reported, never accepted. They decide how much outbound work a run commits the host to, and this endpoint's caller is anonymous, so they come from settings — sending them in the request has no effect. There are two levels of status, and they are not the same question. `status: "queued"` on the **batch** means it is behind the worker; on a **file** it means intact and not started yet. On the review-inbox path a `pending` batch containing `queued` files means something different again: those files are waiting behind a *decision*, not the worker. ### Polling `GET /api/uploads/catalog/{batch_id}` — the same body, live. No credential: the `batch_id` is a `uuid4` handed only to whoever sent the drop, so holding it is the proof of having sent it. An id nobody issued is a `404`. Under autorun that id is already the run, so the rest of this subsection does not apply — go straight to [Stage-by-stage progress](#stage-by-stage-progress). The release hop below is the `UPLOAD_AUTORUN=false` path, and is kept because that mode is still supported and a client that handles both needs no branch. **Your drop id stays valid for the whole lifecycle.** Poll it and read the per-file `status`: | File `status` | `released_to` | Meaning | | --- | --- | --- | | `queued` | `null` | Still in the inbox; nobody has looked yet | | `released` | the run id | Accepted and started. **Follow that id for progress and results** | | `dismissed` | `null` | An admin declined this file | ```jsonc // GET /api/uploads/catalog/{drop_id}, after an admin pressed Start { "status": "retired", "files": [ { "filename": "catalog.csv", "status": "released", "released_to": "8dcef8a2ad94..." } ] } ``` Then `GET /api/uploads/catalog/{released_to}` — also with no credential — for the run itself. A run an admin assembled from several drops lists every file in it, so you may see filenames batched alongside your own. `GET /api/uploads/catalog?limit=20` — your recent submissions, newest first, under `{"batches": [...]}`. This one **does** need a credential: letting an anonymous caller enumerate every sender's drops is a different thing from letting one check the id they hold. Note that production currently has **no API keys configured** at all (`auth.api_keys_count` is `0`), so today this list is reachable only with an admin bearer token. Ask for a key to be issued if you need it; the two endpoints above need none. #### Poll anonymously, even if you hold a credential Ownership is matched on the sender name, exactly. A run an admin assembled from several drops carries a joined list — `"alice, bob"` — so a caller presenting a credential as `alice` gets a **`404` on a run containing their own file**, while the identical request *without* the credential returns `200`. That is not a bug you can work around from the outside; it is the ownership filter doing what it says. Hold the id, send no credential. ### Stage-by-stage progress The poll body carries enough to draw the whole pipeline, not just a percentage. **On the batch:** | Field | Meaning | | --- | --- | | `stage_names` | All eleven names, in order — see [The eleven stages](#the-eleven-stages) | | `runner` | `"inprocess"` or `"dagster"` — who is executing it (§2) | `stage_names` is served rather than left for you to hardcode, deliberately: draw the pipeline from it and your client cannot drift out of step when a stage is added or renamed. Read it once per poll; it is eleven short strings. **On each file:** | Field | Meaning | | --- | --- | | `stage_index` / `stage_name` | Where this file is **right now** (`0` = not started) | | `total_stages` | `11`, so a progress bar needs no constant | | `rows_done` / `rows_total` | Progress *within* the current stage | | `stages[]` | The timeline — one entry per stage this file has entered | Each `stages[]` entry is `{index, name, rows_done, rows_total, started_at, finished_at}`. **`finished_at: null` is the stage running now.** The array persists after the run ends, so a finished file can still show its full timeline — the scalars above only ever describe the present moment, which is why they are not enough on their own. ```jsonc // GET /api/uploads/catalog/{released_to} — mid-run { "batch_id": "8dcef8a2ad94...", "status": "running", "runner": "inprocess", "stage_names": ["Brand Resolution & FSSAI Licence Mapping", "Row Intake & Normalisation", "..."], "files": [ { "filename": "catalog.csv", "status": "running", "stage_index": 6, "stage_name": "Image Search & Contamination Filtering", "total_stages": 11, "rows_done": 120, "rows_total": 400, "stages": [ { "index": 1, "name": "Brand Resolution & FSSAI Licence Mapping", "rows_done": 400, "rows_total": 400, "started_at": 1756612800.1, "finished_at": 1756612801.4 }, { "index": 6, "name": "Image Search & Contamination Filtering", "rows_done": 120, "rows_total": 400, "started_at": 1756612809.7, "finished_at": null } ] } ] } ``` Timestamps are epoch seconds as floats. Stage 6 (image search) and stage 8 (barcode enrichment) reach the network and are by far the slowest; a file sitting on either for minutes is normal, which is what `rows_done` is for. Poll every few seconds. Stop when `status` leaves `queued`/`running` — see the table below. ### Status values | Status | Meaning | | --- | --- | | `pending` | In the review inbox. Nothing has run. `UPLOAD_AUTORUN=false` only | | `retired` | Every file in this drop has been released or dismissed. Read the per-file `status` | | `queued` | Waiting for the worker. Under autorun this is the first status an upload gets | | `running` | In the pipeline now | | `done` | Every file completed | | `partial` | Some landed, some failed. Deliberately not `done` — four of five succeeding must not read as flat success | | `failed` | No file completed | | `interrupted` | A restart cut a run short. An admin resumes it; it never auto-restarts | | `cancelled` | Stopped by an admin before the remaining files began | ### What a finished run tells you Each file in a completed run carries a `result`. Alongside the counts it lists **every row that resolved**, with the identifiers needed to reconcile the sheet against the catalogue: ```jsonc "result": { "rows_total": 2, "inserted": 1, "backfilled": 0, "skipped_existing": 1, "rejected": 0, "brands": ["amul"], "products": [ { "image_id": "amul_amul_butter_100g", "brand": "amul", "product_name": "Amul Butter 100g", "product_sku": "ACME-BUT-100", "sku_source": "sheet", "disposition": "inserted" }, { "image_id": "amul_amul_ghee_1l", "brand": "amul", "product_name": "Amul Ghee 1L", "product_sku": "AMUL-GHE-1-001", "sku_source": "Internal", "disposition": "unchanged" } ], "products_truncated": false } ``` | `disposition` | What happened | | --- | --- | | `inserted` | New product, created by this run | | `backfilled` | Existed already; this sheet filled columns it had left empty | | `unchanged` | Existed already and was complete. Nothing was written | **Join on `image_id`.** It is the key the catalogue deduplicates every product by. Do not match on `product_name` — a name differing by one character is a different `image_id` and therefore a different product, so name matching silently creates duplicates instead of updating. `unchanged` rows are listed too. Re-sending a sheet is the normal case and writes nothing; without them you would get an empty list back for a completely successful upload. `products_truncated` is `true` when a file resolved more than 5,000 rows (`MAX_REPORTED_PRODUCTS`) — the counts stay accurate, only the listing is cut. `products` is returned on the **single-batch read** only. The list endpoints omit it, since twenty runs of several thousand rows each is not a list payload. ### Which SKU wins | Your sheet's `sku` column | Result | `sku_source` | | --- | --- | --- | | Has a value | **Preserved verbatim.** The resolver is never called | `sheet` | | Blank | A deterministic internal SKU is minted, e.g. `AMUL-GHE-1-001` | `Internal` | | Blank, and marketplace lookup enabled | A real marketplace product id | the marketplace name | The third row does not occur in production: `ENABLE_SKU_WEB_LOOKUP` is `false` there, so a blank cell always yields an `Internal` SKU. If your sheet carries its own `sku_source` column, that is preserved too and not overwritten with `sheet`. **A product's SKU is stable once stored.** Re-sending the same product does not renumber it — a backfill keeps the stored value rather than the one that run minted. But that stability is per `image_id`: change the product name and you get a new `image_id`, a new product, and a new SKU. Another reason to join on `image_id`. ### The eleven stages 1. Brand Resolution & FSSAI Licence Mapping 2. Row Intake & Normalisation 3. Title & Category Consistency Guard 4. Pack-Size Variant Explosion & Unit Safety 5. Pricing Band Estimation 6. Image Search & Contamination Filtering 7. Marketplace & Internal SKU Resolution 8. Barcode Retrieval & Enrichment 9. HSN / GST Tax Enrichment 10. Deterministic Product Validation Gate 11. Vector Embedding & Storage ### Limits | Limit | Value | Env var | Exceeded | | --- | --- | --- | --- | | Files per request | 20 | `BATCH_MAX_FILES` | `413` | | Bytes per file | 10 MB | — | `413` | | Rows per file | 2,000 | — | file marked `failed`, others accepted | | Bytes per request | 50 MB | `BATCH_MAX_TOTAL_BYTES` | `413` | | Rows per request | 20,000 | `BATCH_MAX_TOTAL_ROWS` | `413` | | Batches queued behind the running one | 4 | `BATCH_QUEUE_MAX` | `429` — **autorun only** | | Files awaiting review | 200 | `INBOX_MAX_PENDING_FILES` | `429` — `UPLOAD_AUTORUN=false` only | | Bytes awaiting review | 200 MB | `INBOX_MAX_PENDING_BYTES` | `429` — `UPLOAD_AUTORUN=false` only | **Under autorun, expect a `429` occasionally and retry.** One batch runs at a time and only four may wait behind it, so a handful of uploads in quick succession will hit the ceiling. The files *are* staged when this happens and the message names the batch id, so an admin can resume that one instead of you resending — but a client that treats 429 as a hard failure will report a working system as broken. Back off and retry. The two inbox ceilings apply only when `UPLOAD_AUTORUN=false`. They count what is still *awaiting review*, so starting or dismissing a drop frees its share immediately, and a `429` from them stores nothing — resend once an admin has cleared space. Drops nobody acts on are deleted after `BATCH_RETENTION_DAYS` (7). --- ## 2. Orchestration (admin) §1 is the sender's half: drop a file, poll an id. This is the other half — seeing what is waiting, deciding what runs, and choosing who runs it. Every route here is `admin`-only. ### Getting a token ```bash curl -s -X POST https://mcp.nearle.ai.in/api/auth/login \ -H 'Content-Type: application/json' \ -d '{"username":"admin","password":""}' # -> {"access_token":"eyJ...","token_type":"bearer", ...} ``` Then `Authorization: Bearer ` on everything below. The token lasts 12 hours. > **Before you hand this to anyone: there is exactly one admin account.** The second > interactive account is disabled by design (its password hash is deliberately unset), > so there is no way to issue a scoped orchestration login. Sharing this credential > shares full administrative access to the whole application — catalogue writes, > cancel, resume, the training endpoints — not just the routes in this section. If > that is not what you want, keep the orchestration in-house and give the sender only > §1, which needs no credential at all. ### The routes | Method | Path | Purpose | | --- | --- | --- | | `GET` | `/api/admin/catalog-batch/inbox` | Everything awaiting review, grouped by sender | | `POST` | `/api/admin/catalog-batch/from-inbox` | Tick files and run them as one batch | | `GET` | `/api/admin/catalog-batch/batches?limit=20` | Recent batches, slim (no per-product manifest) | | `GET` | `/api/admin/catalog-batch/batches/{batch_id}` | One batch in full, with the stage timeline | | `POST` | `/api/admin/catalog-batch/batches/{batch_id}/resume` | Restart a `queued` or `interrupted` batch | | `POST` | `/api/admin/catalog-batch/batches/{batch_id}/cancel` | Stop before the remaining files begin | | `POST` | `/api/admin/catalog-batch/inbox/dismiss` | Decline files without running them | | `POST` | `/api/admin/catalog-batch/preview` | Parse sheets and report what was read — stores nothing | | `POST` | `/api/admin/catalog-batch/ingest` | Upload and run directly, skipping the inbox | The two `batches` reads return the same shape as §1's poll, so [Stage-by-stage progress](#stage-by-stage-progress) applies unchanged. Use `preview` before `ingest` when a sheet's headers are in doubt: it answers the column-mapping question without writing anything. ### Starting a batch ```bash # 1. See what is waiting curl -s https://mcp.nearle.ai.in/api/admin/catalog-batch/inbox \ -H "Authorization: Bearer $TOKEN" ``` ```jsonc { "pending_count": 2, "submissions": [ { "submission_id": "49a82536866a...", "submitted_by": "priya", "created_at": 1756370000.0, "files": [ { "file_id": "49a82536866a...:0", "filename": "catalog.csv", "rows_total": 412, "size_bytes": 57000 } ] } ] } ``` ```bash # 2. Start the ones you want, as one batch curl -X POST https://mcp.nearle.ai.in/api/admin/catalog-batch/from-inbox \ -H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \ -d '{"file_ids":["49a82536866a...:0"], "runner":"inprocess","use_llm":false,"fetch_images":true}' # -> 202, and a NEW batch_id. Poll that one for the stages. ``` **`file_ids` are compound** — `"{batch_id}:{index}"` — and you take them verbatim from `file_id` in the inbox response. Never assemble one by hand; a bare index is not unique across two submissions. **A partial start is a normal outcome, not an error.** Malformed or stale ids are dropped silently rather than failing the request, because the admin UI polls every five seconds and a tick can legitimately refer to a file another tab started a moment ago. So check the returned batch's `files` for what was *actually* taken. If nothing was, you get a `409` telling you to refresh. The files you started leave the inbox, and the original drop's per-file `released_to` is set to the new batch id — which is how the sender in §1 gets from the id they hold to the run that carries their results. ### Choosing a runner | `runner` | Who executes it | | --- | --- | | `"inprocess"` *(default)* | This container's worker thread. One batch at a time behind a queue. | | `"dagster"` | Nobody, unless a Dagster instance is running to claim it. | > ⚠️ **`runner: "dagster"` does not work in production, and fails silently.** Dagster > is a local development orchestrator: it is absent from `requirements.txt`, from the > production compose file and from the deployed image, which never even copies the > `orchestration/` directory. A batch staged for it is deliberately handed to no one — > it waits for `dagster dev` to pick it up, which in production is never. You get a > batch parked at `queued` with the detail *"Waiting for the Dagster orchestrator to > pick this batch up"*, indefinitely, which reads exactly like a hang. > > **Against `https://mcp.nearle.ai.in`, always send `"inprocess"`** — or omit the > field, which means the same thing. To rescue one already stuck, `POST > /api/admin/catalog-batch/batches/{batch_id}/resume` hands it to the in-process > worker. `use_llm` and `fetch_images` are chosen here rather than by the sender, because they commit the host to outbound work. `fetch_images: true` makes stage 6 substantially slower. --- ## 3. Operator imports Three endpoints that write straight to the store, sales and nutrition tables — no pipeline, no review, no polling. Synchronous, with a per-row report. All three share one response shape. **Form field name is `file`** (singular). Each also answers at an `/upload` alias — `/api/upload/stores/upload`, `/api/upload/analytics/upload`, `/api/upload/nutrition/upload` — carrying the same guard. Prefer the short form. ```jsonc // success { "status": "success", "filename": "sheet.csv", "rows_total": 1, "rows_imported": 1, "rows_skipped": 0, "errors": [], "message": "Successfully imported 1 store inventory items." } // partial — row 3 had no brand { "status": "partial", "rows_total": 2, "rows_imported": 1, "rows_skipped": 1, "errors": [ { "row": 3, "error": "missing required field(s): brand" } ], "message": "Imported 1 store inventory items; skipped 1 row(s) that were missing required fields - see 'errors'." } ``` `row` is the number your spreadsheet shows — 1-based, counting the header — so row 3 is the second data row. At most **50** errors are returned; the rest are still skipped. ### `POST /api/upload/stores` | Column | | When absent | | --- | --- | --- | | `store_id` | **required** | Row skipped and reported | | `brand` | **required** | Row skipped and reported (lowercased on write) | | `product_name` | **required** | Row skipped and reported | | `image_id` | optional | Derived from brand + product name, so a re-upload lands on the same key | | `store_name` | optional | Falls back to the `store_id`, underscores replaced and title-cased | | `category`, `city` | optional | Stored NULL | | `available_stock`, `reserved_stock`, `reorder_level`, `safety_stock` | optional | Stored `0` — the schema's own default; columns are NOT NULL | | `mrp`, `cost_price`, `selling_price` | **all three** | Prices written only when all three present. Inventory still imports; counted in `prices_skipped` | Extra response fields: `stores_affected`, `prices_written`, `prices_skipped`. ### `POST /api/upload/analytics` > **This endpoint never worked before.** Its insert named a `total_price` column that doesn't exist on `order_items` (the column is `line_total`) and omitted the NOT NULL `store_id`. Every call raised and returned `500 Database import failed`. If you have an integration that gave up on it, retry. | Column | | When absent | | --- | --- | --- | | `store_id`, `brand`, `customer_id`, `quantity`, `unit_price` | **required** | Row skipped and reported | | `image_id` **or** `product_name` | one of | Row skipped. `image_id` derived from the name when only the name is given | | `order_id` | optional | Generated. Each row then becomes its own order | | `total_price` | optional | Derived as `quantity × unit_price`. Accepted under this name; stored in `line_total` | | `order_date` | optional | Import time, counted in `dates_defaulted_to_now` so you can see how much of the file isn't really dated | | `payment_method`, `delivery_status` | optional | Stored NULL | Extra response fields: `total_revenue`, `dates_defaulted_to_now`. `customer_id` is required rather than defaulted: it used to fall back to a shared `cust_imported`, which merges every buyer in a file into one customer and corrupts exactly the per-customer models this table feeds. Rows sharing an `order_id` become one order, whose `order_value` is recomputed as the **sum of its lines**. **Import each file once.** `order_items` has a surrogate key and nothing to conflict on, so re-uploading appends the lines again. ### `POST /api/upload/nutrition` Only `brand`, plus one of `image_id` / `product_name`, is required. Every nutrient column is optional and stays NULL when omitted. Recognised: `calories_kcal`, `protein_g`, `carbohydrates_g`, `total_sugar_g`, `dietary_fiber_g`, `total_fat_g`, `sodium_mg`, `calcium_mg`, `iron_mg`, `vitamin_c_mg`, `health_score`, `diet_tags`, `allergens` (plus `category`). What you send decides the row's `data_status`, which every downstream reader must check before treating a NULL as zero: | Status | Awarded when | | --- | --- | | `verified` | All seven core nutrients present — calories, protein, carbohydrates, sugar, fibre, fat, sodium | | `partial` | Some of the seven present | | `unavailable` | None present. Row still created, carrying only its identity | ```jsonc // two of seven nutrients { "rows_imported": 1, "data_status_counts": { "verified": 0, "partial": 1, "unavailable": 0 } } // full core set { "rows_imported": 1, "data_status_counts": { "verified": 1, "partial": 0, "unavailable": 0 } } ``` > **Never write `"None"` in the allergens column.** An omitted allergens column stores NULL with `allergen_source: "unavailable"`. A literal `None` stores the string as an allergen name. The two must stay distinguishable: one says nobody told us, the other claims the product is allergen-free — and that claim is served by the public `/api/nutrition` endpoints. Upserts use `COALESCE` throughout, so a later, thinner sheet can never blank a value an earlier trusted source established, and a row already `verified` is never downgraded by a `partial` upload. --- ## 4. Errors | Code | Cause | What to do | | --- | --- | --- | | `400` | File unparseable, no data rows, or (catalogue drop) no product-name column in any file | Read `detail`; it names the headers it found | | `401` | Credential present but invalid; or absent on an endpoint that needs one | Check the table in Authentication | | `403` | Valid credential, wrong permission | Check the role table | | `413` | Over a file-count, byte or row ceiling | Split the drop | | `422` | **Operator imports only:** not one row could be imported | `detail.errors` lists row numbers and what each was missing | | `429` | **Catalogue drop, autorun:** four batches already queued | Normal under load. Back off and retry; `detail` names the staged batch id | | `429` | **Catalogue drop, `UPLOAD_AUTORUN=false`:** the review inbox is full | Nothing was stored. Ask an admin to clear it, then resend | ```jsonc // 400 — wrong columns { "detail": "None of the uploaded files could be ingested. wrong.csv: No product name column was found. Headers read: colour, size. Recognised fields: size_variants." } // 422 — nothing importable { "detail": { "message": "No store inventory items could be imported from 'sheet.csv'. Check the column headers against the sample template (GET /api/upload/template/...).", "rows_total": 1, "errors": [ { "row": 2, "error": "missing required field(s): store_id, brand" } ] } } ``` Note the shape difference: on `400`, `413` and `429`, `detail` is a **string**; on `422` it is an **object**. A client that renders `detail` directly will print `[object Object]` for a 422 unless it handles both. Sample sheets with the correct headers: `GET /api/upload/template/stores`, `/analytics`, `/nutrition` — no credential needed. --- ## 5. If you already integrated | Was | Is now | | --- | --- | | `POST /api/uploads/catalog` needed an `X-API-Key` | No credential. Send the file; drop the header | | Files waited for review (28 Aug – 31 Aug 2026) | **They run on arrival again.** Expect `queued`, not `pending`, and no `released_to` hop — the id you get back is the run. `UPLOAD_AUTORUN=false` restores the review inbox | | The drop id 404'd once an admin started it | It stays valid. The file reads `released` and carries `released_to` | | The result was counts only | It also lists `products` with `image_id` / `product_sku` / `disposition` | | Progress was one `stage_index` scalar | Each file also carries a `stages[]` timeline, and the batch carries `stage_names` and `runner` — see [Stage-by-stage progress](#stage-by-stage-progress) | | `?use_llm` / `?fetch_images` on the drop | Still ignored. Under autorun they come from `UPLOAD_AUTORUN_USE_LLM` / `UPLOAD_AUTORUN_FETCH_IMAGES`; the response reports what was used | | `429` meant the review inbox was full | Under autorun it means the worker queue is full again — retryable, and the files were staged | | `200` with `rows_imported: 0` | `422` with per-row reasons. Handle as a client error, not a server one | | Missing `customer_id` became `cust_imported` | Row is skipped. Supply a real customer id | | Missing `cost_price` invented as `mrp × 0.7` | Prices skipped for that row; inventory still lands. Send all three | | Missing nutrients and allergens got plausible defaults | Stored NULL, and the row is `partial`/`unavailable` rather than `verified` | --- ## 6. API keys are now optional Nobody needs a key to send catalogue spreadsheets. Issue one only for a machine client that wants the credentialed reads (`GET /api/uploads/catalog`) or the operator imports. ```bash python scripts/make_auth_secrets.py --api-key catalog-drop:uploader ``` Put the resulting `name:role:secret` triple in the **Dokploy Environment tab**, not in `.env.production`, and restart the service. Two reasons, both recorded in `.env.production`'s own comments: 1. `.env.production` is committed. A per-consumer key is the one credential that gets issued and revoked often, and it does not belong in git. 2. `backend/Dockerfile` does `COPY .env.production .env`, so a value there is baked at **build** time — issuing or revoking would mean rebuilding an image that installs CPU torch, a build that has already failed once on disk space. `settings.py` calls `load_dotenv()` **without** `override=True`, so the process environment wins. Container environment is fixed at creation, so the service must be **recreated**, not merely restarted. `/api/health` will then report `api_keys_source: "process-env"`. Constraints enforced at boot, before any request is served: - Format `name:role:secret`, comma-separated between entries. - `role` is `admin`, `user` or `uploader`. Keys never expire — treat one as a long-lived secret and rotate it deliberately. Two entries can be live at once, which is how you rotate without a cutover window. - The secret must be at least **32 characters**, because `/api/health` publishes a digest. - **Name the key for its function, not the person holding it.** `/api/health` is public and reports `{name, role, fingerprint}` for every configured key. --- ## 7. What an open drop costs **With `UPLOAD_AUTORUN=true`, it costs CPU and it reaches the catalogue.** This is the part to be clear-eyed about: the endpoint takes no credential, so anyone who can reach the host can cause products to be written to the live catalogue, and an ingest is an upsert with no undo. That was chosen knowingly — the requirement was uploads that run without manual intervention — but it should never be a surprise to whoever operates this next. What still bounds it is throughput, not identity: - the per-request ceilings in [Limits](#limits) — 20 files, 50 MB, 20,000 rows; - one worker thread running a single batch at a time, with `BATCH_QUEUE_MAX` waiting behind it and a `429` past that. A sender can occupy the ingestion worker — that is what it is for — but cannot multiply it, and cannot touch the request path the healthcheck reads; - `BATCH_RETENTION_DAYS`, which reclaims staged bytes either way. Nothing here bounds *who*. **If the host is reachable from the open internet, put an IP allow-list on this route at the proxy** — with autorun on, that is no longer defence in depth, it is the only control over who may write. **With `UPLOAD_AUTORUN=false`** the cost is disk and nothing else until somebody looks: the endpoint queues nothing, so it cannot occupy the worker or reach the catalogue on its own. `INBOX_MAX_PENDING_FILES` and `INBOX_MAX_PENDING_BYTES` bound what unreviewed submissions occupy, past either it answers `429` and stores nothing, and dismissing a drop deletes its bytes immediately. The worst a stranger can then do is fill an inbox an admin has to decline. --- ## 8. Checking what is deployed `GET /api/health` is public and answers `200` even when dependencies are down. ```bash curl -s https://mcp.nearle.ai.in/api/health ``` - **`POST /api/uploads/catalog` with no credential and no file returns `400`/`422`** → the open drop is live. This is what production answers today; a `401` would mean the old, credentialed build had been rolled back. - **`GET /api/uploads/catalog/<32 random hex chars>` returns `404`, not `401`** → the anonymous read is live and the id really is the credential. - **Which mode is running.** There is no settings endpoint, so send one small sheet and read the batch `status` it comes back with: ```bash printf 'Product Name\nAmul Butter 100g\n' > /tmp/probe.csv curl -s -X POST https://mcp.nearle.ai.in/api/uploads/catalog \ -F 'files=@/tmp/probe.csv' -F 'sender=deployment-probe' \ | python -c "import json,sys; print(json.load(sys.stdin)['status'])" # queued -> UPLOAD_AUTORUN is on; that row is being ingested now # pending -> the review inbox is on; nothing runs until an admin starts it ``` Note what the first answer means: the probe row **is ingested**. Use a product name you are willing to see in the catalogue, or run this against a staging host. - **The served schema carries the stage timeline** → `stages[]` and `stage_names` are there, so a client can render the eleven stages: ```bash curl -s https://mcp.nearle.ai.in/openapi.json \ | python -c "import json,sys; s=json.load(sys.stdin)['components']['schemas']; \ print('stage_names' in s['BatchOut']['properties'], 'stages' in s['BatchFileOut']['properties'])" # -> True True ``` - **`auth.api_keys_count` / `api_keys` / `api_keys_source` present** → the build includes the auth diagnostics. Absent → the deployment predates them. - **`auth.api_keys[].fingerprint`** answers *"is my key on this deployment?"* without anyone sending the secret — the question a `401` cannot answer, since an undeployed key and a wrong key fail identically. Compare against `scripts/make_auth_secrets.py --fingerprint`. `status: "degraded"` with `ollama: false` is expected in production: `USE_OLLAMA` is `false` there.