Files
catalogue_backend/docs/INGESTION_API.md
2026-08-28 17:13:31 +05:30

21 KiB
Raw Blame History

Sending spreadsheets to the catalogue

Base: https://mcp.nearle.ai.in · Formats: .xlsx .xls .csv .tsv

Four endpoints accept a spreadsheet. One is open — no credential needed — and stages files for review; three are operator imports that require a credential and write directly to the store, sales and nutrition tables.

Deployment status. The review-inbox behaviour in §1 is committed but not yet deployed. Until it is, POST /api/uploads/catalog still answers 401 without a credential. Confirm before integrating — see §7 Checking what is deployed.


Read this first

These endpoints used to fill a missing column with a plausible constant. A sheet whose headers didn't match still returned 200 OK — having written rows attributed to brand amul in store store_mumbai_1 at MRP 100. Nutrition rows were worse: absent nutrients became invented numbers and an absent allergens column became allergens = ['None'], stored as data_status: "verified".

The rule now: nothing is invented. A missing value is stored as NULL. A row missing the fields that identify it is skipped and reported back with its row number. An upload where nothing could be imported is a 422, not a 200 with rows_imported: 0.

A blank allergens cell records "we were not told" — which is not the claim "contains no allergens". Leave the column out rather than writing None.

Arithmetic is not invention. line_total is still derived from quantity × unit_price when your sheet omits it, because both were supplied. That distinction is the whole design: compute from given values, never substitute for absent ones.


Authentication

The catalogue drop (§1) needs no credential. Anyone who can reach the host can send it a spreadsheet. That is only safe because nothing sent there runs on arrival: files wait in an admin review inbox, so the cost of an unwanted drop is disk until somebody declines it — never products in the live catalogue.

The three operator imports (§2) write straight to the database and stay credentialed. Two credential types are accepted there:

X-API-Key: <your secret>          # machine clients
Authorization: Bearer <token>     # from POST /api/auth/login, 12h lifetime

Send one or the other. Sending both is not rejected — the bearer token is tried first and the API key ignored — but relying on that is a way to spend an hour debugging the wrong credential.

Endpoint Permission Who holds it
POST /api/uploads/catalog none anyone
GET /api/uploads/catalog/{batch_id} none — the id is the credential anyone holding that id
GET /api/uploads/catalog (list) upload_catalog uploader, admin
/api/upload/stores upload_store_inventory user, admin
/api/upload/analytics manage_analytics admin
/api/upload/nutrition manage_nutrition admin

admin is a superuser: Principal.has_permission grants it every permission rather than requiring each to be listed. On the credentialed endpoints, no credential → 401 and a valid credential without the permission → 403.

A credential that is present but wrong is a 401 everywhere, including on the open drop. Anonymous is allowed; mistyped is not — a typo must not silently downgrade to an anonymous submission the sender then cannot find.


1. The catalogue drop

This is the one to give a colleague. No credential, up to 20 sheets per request, answers immediately with a batch_id.

Nothing you send here runs on arrival. The files are stored and an admin sees them in the review inbox; only when they select the sheets and press Start does anything reach the 11-stage pipeline or the catalogue.

POST /api/uploads/catalog → 202

curl -X POST https://mcp.nearle.ai.in/api/uploads/catalog \
  -F 'files=@store-catalog.xlsx' \
  -F 'files=@second-store.xlsx' \
  -F 'sender=priya'
Form field Purpose
files required The spreadsheets. Repeat the field for more than one; up to 20.
sender optional A label for the inbox, so the admin can see who sent what. Free text, trimmed to 60 chars. Defaults to anonymous.

use_llm and fetch_images are no longer accepted here. They decide how a run behaves and commit the host to outbound work, so the choice belongs to the admin pressing Start — not to the sender. Passing them is inert.

Which columns are read

Headers match case-insensitively by keyword, so Product Name, ITEM and variant all resolve to the same field. A product-name column is the only hard requirement — a sheet without one is rejected during the request.

Field Headers that match
product_name required product · item · variant · name
brand optional brand · manufacturer · company
category optional category · segment
size_variants optional size · pack · weight · volume · net qty · quantity
final_selling_price optional final price · selling price · price · mrp · rate · cost
barcode optional barcode · bar code · gtin · ean · upc
fssai_license optional any header containing fssai
hsn_code optional any header containing hsn
image_url optional image · photo · picture · url · link
description optional description · desc · detail
title optional title
product_sku optional sku
price_range optional price range · range
providers optional provider · platform · marketplace · available at
highlights optional highlight · feature · benefit
nutrients optional nutrient · nutrition

Rules are evaluated in the order above, so a header matching two of them takes the earlier field. Where two columns map to the same field, the leftmost wins and the other is reported as ignored. Unrecognised columns are listed back, not an error.

Response — 202, one good file and one bad

The bad file is kept as a failed member rather than dropped, so a sender who submitted two files and sees one knows what happened to the other.

{
  "batch_id": "49a82536866a483a9189954d3c749243",
  "status": "pending",
  "detail": "Waiting for review. Nothing runs until an admin starts it.",
  "submitted_by": "priya",
  "created_at": 1756370000.0,
  "updated_at": 1756370000.0,
  "files_total": 2,
  "files_done": 0,
  "files_failed": 1,
  "current_file": null,
  "use_llm": false,
  "fetch_images": false,
  "totals": { "rows_total": 0, "products_built": 0, "inserted": 0,
              "backfilled": 0, "skipped_existing": 0, "rejected": 0 },
  "brands": [],
  "files": [
    { "index": 0, "filename": "catalog.csv", "status": "queued",
      "total_stages": 11, "rows_total": 1, "size_bytes": 57 },
    { "index": 1, "filename": "notes.txt", "status": "failed",
      "detail": "The file has no data rows." }
  ],
  "message": "1 file(s) received and waiting for review. 1 could not be read - see 'files' for the reason on each, and resend those."
}

Note the two levels of status. The batch is pending — waiting on a person. An individual file inside it reads queued, meaning it is intact and eligible to be run; it is waiting behind a decision, not behind the worker.

Polling

GET /api/uploads/catalog/{batch_id} — the same body, live. No credential: the batch_id is a uuid4 handed only to whoever sent the drop, so holding it is the proof of having sent it. An id nobody issued is a 404.

Expect pending until an admin acts. After that the batch you sent stops existing under that id — the files move into a run with an id of its own, and your drop reports 404. A 404 after pending means your files were accepted and started, not that they were lost. Ask the admin for the run id if you need to follow it.

GET /api/uploads/catalog?limit=20 — your recent submissions, newest first, under {"batches": [...]}. This one does need a credential: letting an anonymous caller enumerate every sender's drops is a different thing from letting one check the id they hold.

Status values

Status Meaning
pending In the review inbox. Nothing has run.
queued Approved and waiting for the worker
running In the pipeline now
done Every file completed
partial Some landed, some failed. Deliberately not done — four of five succeeding must not read as flat success
failed No file completed
interrupted A restart cut a run short. An admin resumes it; it never auto-restarts
cancelled Stopped by an admin before the remaining files began

The eleven stages

  1. Brand Resolution & FSSAI Licence Mapping
  2. Row Intake & Normalisation
  3. Title & Category Consistency Guard
  4. Pack-Size Variant Explosion & Unit Safety
  5. Pricing Band Estimation
  6. Image Search & Contamination Filtering
  7. Marketplace & Internal SKU Resolution
  8. Barcode Retrieval & Enrichment
  9. HSN / GST Tax Enrichment
  10. Deterministic Product Validation Gate
  11. Vector Embedding & Storage

Limits

Limit Value Env var Exceeded
Files per request 20 BATCH_MAX_FILES 413
Bytes per file 10 MB — 413
Rows per file 2,000 — file marked failed, others accepted
Bytes per request 50 MB BATCH_MAX_TOTAL_BYTES 413
Rows per request 20,000 BATCH_MAX_TOTAL_ROWS 413
Files awaiting review 200 INBOX_MAX_PENDING_FILES 429 — nothing stored
Bytes awaiting review 200 MB INBOX_MAX_PENDING_BYTES 429 — nothing stored

The last two are the ceiling on the inbox as a whole, across every sender. They count only what is still awaiting review, so starting or dismissing a drop frees its share immediately. A 429 here stores nothing — resend once an admin has cleared space.

Drops nobody acts on are deleted after BATCH_RETENTION_DAYS (7).


2. Operator imports

Three endpoints that write straight to the store, sales and nutrition tables — no pipeline, no review, no polling. Synchronous, with a per-row report. All three share one response shape. Form field name is file (singular).

Each also answers at an /upload alias — /api/upload/stores/upload, /api/upload/analytics/upload, /api/upload/nutrition/upload — carrying the same guard. Prefer the short form.

// success
{
  "status": "success",
  "filename": "sheet.csv",
  "rows_total": 1,
  "rows_imported": 1,
  "rows_skipped": 0,
  "errors": [],
  "message": "Successfully imported 1 store inventory items."
}

// partial — row 3 had no brand
{
  "status": "partial",
  "rows_total": 2,
  "rows_imported": 1,
  "rows_skipped": 1,
  "errors": [
    { "row": 3, "error": "missing required field(s): brand" }
  ],
  "message": "Imported 1 store inventory items; skipped 1 row(s) that were missing required fields - see 'errors'."
}

row is the number your spreadsheet shows — 1-based, counting the header — so row 3 is the second data row. At most 50 errors are returned; the rest are still skipped.

POST /api/upload/stores

Column When absent
store_id required Row skipped and reported
brand required Row skipped and reported (lowercased on write)
product_name required Row skipped and reported
image_id optional Derived from brand + product name, so a re-upload lands on the same key
store_name optional Falls back to the store_id, underscores replaced and title-cased
category, city optional Stored NULL
available_stock, reserved_stock, reorder_level, safety_stock optional Stored 0 — the schema's own default; columns are NOT NULL
mrp, cost_price, selling_price all three Prices written only when all three present. Inventory still imports; counted in prices_skipped

Extra response fields: stores_affected, prices_written, prices_skipped.

POST /api/upload/analytics

This endpoint never worked before. Its insert named a total_price column that doesn't exist on order_items (the column is line_total) and omitted the NOT NULL store_id. Every call raised and returned 500 Database import failed. If you have an integration that gave up on it, retry.

Column When absent
store_id, brand, customer_id, quantity, unit_price required Row skipped and reported
image_id or product_name one of Row skipped. image_id derived from the name when only the name is given
order_id optional Generated. Each row then becomes its own order
total_price optional Derived as quantity × unit_price. Accepted under this name; stored in line_total
order_date optional Import time, counted in dates_defaulted_to_now so you can see how much of the file isn't really dated
payment_method, delivery_status optional Stored NULL

Extra response fields: total_revenue, dates_defaulted_to_now.

customer_id is required rather than defaulted: it used to fall back to a shared cust_imported, which merges every buyer in a file into one customer and corrupts exactly the per-customer models this table feeds. Rows sharing an order_id become one order, whose order_value is recomputed as the sum of its lines.

Import each file once. order_items has a surrogate key and nothing to conflict on, so re-uploading appends the lines again.

POST /api/upload/nutrition

Only brand, plus one of image_id / product_name, is required. Every nutrient column is optional and stays NULL when omitted. Recognised: calories_kcal, protein_g, carbohydrates_g, total_sugar_g, dietary_fiber_g, total_fat_g, sodium_mg, calcium_mg, iron_mg, vitamin_c_mg, health_score, diet_tags, allergens (plus category).

What you send decides the row's data_status, which every downstream reader must check before treating a NULL as zero:

Status Awarded when
verified All seven core nutrients present — calories, protein, carbohydrates, sugar, fibre, fat, sodium
partial Some of the seven present
unavailable None present. Row still created, carrying only its identity
// two of seven nutrients
{ "rows_imported": 1,
  "data_status_counts": { "verified": 0, "partial": 1, "unavailable": 0 } }

// full core set
{ "rows_imported": 1,
  "data_status_counts": { "verified": 1, "partial": 0, "unavailable": 0 } }

Never write "None" in the allergens column. An omitted allergens column stores NULL with allergen_source: "unavailable". A literal None stores the string as an allergen name. The two must stay distinguishable: one says nobody told us, the other claims the product is allergen-free — and that claim is served by the public /api/nutrition endpoints.

Upserts use COALESCE throughout, so a later, thinner sheet can never blank a value an earlier trusted source established, and a row already verified is never downgraded by a partial upload.


3. Errors

Code Cause What to do
400 File unparseable, no data rows, or (catalogue drop) no product-name column in any file Read detail; it names the headers it found
401 Credential present but invalid; or absent on an endpoint that needs one Check the table in Authentication
403 Valid credential, wrong permission Check the role table
413 Over a file-count, byte or row ceiling Split the drop
422 Operator imports only: not one row could be imported detail.errors lists row numbers and what each was missing
429 Catalogue drop: the review inbox is full Nothing was stored. Ask an admin to clear it, then resend
// 400 — wrong columns
{ "detail": "None of the uploaded files could be ingested. wrong.csv: No product name column was found. Headers read: colour, size. Recognised fields: size_variants." }

// 422 — nothing importable
{ "detail": {
    "message": "No store inventory items could be imported from 'sheet.csv'. Check the column headers against the sample template (GET /api/upload/template/...).",
    "rows_total": 1,
    "errors": [ { "row": 2, "error": "missing required field(s): store_id, brand" } ] } }

Note the shape difference: on 400, 413 and 429, detail is a string; on 422 it is an object. A client that renders detail directly will print [object Object] for a 422 unless it handles both.

Sample sheets with the correct headers: GET /api/upload/template/stores, /analytics, /nutrition — no credential needed.


4. If you already integrated

Was Is now
POST /api/uploads/catalog needed an X-API-Key No credential. Send the file; drop the header
Files ran on arrival Files wait for review. Expect pending, not queued
?use_llm / ?fetch_images on the drop Ignored. The admin chooses at Start
429 meant the worker queue was full 429 now means the review inbox is full
200 with rows_imported: 0 422 with per-row reasons. Handle as a client error, not a server one
Missing customer_id became cust_imported Row is skipped. Supply a real customer id
Missing cost_price invented as mrp × 0.7 Prices skipped for that row; inventory still lands. Send all three
Missing nutrients and allergens got plausible defaults Stored NULL, and the row is partial/unavailable rather than verified

5. API keys are now optional

Nobody needs a key to send catalogue spreadsheets. Issue one only for a machine client that wants the credentialed reads (GET /api/uploads/catalog) or the operator imports.

python scripts/make_auth_secrets.py --api-key catalog-drop:uploader

Put the resulting name:role:secret triple in the Dokploy Environment tab, not in .env.production, and restart the service. Two reasons, both recorded in .env.production's own comments:

  1. .env.production is committed. A per-consumer key is the one credential that gets issued and revoked often, and it does not belong in git.
  2. backend/Dockerfile does COPY .env.production .env, so a value there is baked at build time — issuing or revoking would mean rebuilding an image that installs CPU torch, a build that has already failed once on disk space. settings.py calls load_dotenv() without override=True, so the process environment wins.

Container environment is fixed at creation, so the service must be recreated, not merely restarted. /api/health will then report api_keys_source: "process-env".

Constraints enforced at boot, before any request is served:

  • Format name:role:secret, comma-separated between entries.
  • role is admin, user or uploader. Keys never expire — treat one as a long-lived secret and rotate it deliberately. Two entries can be live at once, which is how you rotate without a cutover window.
  • The secret must be at least 32 characters, because /api/health publishes a digest.
  • Name the key for its function, not the person holding it. /api/health is public and reports {name, role, fingerprint} for every configured key.

6. What an open drop costs

Disk, and nothing else, until somebody looks at it. The endpoint queues nothing, so it cannot occupy the ingestion worker and cannot reach the catalogue on its own — the review gate is what makes accepting files from anyone acceptable.

INBOX_MAX_PENDING_FILES and INBOX_MAX_PENDING_BYTES bound what unreviewed submissions can occupy; past either the endpoint answers 429 and stores nothing. Dismissing a drop deletes its bytes immediately, and retention reclaims anything nobody ever looks at.

If the host is reachable from the open internet, consider an IP allow-list at the proxy as defence in depth. The controls above bound the damage; they do not stop a stranger from filling the inbox with sheets an admin then has to decline.


7. Checking what is deployed

GET /api/health is public and answers 200 even when dependencies are down.

curl -s https://mcp.nearle.ai.in/api/health
  • POST /api/uploads/catalog with no credential and no file returns 400/422 → the open drop is live. A 401 means the old, credentialed build is still running.
  • auth.api_keys_count / api_keys / api_keys_source present → the build includes the auth diagnostics. Absent → the deployment predates them.
  • auth.api_keys[].fingerprint answers "is my key on this deployment?" without anyone sending the secret — the question a 401 cannot answer, since an undeployed key and a wrong key fail identically. Compare against scripts/make_auth_secrets.py --fingerprint.

status: "degraded" with ollama: false is expected in production: USE_OLLAMA is false there.