Brand Ingestion

This commit is contained in:
sriram
2026-09-07 17:52:45 +05:30
parent 75dd3eb3ce
commit 2749bee1a3
11 changed files with 5069 additions and 10 deletions

View File

@@ -267,6 +267,34 @@ BRAND_SYNC_INTERVAL_SECONDS=300
#ACTIVE_BRANDS=Amul,Cadbury,Hindustan Unilever,Own Products
# ---------------------------------------------------------------------------
# Brand discovery (a brand NAME -> the 11-stage pipeline)
# ---------------------------------------------------------------------------
# POST /api/admin/brand-discovery/preview finds a brand's products, and
# /ingest stages the ones an admin approved as an ordinary catalog batch.
# Every value below has a working default; none of these need to be set.
#
# Open Food Facts is the primary source and the language model is the
# supplement. OFF returns real products with real barcodes and pack sizes;
# the default OLLAMA_MODEL_NAME (qwen2.5:1.5b) will invent plausible ones, and
# nothing downstream can tell a well-formed fiction from a real product. Turn
# BRAND_DISCOVERY_USE_OFF off and the result rests on the model alone.
#
# NOTE: discovering a brand that is not in ACTIVE_BRANDS writes a complete
# catalog that no endpoint can read. The ingest route refuses with a 409 and
# names the line to add here; it is a config change plus a restart, never a
# re-ingest.
#BRAND_DISCOVERY_USE_OFF=true
#BRAND_DISCOVERY_USE_LLM=true
#BRAND_DISCOVERY_MAX_PRODUCTS=200
# Pack sizes kept per product when only the language model offers any. Stage 6
# runs an image search per exploded row, so this multiplies the slowest stage.
#BRAND_DISCOVERY_MAX_SIZES=3
# Wall-clock ceiling on the language-model half, checked between prompts. Open
# Food Facts runs first and is never subject to it.
#BRAND_DISCOVERY_DEADLINE_SECONDS=300
USE_S3=true
S3_ACCESS_KEY=your-do-spaces-key
S3_SECRET_KEY=your-do-spaces-secret