Backend upload-automation file

This commit is contained in:
sriram
2026-08-31 14:59:14 +05:30
parent 3b1352b99d
commit 3df2dc5991
9 changed files with 569 additions and 103 deletions

View File

@@ -179,8 +179,46 @@ BATCH_RETENTION_DAYS=7
# waits for someone to press Resume. Auto-resuming means a container stuck in a
# restart loop re-runs the heaviest work in the app on every boot, which is how
# a slow start becomes an unrecoverable spiral.
#
# Worth knowing alongside UPLOAD_AUTORUN below: with uploads running unattended,
# a redeploy in the middle of one leaves that batch "interrupted" and waiting
# for a human. It is the one place manual intervention comes back.
BATCH_AUTO_RESUME=false
# ---------------------------------------------------------------------------
# Unattended ingestion
# ---------------------------------------------------------------------------
# Whether POST /api/uploads/catalog runs the pipeline on arrival (true) or parks
# the files in the admin review inbox for someone to start by hand (false).
#
# READ THIS BEFORE CHANGING IT. That endpoint takes NO credential - it was
# opened on purpose so colleagues could send spreadsheets without one being
# issued to them. With autorun on, "anyone who can reach this host" and "anyone
# who can write to the live catalogue" are the same set of people, and an ingest
# is an upsert with no undo.
#
# What still bounds it is throughput, not identity: the per-request ceilings
# above, and BATCH_QUEUE_MAX behind a single worker. A sender can occupy the
# ingestion worker; they cannot multiply it.
#
# Set false and the review inbox comes back with no code change - the INBOX_*
# ceilings below apply only on that path.
UPLOAD_AUTORUN=true
# How an auto-started run behaves. Deliberately NOT accepted from the request:
# the sender is anonymous, and letting an anonymous caller switch on the
# expensive outbound stages is the one thing this endpoint must not allow.
#
# Images on, because a product landing without one is the failure this endpoint
# exists to avoid. Stage 6 is the slowest stage and reaches the network, but
# only one batch runs at a time, so nothing else competes with it.
#
# LLM off, because use_llm gates only description generation in stage 2, and
# production runs USE_OLLAMA=false - turning it on there buys nothing and costs
# a connection timeout per row.
UPLOAD_AUTORUN_FETCH_IMAGES=true
UPLOAD_AUTORUN_USE_LLM=false
USE_OLLAMA=true
OLLAMA_BASE_URL=http://localhost:11434
OLLAMA_MODEL_NAME=qwen2.5:1.5b