Backend- file ingestion API Updates

This commit is contained in:
sriram
2026-08-28 07:49:56 +05:30
parent f698720ee2
commit a54bd43f8b
17 changed files with 2306 additions and 138 deletions

View File

@@ -144,10 +144,21 @@ If you find yourself writing business logic in this package, it belongs in
| `embedding_refresh_job` | embed rows lacking a vector, verify the index |
| `nutrition_enrichment_job` | fetch nutrition, score it, refit the two models |
| `ml_training_job` | rebuild the store dataset, fit the served models, report |
| `batch_ingestion_job` | a batch of uploaded spreadsheets -> the 11 stages -> Postgres |
Four jobs rather than one because they differ in cost and cadence: ingestion is
Five jobs rather than one because they differ in cost and cadence: ingestion is
offline and cheap, embedding drives the transformer, nutrition leaves the
machine, and ML training depends on order history rather than catalog freshness.
machine, ML training depends on order history rather than catalog freshness, and
batch ingestion is event-driven - it exists only when somebody uploads files.
`batch_ingestion_job` shares every stage with the upload path in production;
both call `app/core/batch_ingest.py`. Pick a batch with run config:
```json
{"ops": {"batch_manifest": {"config": {"batch_id": "<id from the Admin UI>"}}}}
```
Left empty it takes the oldest batch still waiting under `BATCH_UPLOAD_DIR`.
From the command line:
@@ -170,6 +181,7 @@ dagster asset materialize -m orchestration.definitions \
| `daily_nutrition_refresh` | 04:00 | slowest, and the only one making outbound calls |
| `weekly_ml_retrain` | Sun 05:00 | training data is a deterministic simulation, so nightly would repeat itself |
| `seed_catalog_sensor` | every 60s | re-ingest a brand when its seed file changes |
| `batch_upload_sensor` | every 60s | run an uploaded batch that is staged and waiting |
Shipping them stopped is deliberate. This machine also runs Postgres, Ollama,
the API and Vite; a schedule that started itself the moment `dagster dev`
@@ -239,6 +251,17 @@ allocates the whole 8GB VPS (ollama 3G, backend 2560M, postgres 1G), so there
is no headroom for a webserver plus daemon. Orchestration is a development
concern here.
**This includes `batch_ingestion_job`.** The Batch Catalog Ingestion screen on
the site does not talk to Dagster and does not need it running. It stages the
uploaded files, then executes the same functions this job wraps
(`app/core/batch_ingest.py`) on a single bounded worker thread inside the API
process - see `app/core/batch_worker.py` for why one worker rather than a
thread per upload. The job here exists so the graph is inspectable and a batch
can be re-run from the Launchpad on a development machine.
Do not turn `batch_upload_sensor` on next to a running API: both would pick up
the same staged batch. That is why it ships STOPPED like the rest.
Port 3030, not Dagster's default 3000 - `serve.py` binds `PORTS=3000,8000` and
the frontend nginx also listens on 3000.