Backend- file ingestion API Updates
This commit is contained in:
@@ -144,10 +144,21 @@ If you find yourself writing business logic in this package, it belongs in
|
||||
| `embedding_refresh_job` | embed rows lacking a vector, verify the index |
|
||||
| `nutrition_enrichment_job` | fetch nutrition, score it, refit the two models |
|
||||
| `ml_training_job` | rebuild the store dataset, fit the served models, report |
|
||||
| `batch_ingestion_job` | a batch of uploaded spreadsheets -> the 11 stages -> Postgres |
|
||||
|
||||
Four jobs rather than one because they differ in cost and cadence: ingestion is
|
||||
Five jobs rather than one because they differ in cost and cadence: ingestion is
|
||||
offline and cheap, embedding drives the transformer, nutrition leaves the
|
||||
machine, and ML training depends on order history rather than catalog freshness.
|
||||
machine, ML training depends on order history rather than catalog freshness, and
|
||||
batch ingestion is event-driven - it exists only when somebody uploads files.
|
||||
|
||||
`batch_ingestion_job` shares every stage with the upload path in production;
|
||||
both call `app/core/batch_ingest.py`. Pick a batch with run config:
|
||||
|
||||
```json
|
||||
{"ops": {"batch_manifest": {"config": {"batch_id": "<id from the Admin UI>"}}}}
|
||||
```
|
||||
|
||||
Left empty it takes the oldest batch still waiting under `BATCH_UPLOAD_DIR`.
|
||||
|
||||
From the command line:
|
||||
|
||||
@@ -170,6 +181,7 @@ dagster asset materialize -m orchestration.definitions \
|
||||
| `daily_nutrition_refresh` | 04:00 | slowest, and the only one making outbound calls |
|
||||
| `weekly_ml_retrain` | Sun 05:00 | training data is a deterministic simulation, so nightly would repeat itself |
|
||||
| `seed_catalog_sensor` | every 60s | re-ingest a brand when its seed file changes |
|
||||
| `batch_upload_sensor` | every 60s | run an uploaded batch that is staged and waiting |
|
||||
|
||||
Shipping them stopped is deliberate. This machine also runs Postgres, Ollama,
|
||||
the API and Vite; a schedule that started itself the moment `dagster dev`
|
||||
@@ -239,6 +251,17 @@ allocates the whole 8GB VPS (ollama 3G, backend 2560M, postgres 1G), so there
|
||||
is no headroom for a webserver plus daemon. Orchestration is a development
|
||||
concern here.
|
||||
|
||||
**This includes `batch_ingestion_job`.** The Batch Catalog Ingestion screen on
|
||||
the site does not talk to Dagster and does not need it running. It stages the
|
||||
uploaded files, then executes the same functions this job wraps
|
||||
(`app/core/batch_ingest.py`) on a single bounded worker thread inside the API
|
||||
process - see `app/core/batch_worker.py` for why one worker rather than a
|
||||
thread per upload. The job here exists so the graph is inspectable and a batch
|
||||
can be re-run from the Launchpad on a development machine.
|
||||
|
||||
Do not turn `batch_upload_sensor` on next to a running API: both would pick up
|
||||
the same staged batch. That is why it ships STOPPED like the rest.
|
||||
|
||||
Port 3030, not Dagster's default 3000 - `serve.py` binds `PORTS=3000,8000` and
|
||||
the frontend nginx also listens on 3000.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user