Fix nutrition upload crash, persist runtime writes, serve API on its own domain

/api/upload/nutrition called json.dumps() in a module that never imported
json, so every request to it raised NameError, was swallowed by the broad
except, and came back as "500 Database import failed". Import json.

Persist the three directories the app writes to at runtime. Products added
through the UI are appended to data/seed_catalogs/*.json and retrained models
are written to app/intelligence/artifacts/*.joblib; both live inside the image,
so a redeploy silently discarded them. The paths now come from settings
(DATA_DIR / SEED_CATALOG_DIR / MODEL_ARTIFACTS_DIR) so a volume can be mounted
on them, and catalog_engine.save_catalog resolves against DATA_DIR instead of
a working-directory-relative "data/", which landed somewhere different
depending on where the process was started from.

Mounting those volumes would otherwise have made things worse: Docker seeds a
named volume from the image on first use, but a bind mount starts empty and
just hides what the image shipped. A bind mount on /app/data would have left
the API with no seed catalogs, so the next product added would write a JSON
file containing only that product. The image now keeps pristine copies at
/app/.bundled, and restore_bundled_assets() tops up whatever a freshly mounted
directory is missing at startup without overwriting anything already there.

Configure CORS for the split-domain deployment: the React app is served from
catalogue.nearle.ai.in and calls the API on mcp.catalogue.nearle.ai.in, so the
frontend origin has to be in API_CORS_ORIGINS. A wrong list fails only in the
browser while the server logs a healthy 200, so the effective origins are now
logged at startup with a warning when they are localhost-only.

Fix FRONTEND_DIST, which looked for a sibling "frontend/" directory that is
actually named "catalogue_frontend/", so the single-port unified-serving branch
could never activate even with a build sitting next to it.

Rebuild the Dockerfile on the frontend's multi-stage pattern: dependencies
resolve into a venv in a build stage, the runtime stage copies only that.
Adds PYTHONUNBUFFERED so startup errors reach Dokploy's log pane, a liveness
HEALTHCHECK (/api/health answers 200 even when Postgres is down, so a database
blip cannot restart-loop the container), and an overridable PORT. The CMD execs
uvicorn so SIGTERM reaches it rather than the sh wrapper.

Add "from __future__ import annotations" to ollama_service and image_search,
which used PEP 604 unions in runtime-evaluated signatures and so could not be
imported below Python 3.10.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Suriyakumarvijayanayagam
2026-08-13 12:30:45 +05:30
parent b8d93fbbf2
commit 2493b86ed8
13 changed files with 498 additions and 26 deletions

View File

@@ -55,6 +55,52 @@ def _require(name: str, *, feature_flag: str) -> str:
return value
# ---------------------------------------------------------------------------
# Writable data directories (persistence)
# ---------------------------------------------------------------------------
# Everything the running app WRITES lives under one of these three paths. They
# are settings rather than hard-coded paths because in a container they must be
# mounted on a volume - otherwise every product added through the UI and every
# retrained model is discarded the next time the image is redeployed.
#
# DATA_DIR generated catalogs (catalog_engine.save_catalog)
# SEED_CATALOG_DIR per-brand JSON catalogs, appended to by
# POST /api/user/products/add and /upload-file
# MODEL_ARTIFACTS_DIR *.joblib bundles written by the training endpoints
#
# See BUNDLED_ASSETS_DIR below for how the read-only copies shipped inside the
# image get into these directories the first time a volume is mounted.
_BACKEND_ROOT = Path(__file__).resolve().parents[2]
def _dir(name: str, default: Path) -> Path:
raw = os.getenv(name, "").strip()
return Path(raw).expanduser() if raw else default
DATA_DIR = _dir("DATA_DIR", _BACKEND_ROOT / "data")
SEED_CATALOG_DIR = _dir("SEED_CATALOG_DIR", DATA_DIR / "seed_catalogs")
MODEL_ARTIFACTS_DIR = _dir(
"MODEL_ARTIFACTS_DIR", _BACKEND_ROOT / "app" / "intelligence" / "artifacts"
)
# Pristine copies of the bundled seed catalogs and pre-trained models, placed
# here by the Dockerfile at a path that is never itself mounted over.
#
# This exists because the two ways of mounting a volume behave differently, and
# the difference is silent. Docker copies the image's content into a *named*
# volume the first time it is used, but a *bind* mount starts empty and simply
# hides whatever the image had at that path. Mounting a bind mount on /app/data
# would therefore leave the app with no seed catalogs at all: the next product
# added would write a fresh JSON file containing only that one product, and the
# ML endpoints would report no trained models.
#
# So the image keeps a second, unmounted copy, and restore_bundled_assets()
# (app/infrastructure/persistence.py) fills in whatever the writable directory
# is missing at startup. Empty/absent outside Docker, where nothing is mounted
# and the defaults above already point at the real files.
BUNDLED_ASSETS_DIR = _dir("BUNDLED_ASSETS_DIR", Path("/app/.bundled"))
# ---------------------------------------------------------------------------
# Ollama (local LLM)
# ---------------------------------------------------------------------------