Files
catalogue_backend/app/infrastructure/persistence.py
Suriyakumarvijayanayagam 2493b86ed8 Fix nutrition upload crash, persist runtime writes, serve API on its own domain
/api/upload/nutrition called json.dumps() in a module that never imported
json, so every request to it raised NameError, was swallowed by the broad
except, and came back as "500 Database import failed". Import json.

Persist the three directories the app writes to at runtime. Products added
through the UI are appended to data/seed_catalogs/*.json and retrained models
are written to app/intelligence/artifacts/*.joblib; both live inside the image,
so a redeploy silently discarded them. The paths now come from settings
(DATA_DIR / SEED_CATALOG_DIR / MODEL_ARTIFACTS_DIR) so a volume can be mounted
on them, and catalog_engine.save_catalog resolves against DATA_DIR instead of
a working-directory-relative "data/", which landed somewhere different
depending on where the process was started from.

Mounting those volumes would otherwise have made things worse: Docker seeds a
named volume from the image on first use, but a bind mount starts empty and
just hides what the image shipped. A bind mount on /app/data would have left
the API with no seed catalogs, so the next product added would write a JSON
file containing only that product. The image now keeps pristine copies at
/app/.bundled, and restore_bundled_assets() tops up whatever a freshly mounted
directory is missing at startup without overwriting anything already there.

Configure CORS for the split-domain deployment: the React app is served from
catalogue.nearle.ai.in and calls the API on mcp.catalogue.nearle.ai.in, so the
frontend origin has to be in API_CORS_ORIGINS. A wrong list fails only in the
browser while the server logs a healthy 200, so the effective origins are now
logged at startup with a warning when they are localhost-only.

Fix FRONTEND_DIST, which looked for a sibling "frontend/" directory that is
actually named "catalogue_frontend/", so the single-port unified-serving branch
could never activate even with a build sitting next to it.

Rebuild the Dockerfile on the frontend's multi-stage pattern: dependencies
resolve into a venv in a build stage, the runtime stage copies only that.
Adds PYTHONUNBUFFERED so startup errors reach Dokploy's log pane, a liveness
HEALTHCHECK (/api/health answers 200 even when Postgres is down, so a database
blip cannot restart-loop the container), and an overridable PORT. The CMD execs
uvicorn so SIGTERM reaches it rather than the sh wrapper.

Add "from __future__ import annotations" to ollama_service and image_search,
which used PEP 604 unions in runtime-evaluated signatures and so could not be
imported below Python 3.10.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 12:30:45 +05:30

147 lines
5.3 KiB
Python

"""
First-run restore of the assets bundled into the container image.
The problem this solves
----------------------
Three directories are written to at runtime - the seed catalogs appended to by
``POST /api/user/products/add``, the generated catalogs from the ingestion
pipeline, and the ``*.joblib`` bundles the training endpoints produce. In a
container all three sit inside the image, so a redeploy silently discards
every one of them. They need to be on a volume.
Mounting a volume over them introduces the opposite problem. A *named* Docker
volume is seeded from the image the first time it is used, but a *bind* mount
starts empty and merely hides what the image had underneath. Dokploy offers
both and neither announces which you picked, so a bind mount on ``/app/data``
would leave the API running with no seed catalogs: the next product added would
write a JSON file containing only that product, and every model would report as
untrained.
The fix is to keep a pristine copy inside the image at a path nobody mounts
(``BUNDLED_ASSETS_DIR``, populated by the Dockerfile) and top up the writable
directory from it on startup.
Existing files are never overwritten. That is the whole contract: the bundle
supplies what is missing, and anything the running app has already written wins
over the copy baked into the image. Without that rule every redeploy would
revert user-added products back to the bundled catalog.
"""
from __future__ import annotations
import logging
import shutil
from pathlib import Path
from typing import Tuple
from app.infrastructure.settings import (
BUNDLED_ASSETS_DIR,
DATA_DIR,
MODEL_ARTIFACTS_DIR,
SEED_CATALOG_DIR,
)
logger = logging.getLogger(__name__)
# (subdirectory under BUNDLED_ASSETS_DIR, writable destination)
_BUNDLES: Tuple[Tuple[str, Path], ...] = (
("seed_catalogs", SEED_CATALOG_DIR),
("artifacts", MODEL_ARTIFACTS_DIR),
)
def _restore_one(source: Path, destination: Path) -> int:
"""Copy files missing from ``destination``. Returns how many were copied."""
if not source.is_dir():
return 0
destination.mkdir(parents=True, exist_ok=True)
copied = 0
for item in sorted(source.iterdir()):
if not item.is_file():
continue
target = destination / item.name
if target.exists():
continue
try:
# copy2 rather than copy: it preserves mtime, so "is this artifact
# older than the data it was trained on" stays answerable.
shutil.copy2(item, target)
copied += 1
except OSError as exc:
logger.warning("Could not restore %s -> %s: %s", item, target, exc)
return copied
def restore_bundled_assets() -> None:
"""
Top up the writable directories from the image's read-only bundle.
Safe to call on every boot: it is a no-op once the volume is populated, and
a no-op outside Docker where BUNDLED_ASSETS_DIR does not exist.
"""
# Create these regardless. A volume mounted at DATA_DIR arrives empty, and
# the ingestion pipeline writes into it without creating it first.
for path in (DATA_DIR, SEED_CATALOG_DIR, MODEL_ARTIFACTS_DIR):
try:
path.mkdir(parents=True, exist_ok=True)
except OSError as exc:
logger.error(
"Cannot create writable directory %s: %s. Uploads and trained "
"models will fail to save - check the volume's permissions.",
path,
exc,
)
if not BUNDLED_ASSETS_DIR.is_dir():
logger.debug(
"No bundled asset directory at %s - nothing to restore.", BUNDLED_ASSETS_DIR
)
return
for name, destination in _BUNDLES:
copied = _restore_one(BUNDLED_ASSETS_DIR / name, destination)
if copied:
logger.info(
"Restored %d bundled file(s) into %s (first run on this volume).",
copied,
destination,
)
_warn_if_not_persistent()
def _warn_if_not_persistent() -> None:
"""
Point out that the writable directories are still inside the image.
Only reachable when BUNDLED_ASSETS_DIR exists, i.e. in the container. If the
destinations were never mounted, everything written to them is lost on the
next redeploy - which looks exactly like the app quietly ignoring uploads,
hours later and with nothing in the logs to connect it to.
"""
unmounted = [p for p in (SEED_CATALOG_DIR, MODEL_ARTIFACTS_DIR) if not _is_mount(p)]
if unmounted:
logger.warning(
"These directories are written at runtime but do not look like "
"mount points: %s. Anything saved there - products added through "
"the UI, retrained models - will be discarded on the next "
"redeploy. Mount a volume on each (see backend/README.md).",
", ".join(str(p) for p in unmounted),
)
def _is_mount(path: Path) -> bool:
"""
Whether ``path`` sits on a different device than its parent.
A mounted volume shows up as a device-number change. Falls back to True on
error so a probe failure produces silence rather than a false alarm telling
somebody their correctly-mounted volume is broken.
"""
try:
if path.is_mount():
return True
return path.stat().st_dev != path.parent.stat().st_dev
except OSError:
return True