Two faults compounded into one symptom: with the database unreachable, every
route on the service returned 502 - including /docs, which never touches it.
_connect() passed no connect_timeout. A host that DROPS packets rather than
refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked
until the OS gave up - roughly 130 seconds on Linux. Every caller inherited
that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS,
defaulting to 5. Measured against an unroutable host: /api/health went from
hanging past 25s to answering 200 in 5.07s.
The container healthcheck then probed /api/health, so that hang timed out the
check, the container was marked unhealthy, and the platform stopped routing to
it. That is the part that turned a degraded dependency into a total outage, and
it was introduced with the healthcheck itself.
A healthcheck is a LIVENESS question, because the platform's answer to "no" is
to take the container out of service. It may only ask whether the process is
still serving HTTP. /api/health is a READINESS report - it dials Postgres and
Ollama to say whether they are reachable, and coupling the container's
existence to its dependencies is what made a running API unreachable. It now
probes "/", which is served from memory and does no I/O, so it can fail only if
the app really is gone.
Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200,
and the healthcheck exits 0. With the app stopped it still exits 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The deployment landed on mcp.nearle.ai.in, not the mcp.catalogue.nearle.ai.in
these references were written against. Only comments, README and .env.example
are affected - nothing reads the hostname at runtime - but a wrong host in the
connection snippet is a wrong host somebody pastes into an MCP client.
API_CORS_ORIGINS is unchanged: it takes the FRONTEND's origin
(catalogue.nearle.ai.in), not the API's, so moving the API does not affect it.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds a FastMCP server mounted onto the existing FastAPI app, so it ships in the
same container and answers on the same host rather than needing a process of
its own.
Fifteen tools, all read-only: catalog search and browse, nutrition facts and
health scores, healthier alternatives, per-store inventory and pricing,
discounts, trending and sales analytics. Each wraps a service function the REST
API already reaches through a GET. None of the write or compute endpoints are
exposed, because a tool list is chosen from by a model rather than by a person,
and catalog generation or model retraining is not something to leave one tool
call away.
The tools are written by hand rather than generated from the OpenAPI schema.
Mirroring all 63 routes would work, but a model picks a tool by reading its
description, and 63 near-identical generated entries is a worse thing to choose
from than a dozen written to be told apart.
Authentication reuses the access token from POST /api/auth/login - no separate
MCP credential, the same Principal and expiry as the REST API. The check lives
in one middleware rather than at the top of each tool, so a tool added later
cannot be left unguarded by forgetting a line. Note that this requires passing
include={"authorization"} to get_http_headers(), which strips that header by
default to avoid forwarding it downstream; without it the header is invisible
and every request looks unauthenticated, valid ones included.
Mounting a sub-app does not run its lifespan - only the outermost app's is
executed - so the MCP app's lifespan is chained through the FastAPI one. Without
that the endpoint accepts a connection and then fails on the first message with
a session manager that was never started.
Adds GET /api/mcp/info and POST /api/mcp/tools/{name} for the admin UI. The MCP
endpoint speaks streamable HTTP with session handling, so rendering a tool list
in the browser would otherwise mean shipping a full MCP client in React.
fastmcp needs Python 3.10+. The image is 3.11; on anything older the import
fails and the REST API starts without the MCP endpoint instead of not starting.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
/api/upload/nutrition called json.dumps() in a module that never imported
json, so every request to it raised NameError, was swallowed by the broad
except, and came back as "500 Database import failed". Import json.
Persist the three directories the app writes to at runtime. Products added
through the UI are appended to data/seed_catalogs/*.json and retrained models
are written to app/intelligence/artifacts/*.joblib; both live inside the image,
so a redeploy silently discarded them. The paths now come from settings
(DATA_DIR / SEED_CATALOG_DIR / MODEL_ARTIFACTS_DIR) so a volume can be mounted
on them, and catalog_engine.save_catalog resolves against DATA_DIR instead of
a working-directory-relative "data/", which landed somewhere different
depending on where the process was started from.
Mounting those volumes would otherwise have made things worse: Docker seeds a
named volume from the image on first use, but a bind mount starts empty and
just hides what the image shipped. A bind mount on /app/data would have left
the API with no seed catalogs, so the next product added would write a JSON
file containing only that product. The image now keeps pristine copies at
/app/.bundled, and restore_bundled_assets() tops up whatever a freshly mounted
directory is missing at startup without overwriting anything already there.
Configure CORS for the split-domain deployment: the React app is served from
catalogue.nearle.ai.in and calls the API on mcp.catalogue.nearle.ai.in, so the
frontend origin has to be in API_CORS_ORIGINS. A wrong list fails only in the
browser while the server logs a healthy 200, so the effective origins are now
logged at startup with a warning when they are localhost-only.
Fix FRONTEND_DIST, which looked for a sibling "frontend/" directory that is
actually named "catalogue_frontend/", so the single-port unified-serving branch
could never activate even with a build sitting next to it.
Rebuild the Dockerfile on the frontend's multi-stage pattern: dependencies
resolve into a venv in a build stage, the runtime stage copies only that.
Adds PYTHONUNBUFFERED so startup errors reach Dokploy's log pane, a liveness
HEALTHCHECK (/api/health answers 200 even when Postgres is down, so a database
blip cannot restart-loop the container), and an overridable PORT. The CMD execs
uvicorn so SIGTERM reaches it rather than the sh wrapper.
Add "from __future__ import annotations" to ollama_service and image_search,
which used PEP 604 unions in runtime-evaluated signatures and so could not be
imported below Python 3.10.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>