40 Commits

Author SHA1 Message Date
sriram
c0601b65fe Brand Discovery-LLM Updates 2026-09-29 16:42:49 +05:30
sriram
c0489d89d6 Updates on Image search using vectors 2026-09-28 15:44:01 +05:30
sriram
6628207810 Non-consumable products updation 2026-09-25 11:48:52 +05:30
sriram
fe208f4715 Image capturing Flow updates 2026-09-24 17:18:38 +05:30
sriram
f933ea10a1 Image vector to product details 2026-09-19 15:39:53 +05:30
sriram
bc786b1c49 GET api update for image vector 2026-09-18 17:50:14 +05:30
sriram
d6296bd1f0 imag vector generation with dimentionality reduction 2026-09-18 15:26:25 +05:30
sriram
afa0bfa743 image vector dimensionality reduction 2026-09-17 14:21:48 +05:30
sriram
deae694a1f Image Status check 2026-09-16 17:00:29 +05:30
sriram
9c8dbf1759 Image vector embedding 2026-09-16 16:33:39 +05:30
sriram
ce4fa70dee Brand valid image generation 2026-09-11 15:47:56 +05:30
sriram
e1a5962f82 unpopular brand generation 2026-09-10 17:51:05 +05:30
sriram
d5a23f6456 product generation with validation check 2026-09-10 16:17:28 +05:30
sriram
10b24c6348 Catalog feature updates on column fields 2026-09-08 15:18:29 +05:30
sriram
2749bee1a3 Brand Ingestion 2026-09-07 17:52:45 +05:30
sriram
75dd3eb3ce Valid Barcode Generation 2026-09-07 15:45:25 +05:30
sriram
d3709c1a4f product json updates 2026-09-05 12:14:16 +05:30
sriram
7ac3571b59 Backend scores updates on products 2026-09-05 11:53:13 +05:30
sriram
bb30edaffa image repairment in own products 2026-09-04 15:44:07 +05:30
sriram
8059120ce3 image repair on existing products 2026-09-04 13:23:43 +05:30
sriram
0c3e23fac4 Health score updates in backend 2026-09-04 12:06:40 +05:30
sriram
d5ec5755bf Brand image display cards 2026-09-02 13:34:40 +05:30
sriram
7b3fbc47b4 Backend catalog recent updates 2026-09-01 18:02:11 +05:30
sriram
6c7a886659 New updates on DB and JSON 2026-09-01 13:55:15 +05:30
sriram
164ef90b31 Automate checklist-backend updates 2026-08-31 15:32:30 +05:30
sriram
8394316907 backend store_catalog updates 2026-08-31 11:03:56 +05:30
sriram
998df898db Dagster Orchestration 2026-08-29 14:49:45 +05:30
Suriyakumarvijayanayagam
52f5d3be1d sheet upload fix 2026-08-28 11:22:53 +05:30
Suriyakumarvijayanayagam
a9ab1fcf71 Declare tenacity - undeclared import crashed the container on boot
app/services/enrichment/barcode/retry.py imports tenacity at module scope,
and that module is on app/main.py's import path (main -> store_catalog router
-> store_catalog_pipeline -> barcode enrichment -> sources -> retry). tenacity
was in no requirements file, so the deployed image exited 1 during startup:

    File "/app/app/main.py", line 33, in <module>
        from app.api.routers import store_catalog
    ...
    File "/app/app/services/enrichment/barcode/retry.py", line 16
        from tenacity import (
    ModuleNotFoundError: No module named 'tenacity'

This surfaced as "100% CPU", not as a crash, which is why it was mis-read.
uvicorn never bound a socket, Swarm restarted the task, and each restart
re-ran the ~21s of eager pandas/scipy/sklearn imports that the analytics,
recommendations and nutrition routers pull in at module scope. On this
1-vCPU host that loop pins the only core indefinitely.

retry.py's own docstring asserted tenacity was "already a project dependency
(see requirements.txt)". It never was - corrected to say the opposite, and to
record that it is a hard startup dependency rather than an optional extra.

Verified on the deployed image, not just locally: with tenacity present the
container reaches health=healthy with restarts=0, and /, /api/health,
/api/brands, /api/system/status and /docs all return 200. Idle cost is
0.16% CPU / 168MiB. An AST scan of every import in app/, cli/, scripts/ and
serve.py against the image finds no other missing module (playwright is the
one remaining absence and is deliberate - commented out in requirements.txt
and imported lazily inside a function).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7bVBpxH4AbJtK7Kp3MzDR
2026-08-25 15:23:06 +05:30
sriram
7bf8dc6922 Add Dagster orchestration and reduce active brands in backend 2026-08-20 16:39:54 +05:30
sriram
fbb1356e47 updates on catalog search and suggestions in backend 2026-08-20 13:03:16 +05:30
sriram
f41973e1cf Updates on backend file 2026-08-19 14:39:31 +05:30
sriram
7a4583372f backend stores data file enrichment pipeline 2026-08-18 16:58:36 +05:30
sriram
dba8d36175 backend updates for bulk product uploads from user 2026-08-17 15:26:49 +05:30
sriram
1b56aae396 backend env updates and brand json files 2026-08-14 12:14:58 +05:30
sriram
86bb1ebd0a Merge branch 'master' of https://gitapp.workolik.com/nearle_daily/catalogue_backend 2026-08-13 16:11:24 +05:30
sriram
241fd237f8 backend apis updation 2026-08-13 15:54:29 +05:30
Suriyakumarvijayanayagam
c078bced04 Stop an unreachable database from presenting as Bad Gateway
Two faults compounded into one symptom: with the database unreachable, every
route on the service returned 502 - including /docs, which never touches it.

_connect() passed no connect_timeout. A host that DROPS packets rather than
refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked
until the OS gave up - roughly 130 seconds on Linux. Every caller inherited
that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS,
defaulting to 5. Measured against an unroutable host: /api/health went from
hanging past 25s to answering 200 in 5.07s.

The container healthcheck then probed /api/health, so that hang timed out the
check, the container was marked unhealthy, and the platform stopped routing to
it. That is the part that turned a degraded dependency into a total outage, and
it was introduced with the healthcheck itself.

A healthcheck is a LIVENESS question, because the platform's answer to "no" is
to take the container out of service. It may only ask whether the process is
still serving HTTP. /api/health is a READINESS report - it dials Postgres and
Ollama to say whether they are reachable, and coupling the container's
existence to its dependencies is what made a running API unreachable. It now
probes "/", which is served from memory and does no I/O, so it can fail only if
the app really is gone.

Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200,
and the healthcheck exits 0. With the app stopped it still exits 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:52:43 +05:30
Suriyakumarvijayanayagam
2493b86ed8 Fix nutrition upload crash, persist runtime writes, serve API on its own domain
/api/upload/nutrition called json.dumps() in a module that never imported
json, so every request to it raised NameError, was swallowed by the broad
except, and came back as "500 Database import failed". Import json.

Persist the three directories the app writes to at runtime. Products added
through the UI are appended to data/seed_catalogs/*.json and retrained models
are written to app/intelligence/artifacts/*.joblib; both live inside the image,
so a redeploy silently discarded them. The paths now come from settings
(DATA_DIR / SEED_CATALOG_DIR / MODEL_ARTIFACTS_DIR) so a volume can be mounted
on them, and catalog_engine.save_catalog resolves against DATA_DIR instead of
a working-directory-relative "data/", which landed somewhere different
depending on where the process was started from.

Mounting those volumes would otherwise have made things worse: Docker seeds a
named volume from the image on first use, but a bind mount starts empty and
just hides what the image shipped. A bind mount on /app/data would have left
the API with no seed catalogs, so the next product added would write a JSON
file containing only that product. The image now keeps pristine copies at
/app/.bundled, and restore_bundled_assets() tops up whatever a freshly mounted
directory is missing at startup without overwriting anything already there.

Configure CORS for the split-domain deployment: the React app is served from
catalogue.nearle.ai.in and calls the API on mcp.catalogue.nearle.ai.in, so the
frontend origin has to be in API_CORS_ORIGINS. A wrong list fails only in the
browser while the server logs a healthy 200, so the effective origins are now
logged at startup with a warning when they are localhost-only.

Fix FRONTEND_DIST, which looked for a sibling "frontend/" directory that is
actually named "catalogue_frontend/", so the single-port unified-serving branch
could never activate even with a build sitting next to it.

Rebuild the Dockerfile on the frontend's multi-stage pattern: dependencies
resolve into a venv in a build stage, the runtime stage copies only that.
Adds PYTHONUNBUFFERED so startup errors reach Dokploy's log pane, a liveness
HEALTHCHECK (/api/health answers 200 even when Postgres is down, so a database
blip cannot restart-loop the container), and an overridable PORT. The CMD execs
uvicorn so SIGTERM reaches it rather than the sh wrapper.

Add "from __future__ import annotations" to ollama_service and image_search,
which used PEP 604 unions in runtime-evaluated signatures and so could not be
imported below Python 3.10.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 12:30:45 +05:30
sriram
c2af4556c6 updates on the backend 2026-08-11 19:16:01 +05:30