Commit Graph

58 Commits

Author SHA1 Message Date
sriram
1c20bfc408 Brand Images Repairs 2026-09-02 15:52:31 +05:30
sriram
d5ec5755bf Brand image display cards 2026-09-02 13:34:40 +05:30
sriram
5399fea4cc Feature Updation 2026-09-02 12:25:36 +05:30
sriram
1de518feed Barcode Generation updates in db 2026-09-02 11:04:55 +05:30
sriram
7b3fbc47b4 Backend catalog recent updates 2026-09-01 18:02:11 +05:30
sriram
68bff007ea New updates done in JSON 2026-09-01 13:57:50 +05:30
sriram
6c7a886659 New updates on DB and JSON 2026-09-01 13:55:15 +05:30
sriram
183b65b3bd Backend fix checkbox 2026-09-01 11:25:04 +05:30
sriram
a6e78fee95 Feature Elimination backend 2026-08-31 16:31:43 +05:30
sriram
164ef90b31 Automate checklist-backend updates 2026-08-31 15:32:30 +05:30
sriram
3df2dc5991 Backend upload-automation file 2026-08-31 14:59:14 +05:30
sriram
3b1352b99d md file document updation 2026-08-31 12:10:54 +05:30
sriram
8f25842281 Dockerfile update 2026-08-31 11:36:03 +05:30
sriram
8394316907 backend store_catalog updates 2026-08-31 11:03:56 +05:30
sriram
998df898db Dagster Orchestration 2026-08-29 14:49:45 +05:30
sriram
27d53fa957 upload-catalog-integration 2026-08-29 11:21:11 +05:30
sriram
5aa2669f7d upload files without API-Key 2026-08-28 17:13:31 +05:30
Suriyakumarvijayanayagam
52f5d3be1d sheet upload fix 2026-08-28 11:22:53 +05:30
sriram
0e75d32f61 ingestion updates 2026-08-28 08:57:10 +05:30
sriram
a54bd43f8b Backend- file ingestion API Updates 2026-08-28 07:49:56 +05:30
sriram
f698720ee2 excel file update 2026-08-27 16:41:25 +05:30
Suriyakumarvijayanayagam
a9ab1fcf71 Declare tenacity - undeclared import crashed the container on boot
app/services/enrichment/barcode/retry.py imports tenacity at module scope,
and that module is on app/main.py's import path (main -> store_catalog router
-> store_catalog_pipeline -> barcode enrichment -> sources -> retry). tenacity
was in no requirements file, so the deployed image exited 1 during startup:

    File "/app/app/main.py", line 33, in <module>
        from app.api.routers import store_catalog
    ...
    File "/app/app/services/enrichment/barcode/retry.py", line 16
        from tenacity import (
    ModuleNotFoundError: No module named 'tenacity'

This surfaced as "100% CPU", not as a crash, which is why it was mis-read.
uvicorn never bound a socket, Swarm restarted the task, and each restart
re-ran the ~21s of eager pandas/scipy/sklearn imports that the analytics,
recommendations and nutrition routers pull in at module scope. On this
1-vCPU host that loop pins the only core indefinitely.

retry.py's own docstring asserted tenacity was "already a project dependency
(see requirements.txt)". It never was - corrected to say the opposite, and to
record that it is a hard startup dependency rather than an optional extra.

Verified on the deployed image, not just locally: with tenacity present the
container reaches health=healthy with restarts=0, and /, /api/health,
/api/brands, /api/system/status and /docs all return 200. Idle cost is
0.16% CPU / 168MiB. An AST scan of every import in app/, cli/, scripts/ and
serve.py against the image finds no other missing module (playwright is the
one remaining absence and is deliberate - commented out in requirements.txt
and imported lazily inside a function).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01N7bVBpxH4AbJtK7Kp3MzDR
2026-08-25 15:23:06 +05:30
sriram
1b347f91db Production Login page passcode updates 2026-08-24 22:55:06 +05:30
sriram
2482fe43e3 Admin login pasword changes 2026-08-24 13:26:41 +05:30
sriram
bbeb859d05 login passcode changes 2026-08-24 13:07:21 +05:30
sriram
b267dbab74 brand image updates 2026-08-20 17:52:32 +05:30
sriram
7bf8dc6922 Add Dagster orchestration and reduce active brands in backend 2026-08-20 16:39:54 +05:30
sriram
fbb1356e47 updates on catalog search and suggestions in backend 2026-08-20 13:03:16 +05:30
sriram
e224043e26 update backend brand-products json files 2026-08-19 16:49:18 +05:30
sriram
f41973e1cf Updates on backend file 2026-08-19 14:39:31 +05:30
sriram
a7a2430694 backend changes on test files of catalog 2026-08-18 18:42:28 +05:30
sriram
7a4583372f backend stores data file enrichment pipeline 2026-08-18 16:58:36 +05:30
sriram
691efef880 backend update auth settings and add brand catalogs 2026-08-18 12:21:37 +05:30
sriram
dba8d36175 backend updates for bulk product uploads from user 2026-08-17 15:26:49 +05:30
5e373ea20c Trim ~350MB of unreachable payload from the backend image
The image was 2.45GB, of which the venv is 1.73GB. The build was already
multi-stage and already installed CPU-only torch, so the remaining weight was
not build tooling - it was payload inside the installed packages that the
running service can never execute.

Removed in the build stage, before the runtime stage copies /opt/venv, so the
bytes never enter the final image:

  - bundled test suites (~237MB; torch/test is 83MB, pandas/tests 40MB)
  - torch/include (62MB), C++ headers for compiling against libtorch
  - torch/bin (50MB), gtest binaries and protoc; torch_shm_manager is kept

pytest and httpx2 move to requirements-dev.txt: the image copies app/, cli/,
scripts/, data/ and serve.py, never tests/, so the test stack was unusable
there regardless.

sympy was checked and deliberately kept - 'import sentence_transformers' does
pull it in through torch.fx, so removing it would break embeddings.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-16 09:52:52 +05:30
Suriyakumarvijayanayagam
dce2405e50 Stop /api/system/status 500ing when Ollama is switched off
_ensure_client() returns None - not False - when USE_OLLAMA is false, because
it returns before it ever probes. SystemStatusOut.ollama_connected is typed
bool, so pydantic rejected the None and the endpoint answered 500 on exactly
the configuration this deployment runs.

Found by smoke-testing the live host: every other read route answered 200 and
this one alone was a server error, which read like a database problem and was
not one. "Ollama is off" now reports as ollama_connected: false.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-14 18:29:16 +05:30
sriram
c17e1c7e13 env production updation 2026-08-14 16:41:18 +05:30
sriram
7153e12bdb brands json files insertion 2026-08-14 15:47:53 +05:30
Suriyakumarvijayanayagam
94f43a326e Point CORS at the frontend's real domain - it is spelled "catalouge"
The frontend loaded but could not call the API at all: every preflight from the
browser came back with no Access-Control-Allow-Origin header, so each request
was blocked client-side while the server logged nothing wrong.

API_CORS_ORIGINS was set to catalogue.nearle.ai.in. The host Traefik actually
serves is catalouge.nearle.ai.in - the o and u transposed. The correctly spelled
domain does not resolve at all, which is why checking the "frontend" only ever
returned a connection error and looked like a network problem.

Both spellings are now listed, so this keeps working if the typo is corrected in
Dokploy later. Verified against the live API: a preflight from
https://catalouge.nearle.ai.in previously returned 400 with no allow-origin.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 15:42:34 +05:30
Suriyakumarvijayanayagam
c101f2c8ba Install CPU-only torch and drop playwright - 9.2GB image on a GPU-less VPS
Measured inside the image: 2.7GB of nvidia/ CUDA libraries and 691MB of
triton/, on a single-CPU VPS with no GPU. sentence-transformers pulls torch in
transitively, and pip's default Linux wheel bundles the whole CUDA stack because
it cannot know the target has none. That is ~3.4GB of code that can never
execute, in a 9.2GB image on a 48GB disk shared with a dozen other services -
and a build here has already failed once on "no space left on device".

torch now installs first from PyTorch's CPU index, so the requirements.txt pass
finds it satisfied and leaves it alone.

playwright is commented out rather than deleted. It is the last-resort
image-search tier and the Dockerfile never installs its browser binary, so in a
container the tier is skipped at runtime regardless while the package still
costs 137MB. Its import is lazy, so absence takes the existing "not installed"
path rather than breaking anything.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-14 13:06:03 +05:30
sriram
1b56aae396 backend env updates and brand json files 2026-08-14 12:14:58 +05:30
sriram
21c675cb55 test file update in backend 2026-08-13 17:03:40 +05:30
sriram
86bb1ebd0a Merge branch 'master' of https://gitapp.workolik.com/nearle_daily/catalogue_backend 2026-08-13 16:11:24 +05:30
sriram
241fd237f8 backend apis updation 2026-08-13 15:54:29 +05:30
Suriyakumarvijayanayagam
ea0c00d68e Ship config as .env.production - Dokploy was overwriting the committed .env
The container was never getting its configuration, so it started with nothing
set and Traefik reported a Bad Gateway on every route.

Dokploy writes its own .env into the build context from the service's
Environment tab AFTER cloning the repository. That tab is empty, so it wrote a
zero-byte file over the committed one, and `COPY .env .` faithfully copied the
empty result into the image. The checkout showed it exactly: every file
timestamped 08:33, and .env alone at 08:34 with a size of 0. Inside the running
container, /app/.env was 0 bytes.

Nothing about this is visible from the outside. The build log shows the COPY
succeeding, the image is produced, and the platform reports only a 502.

Dokploy does not manage .env.production, so the config now travels under that
name and the Dockerfile copies it to /app/.env in the image. Anything set in the
Environment tab still wins at runtime, because settings.py calls load_dotenv()
without override=True.

Verified by reconstructing the build context the way Dokploy does - git archive
of HEAD, then an empty .env written over it - applying .dockerignore and the
COPY lines, and booting the result with an empty environment: /app/.env is 4938
bytes, the sign-in passwords are absent, and scripts/check_deploy.sh reports 6
passed, 0 failed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 14:13:05 +05:30
Suriyakumarvijayanayagam
5f66a797d8 Add a post-deploy smoke test
Separates three failures that are indistinguishable from a browser, and which
this deployment hit in sequence:

  502 on every route      the container is not running - it exited at startup,
                          so nothing reached the app and no route is special
  200 + database:false    the API is healthy, Postgres is not
  200 + database:true     working

It also catches AUTH_ALLOW_ANY_LOGIN being left on, by asserting that a
deliberately wrong password is rejected. That setting is a convenience locally
and a total auth bypass on a published host, and nothing else about a running
deployment looks different when it is true.

    ./scripts/check_deploy.sh
    ./scripts/check_deploy.sh http://localhost:3000

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 14:03:13 +05:30
Suriyakumarvijayanayagam
8169cafd9c Copy .env into the image - it was never reaching the container
The configuration was committed but excluded twice over: .dockerignore listed
.env, and the Dockerfile's explicit COPY lines never mentioned it. So the image
built cleanly, the container started with no configuration at all, and exited on
the first required setting - the same RuntimeError and the same Bad Gateway that
committing .env was meant to fix.

Silent in both directions. The build log shows every COPY succeeding, and the
platform reports only a 502, because the process is gone before it can say
anything. Nothing about "build completed" hints that the container has no
configuration.

Verified by reconstructing the image filesystem from the Dockerfile's COPY
lines with .dockerignore applied, then booting from it with an empty
environment: /app/.env is present, SIGNIN_PASSWORDS.txt is not, and /, /docs and
/api/health all answer 200. That reconstruction is the check the earlier "boots
from .env alone" testing was missing - it ran on the host, where .env sits next
to app/ whether the image would have contained it or not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:58:32 +05:30
Suriyakumarvijayanayagam
c078bced04 Stop an unreachable database from presenting as Bad Gateway
Two faults compounded into one symptom: with the database unreachable, every
route on the service returned 502 - including /docs, which never touches it.

_connect() passed no connect_timeout. A host that DROPS packets rather than
refusing them, which is what a firewall or a wrong DB_HOST looks like, blocked
until the OS gave up - roughly 130 seconds on Linux. Every caller inherited
that, /api/health included. Now bounded by DB_CONNECT_TIMEOUT_SECONDS,
defaulting to 5. Measured against an unroutable host: /api/health went from
hanging past 25s to answering 200 in 5.07s.

The container healthcheck then probed /api/health, so that hang timed out the
check, the container was marked unhealthy, and the platform stopped routing to
it. That is the part that turned a degraded dependency into a total outage, and
it was introduced with the healthcheck itself.

A healthcheck is a LIVENESS question, because the platform's answer to "no" is
to take the container out of service. It may only ask whether the process is
still serving HTTP. /api/health is a READINESS report - it dials Postgres and
Ollama to say whether they are reachable, and coupling the container's
existence to its dependencies is what made a running API unreachable. It now
probes "/", which is served from memory and does no I/O, so it can fail only if
the app really is gone.

Verified with an unroutable DB host: /, /docs and /openapi.json all answer 200,
and the healthcheck exits 0. With the app stopped it still exits 1.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:52:43 +05:30
Suriyakumarvijayanayagam
d81c3ea18e Fill .env with the real deployment configuration
Adds the live database, DigitalOcean Spaces and Google CSE settings supplied by
the repo owner, so the container needs nothing set in the Dokploy UI.

Three deliberate departures from the development .env this came from:

AUTH_ALLOW_ANY_LOGIN is false, not true. In development it is a convenience -
the password field is not checked, so any username signs in and `admin` reaches
the admin pages. On a host published to the internet it means anyone who finds
mcp.nearle.ai.in becomes admin by typing anything at all. The development file's
own comment says to turn it off before the backend leaves the laptop.

The auth secrets are the freshly generated ones, not the development values.
Those hashes are for the passwords DevAdmin!2026 and DevUser!2026, which sit in
plaintext in test_login_fix.py in this same repository - committing them would
have published working admin credentials alongside the hash that accepts them.
Verified: DevAdmin!2026 is now rejected with a 401.

USE_OLLAMA is false. The development value http://localhost:11434 cannot work
from inside a container, where localhost is the container rather than the VPS
host. Left on with nothing listening, /api/chat fails and every healthcheck
takes ~3s longer, because the health handler probes Ollama with a 3s timeout.
Enable it by pointing OLLAMA_BASE_URL at something the container can reach.

DB_NAME is stated explicitly rather than relying on settings.py's default,
which is what the development file was leaning on.

Verified booting from this file alone, with no environment variables: binds
3000 and 8000, mounts MCP, correct password returns a token, and both the
any-password bypass and the old development password return 401.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:44:26 +05:30
Suriyakumarvijayanayagam
56fdb82d0a Commit .env so the deployment carries its own configuration
At the repo owner's instruction, to stop the deploy depending on re-entering
config in the Dokploy UI - which is how the container ended up exiting at
startup on missing AUTH_SECRET_KEY and returning Bad Gateway.

The two database secrets are deliberately NOT in the file. settings.py calls
load_dotenv() without override=True, so a real environment variable wins over
the file; DB_HOST and DB_PASSWORD are set in Dokploy and never enter git. Two
fields to fill instead of seven.

The auth secrets ARE committed, which is worth being explicit about:
AUTH_SECRET_KEY signs every access token, so anyone with read access to this
repository can mint a valid admin token, and git history retains it after any
rotation. .gitignore records the same warning next to the exception that allows
the file. Regenerate with scripts/make_auth_secrets.py and redeploy if that
stops being an acceptable trade.

The generated sign-in passwords are written to SIGNIN_PASSWORDS.txt, which
stays ignored - only the PBKDF2 digests are in .env, and those cannot be
reversed.

Verified end to end: the app boots on 3000 and 8000 with DB_HOST/DB_PASSWORD
supplied as environment variables, and a login with the generated admin
password returns a token while a wrong password returns 401.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-08-13 13:39:20 +05:30