Commit Graph

11 Commits

Author SHA1 Message Date
781757d377 fix(deploy): harden runtime configuration
Production answered 502 on every url, /favicon.ico included, because the
container had no AUTH_SECRET and the boot check called process.exit(1): the
container died, so Traefik had no upstream and the one line explaining it was
trapped inside a restart-looping container. The exit is already gone (0dc865b).
This makes the configuration contract itself hard to get wrong.

One required variable, one resolver, two probe endpoints.

shared/config/authSecret.ts is now the only place the signing secret is
resolved. sessionToken.ts (HMAC of the identity cookie) and tokenStore.ts
(AES-256-GCM key for the platform token bundle) each read process.env
independently before, under rules that disagreed — one accepted a
whitespace-only value the other rejected. It accepts AUTH_SECRET, or
AUTH_SECRET_FILE for the Docker/Swarm secret convention when a dashboard field
mangles a value, trims both, and is a pure function of the environment.

That purity is load-bearing. Next 16 compiles proxy.ts for the NODE runtime
(its own docs: "Proxy defaults to using the Node.js runtime"), confirmed in the
build output — the proxy is in .next/server/chunks, not .next/server/edge. But
the proxy entry and the route entries are still SEPARATE BUNDLES with their own
copy of this module, so a secret invented in module scope would differ between
them, the proxy would reject every cookie the login route signed, and /login
would redirect forever. There is no generated fallback and there must not be.

LOYALY_API_BASE is no longer required in production. It accepted exactly one
origin, so an unset value could never have meant another, and requiring it added
a failure mode without adding a choice. Verified against the installed @next/env:
a real variable set to the EMPTY STRING is left empty and .env is NOT consulted,
so one blank dashboard field defeated the value shipped in the image and took
production down with "required in production". Any other host set explicitly is
still rejected by name, platform.loyaly.ai included.

/api/health and /api/ready are split. Health was returning 503 on a
configuration fault — readiness semantics on the name every orchestrator probes
by default. A Dockerfile HEALTHCHECK pointed there for one commit, and because
Dokploy runs applications as Swarm services, Swarm removed the task from the
load balancer and rescheduled it: the container was up, serving a 503 that named
the fault, and nothing could reach it to read that 503. Health is now liveness
and always 200 while the process answers; ready is readiness and 503 while a
variable is missing, for a DEPLOY gate (Order start-first + FailureAction
rollback) where failing keeps the previous good task serving. No HEALTHCHECK is
reintroduced.

Diagnostics answer the question that could not be answered from outside the
container: whether the variable never arrived or arrived empty, the secret's
source and length (never its value), and any environment variable whose NAME is
a near-miss for AUTH_SECRET — wrapped (NEXT_PUBLIC_AUTH_SECRET) or mistyped
(AUTH_SECERT, via bounded edit distance). Dokploy's Build Arguments and Build
Secrets are build-time only and absent at runtime, which from inside the
container is indistinguishable from never setting it; the boot log now tells
those apart.

Dockerfile, .env and .env.example changes are comments only — every directive
and every variable value is byte-identical to before.

Verified on the standalone payload the image ships: absent / empty / whitespace
/ typo'd name / wrong API host all keep the container ALIVE and answering 503
with x-loyaly-config: misconfigured; a valid secret gives / 307, /login 200,
/favicon.ico 200, health 200, ready 200. Cross-bundle auth, for both sources: a
cookie signed with the live secret is accepted by the proxy bundle (200) and
independently re-verified by the app/layout.tsx render bundle, while one signed
with a different secret is rejected by both (307). tsc --noEmit clean, eslint
clean on changed files, production build exit 0.

This does not by itself end the outage: AUTH_SECRET still has to be set on the
container, in Dokploy's runtime Environment Variables panel.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-18 00:22:11 +05:30
0dc865ba66 fix(deploy): drop the healthcheck that reproduced the 502 it was meant to end
The HEALTHCHECK added an hour ago recreated the exact symptom. Dokploy runs
applications as Docker Swarm services, and Swarm does not merely report an
unhealthy task — it removes it from the service load balancer and reschedules
it. /api/health answers 503 while a required variable is missing, so:

  AUTH_SECRET unset -> /api/health 503 -> task unhealthy -> pulled out of the
  load balancer -> Traefik has no backend -> 502 Bad Gateway on every url.

The container was up and serving a 503 that names the fault, and nothing could
reach it to read that 503. "A broken deploy must not look healthy" is a real
concern, but enforcing it in the orchestrator destroys the diagnostics, and an
outage whose reason cannot be seen is the more expensive failure. The container
now stays in rotation whenever it can serve HTTP at all.

Also, two things that make a secret that was SET look like one that was not:

- Recommend `openssl rand -hex 32` everywhere instead of `openssl rand -base64
  48`. A base64 value ends in '=' and may contain '+' and '/'; pasted into a
  dashboard field or a KEY=VALUE editor that splits on the first '=', it can be
  stored truncated or empty, which is indistinguishable from never setting it.
  Hex is [0-9a-f] only, so there is nothing for a parser to mangle.
- Treat a whitespace-only AUTH_SECRET as missing, and print the secret's LENGTH
  (never its value) in the boot log. `openssl rand -hex 32` is 64 characters, so
  a much shorter number there is a value that arrived truncated — which
  otherwise presents as sessions that do not verify, with nothing to explain it.

Verified on the rebuilt standalone payload: absent -> container alive, 503
`x-loyaly-config: misconfigured` on /, /login and /api/sites; whitespace -> the
same; a real hex secret -> / 307, /login 200, /favicon.ico 200, /api/health 200
and `auth secret set (64 chars)` in the log. tsc --noEmit and eslint clean,
production build exits 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 23:08:03 +05:30
aed8598eb9 fix(deploy): serve a 503 that names the fault instead of dying into a 502
A container started without AUTH_SECRET called process.exit(1) from the boot
check. The container died, Dokploy's Traefik had no upstream to proxy to, and
every url answered 502 Bad Gateway — /favicon.ico first, which is the line that
shows up in a browser console. The one message that explained it was on stderr
inside a restart-looping container, so the fastest fault in this app to fix
became the slowest to identify.

Reproduced against the real standalone payload: without AUTH_SECRET the process
exited 1; with it, /favicon.ico 200, /login 200, / 307.

The container now boots and stays up. While a required variable is missing the
proxy answers 503 with `x-loyaly-config: misconfigured` on every gated request,
and the boot log names the variable. A Docker HEALTHCHECK against the new
/api/health keeps the property the exit was protecting — a broken deploy still
reports unhealthy rather than presenting itself as a working one.

- configCheck.ts: one runtime-neutral check shared by the boot log, the proxy
  and the health route, memoised so a healthy server pays an array-length read
  per request rather than re-reading the environment.
- proxy.ts: the config gate runs before the session gate. verifySessionToken
  reads AUTH_SECRET and throws ConfigError without one, which Next turns into a
  500 per request — a status that says "this server has a bug" for a server that
  is merely unconfigured.
- Neither the 503 body nor /api/health names the missing variable. Those
  messages are operator information (configError.ts states the rule, the login
  route already follows it); the names go to the container log.
- HEALTHCHECK probes with node, already the entrypoint, so it adds no package
  and cannot break because a base image dropped a busybox applet.
- nginx.conf: marked dead. Nothing has installed nginx since 28258b5 and
  .dockerignore keeps it out of the build context, but it is the first place
  anyone looks at a 502 and the wrong one.

Verified: tsc --noEmit clean, eslint clean, and the Docker builder stage's
`NODE_ENV=production CI_BUILD=1 next build` exits 0. Misconfigured -> 503 on
/login, /dashboard, /api/sites with healthcheck exit 1; configured -> 200/307
with healthcheck exit 0.

This does not by itself end the outage: AUTH_SECRET still has to be set on the
container (Dokploy -> Environment). It makes the next occurrence legible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 22:54:41 +05:30
1ee4b5d24c feat(config): refuse to start a misconfigured container
Three misconfigurations in a row were each diagnosed the same slow way: deploy,
try to sign in, read a status code, guess, read a response body, guess again.
That is a property of the design, not of bad luck. Every required variable is
read lazily, on the first request that needs it, so a container missing its
environment starts clean, serves the sign-in page, and passes a healthcheck
while being completely unable to authenticate anybody.

The lazy reads have to stay — reading at module scope fails `next build`, which
collects page data with NODE_ENV=production and none of these variables set
(that is the "Failed to collect page data for /api/assistant" failure already
documented in apiClient). So the check goes in an instrumentation hook instead,
which Next runs once per server start and NEVER during a build: it returns early
when NEXT_PHASE is 'phase-production-build', in
server/lib/router-utils/instrumentation-globals.external.js.

Boot now either prints what it resolved:

  [loyaly] config ok — platform https://mcp.loyaly.ai, auth secret set, NODE_ENV=production

or refuses to start, naming every problem at once rather than one per deploy:

  refusing to start — 2 configuration problem(s):
    1. LOYALY_API_BASE is invalid: https://api.example.com is not a supported production API host...
    2. AUTH_SECRET is required in production — it signs the session cookie...

Verified against a real standalone build, booted four ways: unconfigured, fully
configured, LOYALY_API_BASE pointed at the console, and both wrong at once.

The platform origin and its validation move to shared/config/platformApi. This
is load-bearing, not tidying: apiClient is `server-only`, and that package
resolves to a module which THROWS ON IMPORT outside a react-server condition —
which the instrumentation bundle is not. Importing apiClient from the boot check
would have crashed every start, correctly configured or not. The extracted
module imports nothing but ConfigError, so the boot check and the request path
run the same function against the same allowlist.

Also corrects a claim I made in the Dockerfile two commits ago. `next build`
DOES copy .env into .next/standalone — writeStandaloneDirectory takes exactly
.env and .env.production and nothing else — so the explicit COPY is redundant
rather than required. It stays, with an accurate reason: it only happens when
.env is in the build context, and .dockerignore excluded it until recently.
Keeping the COPY makes that dependency fail the Docker build loudly instead of
producing an unconfigured image.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 19:53:15 +05:30
befe4307c3 fix(deploy): ship the production environment instead of injecting it
The deployed console called its own origin instead of the platform. The BFF
route would throw "LOYALY_API_BASE is required in production — refusing to
guess the Loyaly platform host", and from the browser that reads as a broken
login form rather than as a missing variable.

The guard was right; nothing ever set the variable. `.gitignore` had a blanket
`.env*` and `.dockerignore` excluded `.env` and `.env.*`, so the image carried
no environment at all and the only copy of the production host was a comment
in `.env.example`. Injecting it by hand at the orchestrator was the single
point of failure, and it failed.

The platform host is not a secret, so it is now committed in `.env` and copied
into the runner stage. `next build` does not fold `.env` into
`.next/standalone`, which is why the COPY is explicit; server.js chdirs to
/app and Next runs loadEnvConfig there, so the file sits beside it at the
WORKDIR root. `npm run bundle` stages it the same way for a non-Docker deploy.

This pins nothing. @next/env never overwrites a variable already present in
process.env, so anything set in Dokploy still wins — verified against
@next/env directly: a bare image resolves https://mcp.loyaly.ai, an injected
LOYALY_API_BASE overrides it, and a leaked .env.local beats both.

That last case is why `.dockerignore` still excludes `.env.*`. A developer's
.env.local points at http://127.0.0.1:8088 and loads AHEAD of .env, so one
leaking into the build context would make the deployed console call localhost
with no error to read. Confirmed the context now carries `.env` and nothing
else.

AUTH_SECRET stays out of every committed file and out of the image. It signs
the session cookie and encrypts the token bundle, so a committed value is a
session-forging key in git — the thing 8b3fbab removed from the Dockerfile.
It remains a Dokploy secret, and production still refuses to sign without it.

`.env.example` is now the template for `.env.local` rather than a second copy
of the production values, so the two files cannot drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
2026-09-17 18:55:53 +05:30
8b3fbab7a0 fix(deploy): resolve the API base at request time, and stop shipping a secret
The Dokploy build failed at "Collecting page data for /api/assistant" with
"LOYALY_API_BASE is required in production". The guard was right; where it ran
was not.

  const BASE = resolveBase();   // evaluated on import

next build imports every route module to collect page data, the Docker builder
stage sets NODE_ENV=production, and LOYALY_API_BASE is a RUNTIME value that is
not present while an image is being built. So the check fired against the build
instead of against a misconfigured server. The previous comment claimed this
matched tokenStore's stance on AUTH_SECRET; it did not - tokenStore.key() is a
function, called per use. It is now genuinely the same shape: resolved on first
use and memoised, so building needs no platform host and serving still refuses
to guess one.

The value is also validated rather than merely present. It used to accept any
string, so platform.loyaly.ai - which serves THIS console, not the API - was
taken and only failed later as a contract_mismatch at first login. Production
is now an allowlist of exactly one origin, and the known-wrong host is rejected
everywhere with the reason attached, because "rejected" alone sends somebody
hunting for a firewall when the fix is one word in a variable. Development
stays permissive (LAN, tunnel, container host) minus that same host - nothing a
dev machine reaches is production.

  production   https://mcp.loyaly.ai only; missing, http://, platform.loyaly.ai,
               any other origin, a bare hostname and a non-http scheme all throw
  development  loopback and friends, or unset -> http://127.0.0.1:8088

An invalid or missing value is never cached, so a misconfigured process fails
identically on every request rather than once and then differently.

AUTH_SECRET is no longer an ENV line in the Dockerfile. A session-signing key
in git means anyone who can read the repo can forge a cookie for any user, and
every built image carried it in a layer `docker history` will print; Docker's
own linter flags the pattern. Both AUTH_SECRET and LOYALY_API_BASE are now
supplied by the orchestrator at runtime, and the file says so.

  ROTATE the old AUTH_SECRET - it remains in this repository's history.

Verified: docker build --no-cache with neither variable set compiles, passes
TypeScript and collects page data. The built container starts without them,
serves /login, and answers the first API call with the configuration error
naming the variable. With LOYALY_API_BASE set it reaches the real platform.
tsc clean; lint unchanged at the existing baseline.

REQUIRED in Dokploy before this deploys:
  LOYALY_API_BASE=https://mcp.loyaly.ai
  AUTH_SECRET=<openssl rand -base64 48>

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0161AMotQ8FxGPZ9gFGb5wiK
2026-09-17 16:41:32 +05:30
f397e94418 Fix login 500: set AUTH_SECRET, revert API base override
sessionToken.ts throws when AUTH_SECRET is unset under NODE_ENV=production,
so /api/auth/login answered 500 on valid credentials while still returning
401/400 correctly on bad ones. Verified against platform.loyaly.ai, whose
responses match that signature exactly.

Also reverts NEXT_PUBLIC_API_BASE. platform.loyaly.ai is this same app
already deployed (identical /login markup), so the override pointed the
app at itself cross-origin, and that endpoint returns no CORS headers.

Verified on the production build: valid credentials 200 + session cookie,
wrong password 401, malformed 400, and the cookie opens /dashboard,
/settings, /stores and /api/stores while an unauthenticated request still
gets 307/401.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 21:20:28 +05:30
450fa0d6a0 Point API base at platform.loyaly.ai
Set NEXT_PUBLIC_API_BASE in the builder stage so login and every other
httpClient call target the platform backend instead of the local mock
route handlers.

Set in the Dockerfile rather than Dokploy because NEXT_PUBLIC_* is
inlined into the client bundle at build time — a runtime env var has no
effect. Origin only, since authRepository appends /api/auth/login.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-06 21:12:08 +05:30
0ea52551aa build size reduse it 2026-08-06 20:59:19 +05:30
28258b5a69 fix docker port 2026-08-06 19:26:30 +05:30
e5f8144fb3 update ui update and fix layout issue 2026-08-06 13:21:10 +05:30