Commit Graph

21 Commits

Author SHA1 Message Date
Suriya
cde7d4b84b Restore dashboard --enable-skip-login
Skip-login was stripped in an earlier "harden security" pass, which is
why the dashboard started demanding a token. Re-added it; it now runs
as the dashboard's own view-only ServiceAccount (get/list/watch), so
opening it needs no token but write access still requires the
admin-user token as before.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 12:37:59 +05:30
Suriya
dd5dfe10f7 Restore dashboard skip-login; remove dead Flux config
Skip-login was stripped in the earlier "harden security" pass, which
is why the dashboard started demanding a token. Re-added
--enable-skip-login; it now runs as the dashboard's own view-only
ServiceAccount (get/list/watch), so opening it needs no token but
write access still requires the admin-user token.

Also removes clusters/production/ (flux-system bootstrap, the
apps-alaska/core/nearle Kustomizations, gitea webhook receiver) since
Flux was removed on the server side - deploys are manual kubectl
apply / deploy-*.sh from here on, and this config had no controller
left to read it.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 12:37:31 +05:30
Suriya
1b629d8ea3 Recreate worker-notifications - lost when core namespace got wiped
This StatefulSet was deployed manually outside git at some point before
this repo's GitOps work began, and was destroyed when deleting the core
Kustomization triggered a full namespace recreation (Flux prune-on-delete
cascades regardless of kubectl's --cascade flag - that only affects
Kubernetes' own owner-reference GC, not Flux's finalizer).

Reconstructed from its own logs (NATS_STREAM=NOTIFICATIONS,
NATS_CONSUMER=notifications-worker, FILTER_SUBJECT=api.v1.notifications.push)
plus the same pattern as its sibling workers. Resource limits and
WORKER_CONCURRENCY are a best-guess match to worker-rider-logs, since the
original values were never version-controlled anywhere. Adding it to git
now so this can't happen again.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 12:25:55 +05:30
Suriya
430bd79bed Revert --token-ttl=0 on dashboard - broke login entirely
Login stopped working the moment this shipped. Suspect this dashboard
version treats token-ttl=0 as "expire immediately" rather than "never
expire". Reverting to the default (900s) to restore working login;
the idle-timeout annoyance can be revisited with a large finite value
instead of 0, tested before shipping.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 11:31:08 +05:30
Suriya
ff372b8b4c Disable dashboard session timeout so login persists
Default --token-ttl is 900s (15min idle), which was forcing re-entry of
the token repeatedly. Setting it to 0 makes a logged-in session
persist instead of expiring back to the login screen.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 11:27:26 +05:30
Suriya
1b38a0c240 Exclude traefik-middlewares.yaml - Middleware CRD not installed on cluster
kubectl api-resources shows no Middleware kind under any API group, so
this manifest could never apply and was blocking the entire core
Kustomization. The Ingress annotation referencing it has been a
pre-existing no-op; CORS is actually handled by nginx-queue-proxy.conf
and each app's own CORS middleware.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 11:00:04 +05:30
Suriya
f883540272 Exclude nearle-ariane.yaml from Flux sync - crash-loops on first deploy
ariane was never actually running on the cluster before this rollout
(not in any prior pod listing); Flux deploying it for the first time
exposed it's broken, unrelated to the GitOps migration itself. Pulling
it out of scope so it stops blocking the nearle Kustomization's health
check for the services that do work (fiesta, jupiter, atlantis, titan).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:58:52 +05:30
Suriya
ffef003a2b Bring nearle stack and routing manifests under Flux management
- manifests/core/kustomization.yaml: add ingress-unified.yaml and
  traefik-middlewares.yaml, which were never in any kustomization and so
  were never actually GitOps-managed - the queue.workolik.com routing fix
  from c38a367 turned out to be live already (likely applied manually
  before this session), but was completely undetected by Flux until now.
  Dropped the redundant top-level `namespace: core` override since every
  existing resource already sets its own namespace explicitly, and the
  two new files span alaska/nearle.
- manifests/nearle/kustomization.yaml + clusters/production/apps-nearle.yaml:
  same GitOps treatment already applied to alaska/core, so a fiesta/jupiter/
  atlantis/titan/ariane version bump in git now auto-deploys instead of
  requiring manual kubectl apply.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:55:47 +05:30
Suriya
a86baac8a8 Trivial commit to test the Gitea -> Flux webhook
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:50:25 +05:30
Suriya
c526d494b0 Fix Receiver type: gitea is not a valid Flux receiver type
The installed Flux version's Receiver CRD rejects type: gitea, which
failed dry-run validation and blocked the entire flux-system
Kustomization batch from applying - including the alaska/core stacks.
Switched to type: generic (no payload parsing needed for a trigger-only
webhook).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:33:10 +05:30
Flux
e7340fb1ce Add Flux sync manifests 2026-07-19 22:58:13 -06:00
Flux
26557b0f9c Add Flux sync manifests 2026-07-19 22:51:38 -06:00
Flux
0e6aa1a4df Add Flux v2.9.2 component manifests 2026-07-19 22:51:27 -06:00
Suriya
dbf2941bbf Remove hardcoded replicas from HPA-managed deliveries StatefulSet
replicas: 4 was fighting the HPA's minReplicas: 8 on every reconcile -
GitOps tooling (Flux) reapplies the manifest on an interval, which would
keep yanking capacity back down between HPA corrections and quietly
undo the burst-headroom fix from c38a367.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:18:06 +05:30
Suriya
eb99e351dc Swap webhook Ingress for a NodePort service
No new A record could be added for a dedicated webhook host, so expose
Flux's webhook-receiver directly via NodePort instead of going through
Traefik/Ingress/DNS - same approach the deliveries LoadBalancer already
uses on 30662.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:14:11 +05:30
Suriya
f58f339b43 Add Flux CD GitOps sync for alaska/core stacks
Prepares clusters/production/ for `flux bootstrap git` - Kustomizations
for manifests/alaska and manifests/core, plus a Gitea push Receiver so
new commits reconcile immediately instead of waiting on the poll
interval. Webhook exposed on a dedicated host (flux-webhook.workolik.com)
to avoid the existing queue.workolik.com Gateway/Ingress overlap.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-07-20 10:10:58 +05:30
Suriya
c38a36709b Fix queue.workolik.com routing conflict and give deliveries HPA burst headroom
queue.workolik.com was served by three separate routing definitions that
didn't agree: nginx-queue-proxy.conf and the classic queue-ingress both sent
everything to the deliveries app, but the Gateway API HTTPRoute
(deliveries-route) had a carve-out sending /live/api/v1/mob/orders and
/live/api/v1/web/products to fiesta's raw backend in the nearle namespace
instead (3 fixed replicas, no autoscaling, no resource limits) - a
completely different capacity profile from deliveries (HPA'd, 4-20
replicas). Depending on which router won for a given request, orders could
land on two backends with very different ability to absorb a burst,
plausibly explaining partial order loss / 429s under concurrent load.
Removed the carve-out so all three routing paths agree: everything goes to
deliveries-service.

Also raised deliveries-hpa minReplicas 4->8 and added an explicit
aggressive scaleUp behavior (no stabilization delay, up to 4 pods or 100%
every 15s). Autoscaling reacts to sustained load over roughly 30-60s
(metric polling + pod scheduling + readiness delay), so it does very
little for a burst that's over in seconds - minReplicas is the actual
defense; the behavior block just makes any further scaling land as fast as
possible.
2026-07-18 16:17:18 +05:30
Suriya
0a8c3b0374 Fix deployment tooling: shell scripts and Terraform
Terraform (validated with the real terraform CLI - was never actually
run against this cluster, no state file existed):
- Delete main.tf: it declared a duplicate kubernetes_namespace.core
  (also in namespaces.tf) and a duplicate provider "kubernetes" block
  (also in providers.tf), both hard errors that would fail
  `terraform plan` immediately.
- Fix workloads.tf references to 6 files deleted in the manifest
  cleanup (jupiter-sts/svc, atlantis-sts/svc, fiesta-sts/svc) - now
  points at the canonical nearle-jupiter/atlantis/fiesta.yaml.
- Fix every kubernetes_manifest resource: they fed multi-document
  YAML (multiple '---'-separated docs per file) straight into
  yamldecode(), which only parses a single document. Rewrote using a
  split-on-'---' + for_each pattern, confirmed safe first by checking
  separator counts exactly match document counts for every affected
  file (no embedded '---' inside any script/config content).
- Add the doormile namespace; rename kubernetes_namespace to
  kubernetes_namespace_v1 (fixes a deprecation warning).
- `terraform validate` now passes clean.

Shell scripts:
- deploy-nearle-stack.sh only applied 4 of the ~13 files in
  manifests/nearle/ - missing the ConfigMap/Secrets fiesta/jupiter/
  titan/ariane need via envFrom, the fiesta gateway script ConfigMap,
  atlantis entirely, and the Gateway/ReferenceGrant/jupiter-cors-proxy
  resources. Now applies every file (verified by diffing the
  directory listing against the script).
- Added deploy-doormile.sh and deploy-ingress.sh - nothing previously
  applied ingress-unified.yaml or traefik-middlewares.yaml at all.
- Rewrote deploy.sh as an orchestrator calling all of the above in
  order (previously referenced a manifests/namespace.yaml layout that
  hasn't existed since before this repo's initial commit).
- Rewrote check-k8s-status.sh to check the real namespaces
  (core/nearle/alaska/doormile/kubernetes-dashboard) instead of a
  'nats-backend' namespace that never existed in this repo.
- Fixed a `cd` bug in setup-jetstream.sh that made it change into
  shfiles/ and then look for scripts/setup_jetstream.py there (a
  child directory that doesn't exist) - it could never have found its
  own target file. Now pulls NATS credentials from the live
  nats-credentials Secret instead of a third hardcoded copy.

Python scripts:
- sync_manifests.py had hardcoded Windows paths (e:\nats\kubernetes\...)
  - replaced with paths relative to the script's own location so it
  actually runs here (or anywhere). Verified by running it.
- setup_jetstream.py created durable consumers under different names
  than worker.py computes at runtime ({NATS_CONSUMER}_{subject}), so
  its max_deliver/ack_wait settings never actually reached the
  consumers workers bind to. Naming now derived with the same logic
  worker.py uses - verified all 10 derived names match workers.yaml
  exactly.
- purge-old-messages.py had hardcoded NATS credentials with no env
  var override at all - fixed to match the pattern used everywhere
  else.
2026-07-18 16:08:15 +05:30
Suriya
91dd240431 Fix worker/gateway logic bugs and duplicate CORS headers
worker.py (both the ConfigMap copy and conf/worker.py):
- Set an explicit ack_wait=60s on the JetStream pull consumer. It was
  previously left at the implicit default (~30s), the same ballpark
  as the outbound HTTP timeout - a slow-but-legitimate external call
  could cause JetStream to redeliver the message to another worker
  while the first was still mid-request, double-processing a
  non-idempotent call (e.g. duplicate order creation).
- Track in-flight tasks and drain them (bounded wait) before closing
  the NATS/HTTP connections on shutdown, instead of cutting them off
  immediately - avoids dropped/duplicated messages on pod restarts.
- Generic exception handler now does nak(delay=5) instead of an
  undelayed nak(), avoiding a tight redelivery loop on a persistent
  bug.
- Missing 'data' field in a message now explicitly drops with a log
  line instead of silently forwarding the entire internal envelope.
- Removed the hardcoded NATS password fallback baked into the source
  (every deployment already supplies it via a Secret at runtime, so
  this was a redundant plaintext copy sitting in a ConfigMap).

app.py (both the ConfigMap copy and conf/app.py):
- Fixed "NATS by connected" typo -> "NATS not connected".
- Same hardcoded-password-fallback removal as worker.py.

CORS:
- conf/nginx-jupiter.conf and the in-cluster jupiter-cors-proxy nginx
  config both add their own CORS headers without stripping any the
  upstream might set, unlike nginx-queue-proxy.conf which does this
  correctly. Added proxy_hide_header for the ACA-* headers in both -
  browsers reject a response with duplicate Access-Control-* values.

docker-compose.yml:
- Added the missing doormile-proxy service (doormile.com -> :8206 ->
  NodePort 30830). nginx-doormile.conf existed but had no service
  wiring it into Traefik, unlike every other app.
2026-07-18 16:07:49 +05:30
Suriya
836c079a05 Fix Kubernetes manifest bugs, dedupe drifted files, harden security
- Rebuild manifests/doormile/miletruth.yaml (was corrupted since the
  initial commit - contained pasted AI/terminal output, truncated env
  var names/values, duplicate keys). Rebuilt from the confirmed-live
  config, secrets sourced via a Secret instead of plaintext values.
- Lock down the Kubernetes Dashboard: remove --enable-skip-login /
  --enable-insecure-login / --insecure-port=9090, remove the extra
  cluster-admin binding on the dashboard's own ServiceAccount, remove
  the now-dead insecure NodePort Service. Token-based login via the
  existing admin-user ServiceAccount is unaffected.
- Fix the duplicate `backendRefs` key under the same HTTPRoute rule in
  alaska.yaml (invalid/redundant YAML).
- Delete 6 redundant duplicate manifests (fiesta-sts/svc,
  atlantis-sts/svc, jupiter-sts/svc) that were partial, stale subsets
  of nearle-fiesta/atlantis/jupiter.yaml - one pair disagreed on the
  fiesta image tag entirely (v1.3.50 vs v1.3.67, neither of which
  matched what's actually live).
- Reconcile nearle-fiesta.yaml and nearle-jupiter.yaml image tags to
  the confirmed-live versions (v1.3.78 / v2.7.55).
- Add allowPrivilegeEscalation:false + drop-all-capabilities to
  fiesta/atlantis/jupiter/titan/ariane and the 5 specialized core
  workers, which previously ran with no securityContext at all.
- Add terminationGracePeriodSeconds:45 to the worker StatefulSets so
  Kubernetes gives the new graceful-shutdown drain (see worker.py
  changes) enough time before SIGKILL.
2026-07-18 16:07:32 +05:30
caac8413e9 Initial commit 2026-07-18 12:00:33 +05:30