replicas: 4 was fighting the HPA's minReplicas: 8 on every reconcile -
GitOps tooling (Flux) reapplies the manifest on an interval, which would
keep yanking capacity back down between HPA corrections and quietly
undo the burst-headroom fix from c38a367.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
No new A record could be added for a dedicated webhook host, so expose
Flux's webhook-receiver directly via NodePort instead of going through
Traefik/Ingress/DNS - same approach the deliveries LoadBalancer already
uses on 30662.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Prepares clusters/production/ for `flux bootstrap git` - Kustomizations
for manifests/alaska and manifests/core, plus a Gitea push Receiver so
new commits reconcile immediately instead of waiting on the poll
interval. Webhook exposed on a dedicated host (flux-webhook.workolik.com)
to avoid the existing queue.workolik.com Gateway/Ingress overlap.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
queue.workolik.com was served by three separate routing definitions that
didn't agree: nginx-queue-proxy.conf and the classic queue-ingress both sent
everything to the deliveries app, but the Gateway API HTTPRoute
(deliveries-route) had a carve-out sending /live/api/v1/mob/orders and
/live/api/v1/web/products to fiesta's raw backend in the nearle namespace
instead (3 fixed replicas, no autoscaling, no resource limits) - a
completely different capacity profile from deliveries (HPA'd, 4-20
replicas). Depending on which router won for a given request, orders could
land on two backends with very different ability to absorb a burst,
plausibly explaining partial order loss / 429s under concurrent load.
Removed the carve-out so all three routing paths agree: everything goes to
deliveries-service.
Also raised deliveries-hpa minReplicas 4->8 and added an explicit
aggressive scaleUp behavior (no stabilization delay, up to 4 pods or 100%
every 15s). Autoscaling reacts to sustained load over roughly 30-60s
(metric polling + pod scheduling + readiness delay), so it does very
little for a burst that's over in seconds - minReplicas is the actual
defense; the behavior block just makes any further scaling land as fast as
possible.
Terraform (validated with the real terraform CLI - was never actually
run against this cluster, no state file existed):
- Delete main.tf: it declared a duplicate kubernetes_namespace.core
(also in namespaces.tf) and a duplicate provider "kubernetes" block
(also in providers.tf), both hard errors that would fail
`terraform plan` immediately.
- Fix workloads.tf references to 6 files deleted in the manifest
cleanup (jupiter-sts/svc, atlantis-sts/svc, fiesta-sts/svc) - now
points at the canonical nearle-jupiter/atlantis/fiesta.yaml.
- Fix every kubernetes_manifest resource: they fed multi-document
YAML (multiple '---'-separated docs per file) straight into
yamldecode(), which only parses a single document. Rewrote using a
split-on-'---' + for_each pattern, confirmed safe first by checking
separator counts exactly match document counts for every affected
file (no embedded '---' inside any script/config content).
- Add the doormile namespace; rename kubernetes_namespace to
kubernetes_namespace_v1 (fixes a deprecation warning).
- `terraform validate` now passes clean.
Shell scripts:
- deploy-nearle-stack.sh only applied 4 of the ~13 files in
manifests/nearle/ - missing the ConfigMap/Secrets fiesta/jupiter/
titan/ariane need via envFrom, the fiesta gateway script ConfigMap,
atlantis entirely, and the Gateway/ReferenceGrant/jupiter-cors-proxy
resources. Now applies every file (verified by diffing the
directory listing against the script).
- Added deploy-doormile.sh and deploy-ingress.sh - nothing previously
applied ingress-unified.yaml or traefik-middlewares.yaml at all.
- Rewrote deploy.sh as an orchestrator calling all of the above in
order (previously referenced a manifests/namespace.yaml layout that
hasn't existed since before this repo's initial commit).
- Rewrote check-k8s-status.sh to check the real namespaces
(core/nearle/alaska/doormile/kubernetes-dashboard) instead of a
'nats-backend' namespace that never existed in this repo.
- Fixed a `cd` bug in setup-jetstream.sh that made it change into
shfiles/ and then look for scripts/setup_jetstream.py there (a
child directory that doesn't exist) - it could never have found its
own target file. Now pulls NATS credentials from the live
nats-credentials Secret instead of a third hardcoded copy.
Python scripts:
- sync_manifests.py had hardcoded Windows paths (e:\nats\kubernetes\...)
- replaced with paths relative to the script's own location so it
actually runs here (or anywhere). Verified by running it.
- setup_jetstream.py created durable consumers under different names
than worker.py computes at runtime ({NATS_CONSUMER}_{subject}), so
its max_deliver/ack_wait settings never actually reached the
consumers workers bind to. Naming now derived with the same logic
worker.py uses - verified all 10 derived names match workers.yaml
exactly.
- purge-old-messages.py had hardcoded NATS credentials with no env
var override at all - fixed to match the pattern used everywhere
else.
worker.py (both the ConfigMap copy and conf/worker.py):
- Set an explicit ack_wait=60s on the JetStream pull consumer. It was
previously left at the implicit default (~30s), the same ballpark
as the outbound HTTP timeout - a slow-but-legitimate external call
could cause JetStream to redeliver the message to another worker
while the first was still mid-request, double-processing a
non-idempotent call (e.g. duplicate order creation).
- Track in-flight tasks and drain them (bounded wait) before closing
the NATS/HTTP connections on shutdown, instead of cutting them off
immediately - avoids dropped/duplicated messages on pod restarts.
- Generic exception handler now does nak(delay=5) instead of an
undelayed nak(), avoiding a tight redelivery loop on a persistent
bug.
- Missing 'data' field in a message now explicitly drops with a log
line instead of silently forwarding the entire internal envelope.
- Removed the hardcoded NATS password fallback baked into the source
(every deployment already supplies it via a Secret at runtime, so
this was a redundant plaintext copy sitting in a ConfigMap).
app.py (both the ConfigMap copy and conf/app.py):
- Fixed "NATS by connected" typo -> "NATS not connected".
- Same hardcoded-password-fallback removal as worker.py.
CORS:
- conf/nginx-jupiter.conf and the in-cluster jupiter-cors-proxy nginx
config both add their own CORS headers without stripping any the
upstream might set, unlike nginx-queue-proxy.conf which does this
correctly. Added proxy_hide_header for the ACA-* headers in both -
browsers reject a response with duplicate Access-Control-* values.
docker-compose.yml:
- Added the missing doormile-proxy service (doormile.com -> :8206 ->
NodePort 30830). nginx-doormile.conf existed but had no service
wiring it into Traefik, unlike every other app.
- Rebuild manifests/doormile/miletruth.yaml (was corrupted since the
initial commit - contained pasted AI/terminal output, truncated env
var names/values, duplicate keys). Rebuilt from the confirmed-live
config, secrets sourced via a Secret instead of plaintext values.
- Lock down the Kubernetes Dashboard: remove --enable-skip-login /
--enable-insecure-login / --insecure-port=9090, remove the extra
cluster-admin binding on the dashboard's own ServiceAccount, remove
the now-dead insecure NodePort Service. Token-based login via the
existing admin-user ServiceAccount is unaffected.
- Fix the duplicate `backendRefs` key under the same HTTPRoute rule in
alaska.yaml (invalid/redundant YAML).
- Delete 6 redundant duplicate manifests (fiesta-sts/svc,
atlantis-sts/svc, jupiter-sts/svc) that were partial, stale subsets
of nearle-fiesta/atlantis/jupiter.yaml - one pair disagreed on the
fiesta image tag entirely (v1.3.50 vs v1.3.67, neither of which
matched what's actually live).
- Reconcile nearle-fiesta.yaml and nearle-jupiter.yaml image tags to
the confirmed-live versions (v1.3.78 / v2.7.55).
- Add allowPrivilegeEscalation:false + drop-all-capabilities to
fiesta/atlantis/jupiter/titan/ariane and the 5 specialized core
workers, which previously ran with no securityContext at all.
- Add terminationGracePeriodSeconds:45 to the worker StatefulSets so
Kubernetes gives the new graceful-shutdown drain (see worker.py
changes) enough time before SIGKILL.