Commit Graph

5 Commits

Author SHA1 Message Date
57c2a52c1e Refuse an HTTP write timeout that would cut off a legal agent run
Some checks failed
CI / test (push) Failing after 5m32s
CI / fixture (push) Failing after 59s
Production answered 502 Bad Gateway on a non-streamed agent run. Nothing about
that was a gateway fault: krow-proxy already had proxy_read_timeout 3600s, and
the API pods were healthy with zero restarts throughout.

HTTP_WRITE_TIMEOUT was 30s. Every shipped agent runs at the `balanced` tier,
whose deadline is 60s, and the `deep` tier allows 120s. So the server aborted
the response on any run over half the time the runtime considered legal, the
proxy saw its upstream vanish mid-response, and it reported the only thing it
could. A gateway error for something no gateway did — which is why it looked
like infrastructure for as long as it did.

Delegation did not cause this; it made it routine. A parent that asks two
subagents takes longer than one answering alone, so a latent misconfiguration
became a reliable one. Verified: the exact request that returned 502 now
answers 200 in 18s.

Streaming is what hid it, and that is the part worth keeping in mind. The chat
panel uses SSE, so the product looked healthy while every non-streaming caller
got 502 on a slow question. A bug only reachable by the callers who do not yet
exist is one nobody reports.

So the value is now derived from the thing that constrains it — the default is
DeepestAgentDeadline plus headroom rather than a number typed once — and
validate() refuses anything below that deadline at startup. A slow,
intermittent, misattributed failure becomes a message on the first boot.

DeepestAgentDeadline is duplicated in internal/config rather than imported,
because internal/runtime already imports internal/config and a cycle to share
one number is a bad trade. TestConfigKnowsTheDeepestAgentDeadline asserts the
two agree, so drift is a build failure rather than a discovery. It also checks
that no tier exceeds it, or the name lies.

ORDERING, and it matters for the next deploy: the check refuses the old 30s, so
a pod carrying this image against an unpatched configmap will not boot.
Production's configmap is already 180s. The handover says so too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-31 12:06:26 +05:30
f7df96c973 agent build 2026-08-28 12:21:44 +05:30
b6f8655909 aravind changes 2026-08-25 16:37:05 +05:30
Suriyakumarvijayanayagam
954ba9076f Add CORS credentials, transactional endpoints, and container deployment
CORS
  cors.go never set Access-Control-Allow-Credentials, so the
  cookie-authenticated API was unreadable from any cross-origin frontend:
  the server answered correctly and the browser blocked the page from
  reading it. Set for allowlisted origins on both the preflight and the
  actual response. Three tests added.

  HTTP_COOKIE_SAMESITE (lax|none|strict, default lax) is new. CORS is only
  half of what a cross-origin browser call needs; SameSite is judged on
  registrable domain, so a frontend on an unrelated domain gets perfect CORS
  headers and still no cookie. "none" is the only value that survives that,
  and validate() refuses it without the Secure flag.

  The "*" rejection now explains itself: browsers refuse Allow-Origin "*"
  together with credentials, so it would break every authenticated call
  rather than loosen anything.

Transactional endpoints (api-contract.md 12.1)
  POST /api/v1/job-applications/{id}/hire
  POST /api/v1/job-postings/{id}/assignments

  Replaces two client-side loops that wrote several records with no
  transaction and no rollback. Each is now one endpoint and one transaction,
  built over repo.Repo so org scoping, derived columns, type casts and error
  translation are not re-derived. Authorization reuses the existing policy
  table rather than adding a parallel one: a workflow is exactly as
  privileged as the writes it performs. 13 tests, including both rollback
  paths.

Bug fix in the repository layer
  repo.bindValue handled int64/int/float64/string but not int32, which is
  what pgx returns for a PostgreSQL `int` column. Nothing previously read a
  record and wrote one of its fields elsewhere, so it never surfaced; the
  hire flow does exactly that and failed with "ai_score must be a number".
  Both KindInt and KindFloat now accept the widths pgx actually produces.

Deployment
  infrastructure/Dockerfile.api  multi-stage, cross-compiling (BUILDPLATFORM
    + GOARCH) so linux/amd64 builds from arm64 are compiled rather than
    emulated. Alpine runtime, non-root uid 10001, 22.1 MB. Ships api, seed,
    setpassword and migrate, plus the migrations, so a Kubernetes
    initContainer can apply the schema from the same image and tag as the
    API. HEALTHCHECK keys on status code, not body, so a "degraded" instance
    is not pulled from rotation during a migration window.

  infrastructure/docker-compose.yml  migrations run to completion before the
    API starts. Assumes a managed PostgreSQL; the local-db overlay adds one
    with TLS enabled so APP_ENV=production is met rather than dodged.

  scripts/drop_public_tables.go  the one-off used to clear an unrelated
    schema from krowdb on 2026-08-24, kept for the record. Build-tagged
    ignore and gated on CONFIRM_DROP=yes.

Verified against PostgreSQL: 16/16 new tests pass, and the image was built,
run and exercised end to end (login, CORS preflight, authenticated reads,
transaction rollback).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CmQiGq73Uyfq7J4yR8Vxxw
2026-08-25 11:33:01 +05:30
7d12ebef3d first commit 2026-08-24 13:06:29 +05:30