Files
krow_backend/go-api/internal/runtime/deadline_test.go
Suriyakumarvijayanayagam 57c2a52c1e
Some checks failed
CI / test (push) Failing after 5m32s
CI / fixture (push) Failing after 59s
Refuse an HTTP write timeout that would cut off a legal agent run
Production answered 502 Bad Gateway on a non-streamed agent run. Nothing about
that was a gateway fault: krow-proxy already had proxy_read_timeout 3600s, and
the API pods were healthy with zero restarts throughout.

HTTP_WRITE_TIMEOUT was 30s. Every shipped agent runs at the `balanced` tier,
whose deadline is 60s, and the `deep` tier allows 120s. So the server aborted
the response on any run over half the time the runtime considered legal, the
proxy saw its upstream vanish mid-response, and it reported the only thing it
could. A gateway error for something no gateway did — which is why it looked
like infrastructure for as long as it did.

Delegation did not cause this; it made it routine. A parent that asks two
subagents takes longer than one answering alone, so a latent misconfiguration
became a reliable one. Verified: the exact request that returned 502 now
answers 200 in 18s.

Streaming is what hid it, and that is the part worth keeping in mind. The chat
panel uses SSE, so the product looked healthy while every non-streaming caller
got 502 on a slow question. A bug only reachable by the callers who do not yet
exist is one nobody reports.

So the value is now derived from the thing that constrains it — the default is
DeepestAgentDeadline plus headroom rather than a number typed once — and
validate() refuses anything below that deadline at startup. A slow,
intermittent, misattributed failure becomes a message on the first boot.

DeepestAgentDeadline is duplicated in internal/config rather than imported,
because internal/runtime already imports internal/config and a cycle to share
one number is a bad trade. TestConfigKnowsTheDeepestAgentDeadline asserts the
two agree, so drift is a build failure rather than a discovery. It also checks
that no tier exceeds it, or the name lies.

ORDERING, and it matters for the next deploy: the check refuses the old 30s, so
a pod carrying this image against an unpatched configmap will not boot.
Production's configmap is already 180s. The handover says so too.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
2026-08-31 12:06:26 +05:30

37 lines
1.4 KiB
Go

package runtime
import (
"testing"
"github.com/krow/krow-backend/go-api/internal/config"
)
// The HTTP server must not cut off a run the runtime considers legal.
//
// config.DeepestAgentDeadline duplicates the deep tier's deadline, because
// internal/runtime already imports internal/config and a cycle to share one
// number is a bad trade. This is the thing that makes the duplicate safe: the
// two drifting apart is a failing test rather than a 502 in production
// months later.
//
// It is not hypothetical. Production ran HTTP_WRITE_TIMEOUT=30s against a
// balanced deadline of 60s, so the server aborted any run over half its
// allowed time and the proxy in front reported 502 — a gateway error for
// something no gateway did.
func TestConfigKnowsTheDeepestAgentDeadline(t *testing.T) {
deepest := LimitsForTier("deep").Deadline
if config.DeepestAgentDeadline != deepest {
t.Fatalf("config.DeepestAgentDeadline is %s but LimitsForTier(\"deep\") is %s — "+
"raise the constant, or the config validation will accept a write timeout "+
"that cuts off a legal run", config.DeepestAgentDeadline, deepest)
}
// And it must genuinely be the largest, or the name lies.
for _, tier := range []string{"fast", "balanced", "deep", "nonsense"} {
if d := LimitsForTier(tier).Deadline; d > config.DeepestAgentDeadline {
t.Errorf("tier %q allows %s, which exceeds DeepestAgentDeadline %s",
tier, d, config.DeepestAgentDeadline)
}
}
}