The platform now runs on Groq by default, through the OpenAI-compatible chat-completions shape. That shape is not one vendor — Gemini, OpenRouter, Together, vLLM and a local Ollama serve it too — so moving again stays configuration rather than code. Two things in the deleted file were not Anthropic's and would have gone with it silently: withRetry / MaxAttempts / retryBackoff were defined in anthropic.go and CALLED BY openai.go. Deleting the file wholesale would have removed the retry policy of the provider that survived, and nothing in openai.go mentions it, so the loss would have been invisible until the next 429. The policy is a property of this platform's runs, not of a vendor's API; it now lives in retry.go where no provider can carry it off. StreamComplete had the same problem and moves to gateway.go, beside the Streamer interface whose comment already referenced it. Three stale-configuration failures are now refused at startup instead of being ignored. Each was verified firing through the real config.Load(): MODEL_PROVIDER=anthropic — named separately from every other wrong value because it used to be correct. Ignoring it gives a stack that believes it is on Claude while every run goes to Groq and is billed there. ANTHROPIC_API_KEY set while MODEL_API_KEY is empty. Ignoring a key an operator did set is the worst version of this: they fail every run on a missing credential they are looking straight at. A leftover claude-* model id, naming the tier that carries it. This is the check the previous commit's error-detail work was diagnosing: such an id is accepted by this process, rejected by the provider, and 400s on EVERY run. "A model is wrong" does not say which of three lines to edit. Defaults ship as a matched pair. defaultBaseURL and the three tier ids are one decision, not four: an id is only meaningful against the service that serves it, and a Groq id on an OpenAI base URL is the same failure from the other side. The tiers also stop being one model — a tier whose cost does not differ is a distinction that buys nothing. Verified end to end against a stub of the wire, driving the real wiring (config.Load in production mode, gateway.New, StreamComplete): streamed deltas, tool-call decoding, the loopback credential exemption, and usage totalling 150 rather than 190 — the cached-prefix subtraction still holds. gofmt clean, go vet clean, 14/14 non-DB packages pass. httpserver still needs a reachable database. NOT verified: the I7 planted-injection eval. Removing this path removed the only model whose refusal behaviour had been measured against it, so the new default is unproven there until `make eval-live` runs with a real key. The Groq model ids should also be confirmed against Groq's current lineup. Flagged in CLAUDE.md §12 and docs/handover.md. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PJvibeSc1JYXjatankqM1g
74 lines
2.8 KiB
Go
74 lines
2.8 KiB
Go
package gateway
|
|
|
|
import (
|
|
"context"
|
|
"errors"
|
|
"time"
|
|
)
|
|
|
|
// MaxAttempts is how many times a transient failure is retried.
|
|
//
|
|
// Three total, not three retries. Past that the problem is not transient and a
|
|
// fourth call is just spending money on the same answer.
|
|
const MaxAttempts = 3
|
|
|
|
// retryBackoff is the pause before each retry.
|
|
//
|
|
// Short, and deliberately so: this sits inside a run that already has a
|
|
// wall-clock deadline, and a backoff long enough to be polite to the API is
|
|
// long enough to spend the caller's whole budget waiting. A run that cannot
|
|
// afford the wait dies on its deadline instead, which is the correct failure.
|
|
var retryBackoff = []time.Duration{400 * time.Millisecond, 1200 * time.Millisecond}
|
|
|
|
// withRetry runs one attempt until it succeeds, fails terminally, or runs out
|
|
// of attempts.
|
|
//
|
|
// THE RETRY IS NOT DEFENSIVE POLISH. Error.Retryable() has existed since this
|
|
// package was written and had ZERO callers — the classification was built and
|
|
// never used, so a 529 "overloaded" killed a run that would have succeeded four
|
|
// hundred milliseconds later. Found by a real overload during live testing,
|
|
// where it presented as "the agent could not finish" with nothing to act on.
|
|
//
|
|
// Only genuinely transient failures qualify: rate limits, timeouts, and 5xx.
|
|
// A 400 is a malformed request and will be malformed again; a 401 is a bad
|
|
// credential and retrying it three times just gets refused three times.
|
|
//
|
|
// The run's context governs. A retry that would outlive the caller's deadline
|
|
// does not happen — the deadline belongs to the run, not to this function, and
|
|
// waiting past it would turn a bounded run into an unbounded one.
|
|
//
|
|
// THIS FILE EXISTS BECAUSE THE POLICY OUTLIVED ITS FIRST PROVIDER. It was
|
|
// written inside the Anthropic implementation and used by both, so deleting
|
|
// that implementation would have deleted the retry policy of the one that
|
|
// remained — silently, because nothing about `openai.go` mentions it. The
|
|
// policy is a property of this platform's runs, not of any vendor's API, so it
|
|
// now lives somewhere no provider can take with it when it goes.
|
|
func withRetry(ctx context.Context, once func() (*Response, error)) (*Response, error) {
|
|
var last error
|
|
for attempt := 0; attempt < MaxAttempts; attempt++ {
|
|
if attempt > 0 {
|
|
pause := retryBackoff[min(attempt-1, len(retryBackoff)-1)]
|
|
select {
|
|
case <-time.After(pause):
|
|
case <-ctx.Done():
|
|
// Out of time. The ORIGINAL failure is returned rather than the
|
|
// context error: "the model was overloaded" is what an operator
|
|
// needs to see, and "context deadline exceeded" would hide it.
|
|
return nil, last
|
|
}
|
|
}
|
|
|
|
resp, err := once()
|
|
if err == nil {
|
|
return resp, nil
|
|
}
|
|
last = err
|
|
|
|
var gwErr *Error
|
|
if !errors.As(err, &gwErr) || !gwErr.Retryable() {
|
|
return resp, err
|
|
}
|
|
}
|
|
return nil, last
|
|
}
|