Run the turn again on another provider when a rate limit lands mid-run
In-place failover covered none of the failures this deployment actually had. gateway.canFailOver will not move a conversation that has called a tool — the assistant turn echoing that call belongs to the provider that issued it, and a vendor which signs its function calls rejects a follow-up carrying somebody else's. But a rate limit lands where the request is BIGGEST, which is the second or third model call, once the catalogue, the retrieved block, the tool results and the whole prior conversation are being re-sent. Every GatewayFailure in agent_runs had already called a tool. The error text says exactly what it was: http 429: Rate limit reached for model `openai/gpt-oss-120b` … on tokens per minute (TPM): Limit 8000, Used 7183 So the loop starts the turn over on the next provider. No transcript is sent, so nothing provider-specific travels and the signature problem cannot arise: the question is simply asked again somewhere with budget left. It costs the work already done, charged to the budget that is not exhausted. gateway.Standby is the whole of what the runtime is told — "there is another one, here it is". No vendor, credential or model id crosses the boundary, and the loop still cannot name a provider. THE RULE THAT MAKES IT SAFE: a run carrying a confirmation never restarts. Re-running re-runs its tools; a read twice is two reads, a write twice is two shifts assigned. I4 makes the test cheap — a write executes only against a resolved token (Registry.gate), so a run with no confirmation cannot have written anything, and one with a confirmation is refused without inspecting what it did. Once, not until the providers run out: a question worth asking twice is not worth asking five times, and each attempt spends a real budget. A terminal error — a rejected credential, a model this deployment cannot use — is not retried anywhere, on the same line canFailOver already draws. Five tests, including both refusals. Verified with teeth: disabling the restart fails the rate-limit case. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -39,6 +39,33 @@ func NewFailover(primary Gateway, rest ...Gateway) Gateway {
|
||||
return &failover{providers: append([]Gateway{primary}, rest...)}
|
||||
}
|
||||
|
||||
// Standby is a gateway that has somewhere else to go.
|
||||
//
|
||||
// The runtime needs this and must NOT learn what a provider is. A mid-run
|
||||
// failure cannot be moved by this package — the conversation is half built and
|
||||
// its tool calls belong to whoever issued them (see canFailOver) — so the only
|
||||
// thing that can rescue it is starting the run again somewhere else, and only
|
||||
// the loop can do that. This is the whole of what the loop is told: "there is
|
||||
// another one, here it is", with no vendor, credential or model id crossing the
|
||||
// boundary.
|
||||
type Standby interface {
|
||||
// Standby returns a gateway beginning at the NEXT provider, and whether
|
||||
// there was one. The receiver is unchanged.
|
||||
Standby() (Gateway, bool)
|
||||
}
|
||||
|
||||
// Standby drops the provider that just failed and returns the rest.
|
||||
//
|
||||
// The remainder keeps its own fallbacks, so a second failure on a three
|
||||
// provider deployment still has somewhere to go. With one provider left there
|
||||
// is no wrapper at all, which is NewFailover's own rule.
|
||||
func (f *failover) Standby() (Gateway, bool) {
|
||||
if len(f.providers) < 2 {
|
||||
return nil, false
|
||||
}
|
||||
return NewFailover(f.providers[1], f.providers[2:]...), true
|
||||
}
|
||||
|
||||
func (f *failover) Complete(ctx context.Context, req Request) (*Response, error) {
|
||||
var last error
|
||||
for i, p := range f.providers {
|
||||
|
||||
Reference in New Issue
Block a user