Documentation v1.1.1

Send Failures and Fallback

·

What happens when a model call times out, is rate limited, runs out of quota or context, and how fallback picks the next model.

Send failure handling

Agent.Send() failures escalate instead of retrying blindly:

Failure Behavior
Timeout Up to MaxSendTimeoutRetries (3) attempts on the same model, spaced by SendTimeoutRetryInterval (15 s)
Rate limit (HTTP 429 or a rate-limit message) Register a 30-minute cooldown, retry the same model after 5 s, 10 s, then 15 s; then switch to the next fallback
Quota exhausted (HTTP 402 / 403 or a quota / billing message) Register a 30-minute cooldown and switch to the next fallback immediately
Context-length exceeded Since v1.1.1, compact history first (the same ExtractOldHistories → ToolHistory pass an exceeded token budget triggers) and resend; when that cannot compact, trim the oldest exchange (compact.TrimFallback) and resend; abort only when nothing is left to trim
No response while streaming Health-probe every 30 s (UnresponsiveProbeInterval), retry a failed probe every 10 s, switch model after 3 failures
Any other error Switch to the next fallback model; abort when no healthy fallback remains

Fallback candidates come from ResolveAgent's ordered list, which since v1.1.0 excludes pass-tier models. nextAgent skips models from the failed model's provider and models whose context window is smaller than the current input, health-checks each candidate (HealthCheckTimeout, 10 s), and rebuilds the list for up to 3 rounds (maxFallbackRounds). A session bound to a specific model (anything other than auto) never falls back.

Switching models clears ToolHistories (or applies compact.RawToolFallback), resets the dedupe map, and restarts the counters — a fresh model never inherits the failed model's partial state.

Timeout layers

Layer Value Catches
Provider HTTP client set inside go-llm-router per provider Transport-level stalls
AgentSendTimeoutSec limits.agent_send_timeout_seconds in config.json, default 600 Exec-layer ceiling via context.WithTimeout
Unresponsive watchdog probe every 30 s; after a failed probe, retry every 10 s; 3 failures switch model A stream that hangs without erroring
Health check 10 s Liveness probe of each fallback candidate
中文