Send Failures and Fallback
What happens when a model call times out, is rate limited, runs out of quota or context, and how fallback picks the next model.
Send failure handling
Agent.Send() failures escalate instead of retrying blindly:
| Failure | Behavior |
|---|---|
| Timeout | Up to MaxSendTimeoutRetries (3) attempts on the same model, spaced by SendTimeoutRetryInterval (15 s) |
| Rate limit (HTTP 429 or a rate-limit message) | Register a 30-minute cooldown, retry the same model after 5 s, 10 s, then 15 s; then switch to the next fallback |
| Quota exhausted (HTTP 402 / 403 or a quota / billing message) | Register a 30-minute cooldown and switch to the next fallback immediately |
| Context-length exceeded | Since v1.1.1, compact history first (the same ExtractOldHistories → ToolHistory pass an exceeded token budget triggers) and resend; when that cannot compact, trim the oldest exchange (compact.TrimFallback) and resend; abort only when nothing is left to trim |
| No response while streaming | Health-probe every 30 s (UnresponsiveProbeInterval), retry a failed probe every 10 s, switch model after 3 failures |
| Any other error | Switch to the next fallback model; abort when no healthy fallback remains |
Fallback candidates come from ResolveAgent's ordered list, which since v1.1.0 excludes pass-tier models. nextAgent skips models from the failed model's provider and models whose context window is smaller than the current input, health-checks each candidate (HealthCheckTimeout, 10 s), and rebuilds the list for up to 3 rounds (maxFallbackRounds). A session bound to a specific model (anything other than auto) never falls back.
Switching models clears ToolHistories (or applies compact.RawToolFallback), resets the dedupe map, and restarts the counters — a fresh model never inherits the failed model's partial state.
Timeout layers
| Layer | Value | Catches |
|---|---|---|
| Provider HTTP client | set inside go-llm-router per provider |
Transport-level stalls |
AgentSendTimeoutSec |
limits.agent_send_timeout_seconds in config.json, default 600 |
Exec-layer ceiling via context.WithTimeout |
| Unresponsive watchdog | probe every 30 s; after a failed probe, retry every 10 s; 3 failures switch model | A stream that hangs without erroring |
| Health check | 10 s | Liveness probe of each fallback candidate |