# Send Failures and Fallback

What happens when a model call times out, is rate limited, runs out of quota or context, and how fallback picks the next model.

## Send failure handling

`Agent.Send()` failures escalate instead of retrying blindly:

| Failure | Behavior |
|---|---|
| Timeout | Up to `MaxSendTimeoutRetries` (3) attempts on the same model, spaced by `SendTimeoutRetryInterval` (15 s) |
| Rate limit (HTTP 429 or a rate-limit message) | Register a 30-minute cooldown, retry the same model after 5 s, 10 s, then 15 s; then switch to the next fallback |
| Quota exhausted (HTTP 402 / 403 or a quota / billing message) | Register a 30-minute cooldown and switch to the next fallback immediately |
| Context-length exceeded | Since v1.1.1, compact history first (the same `ExtractOldHistories` → `ToolHistory` pass an exceeded token budget triggers) and resend; when that cannot compact, trim the oldest exchange (`compact.TrimFallback`) and resend; abort only when nothing is left to trim |
| No response while streaming | Health-probe every 30 s (`UnresponsiveProbeInterval`), retry a failed probe every 10 s, switch model after 3 failures |
| Any other error | Switch to the next fallback model; abort when no healthy fallback remains |

Fallback candidates come from `ResolveAgent`'s ordered list, which since v1.1.0 excludes `pass`-tier models. `nextAgent` skips models from the failed model's provider and models whose context window is smaller than the current input, health-checks each candidate (`HealthCheckTimeout`, 10 s), and rebuilds the list for up to 3 rounds (`maxFallbackRounds`). A session bound to a specific model (anything other than `auto`) never falls back.

Switching models clears `ToolHistories` (or applies `compact.RawToolFallback`), resets the dedupe map, and restarts the counters — a fresh model never inherits the failed model's partial state.

## Timeout layers

| Layer | Value | Catches |
|---|---|---|
| Provider HTTP client | set inside `go-llm-router` per provider | Transport-level stalls |
| `AgentSendTimeoutSec` | `limits.agent_send_timeout_seconds` in `config.json`, default `600` | Exec-layer ceiling via `context.WithTimeout` |
| Unresponsive watchdog | probe every 30 s; after a failed probe, retry every 10 s; 3 failures switch model | A stream that hangs without erroring |
| Health check | 10 s | Liveness probe of each fallback candidate |
