Model Routing
How fallback priority, tiers and the dispatcher model decide which model answers a request.
Model priority and tiers
Registered models are stored as an ordered list under models in config.json. That order is the fallback priority: after the selected model fails, every other entry is tried top to bottom — pass models excluded — and the last non-pass entry is the final line of defense. Fallback skips models on the failed model's provider and models whose context window cannot hold the input, retries the list up to three rounds, and is off for a session pinned to a model. Before v1.0.20 pass models were always moved to the end; from v1.0.20 to v1.0.27 fallback still tried them at their place in the order. POST /v1/model/priority moves the listed names to the front in the given order and keeps the rest after them.
Each model can carry a tier in model_tag ({"<provider>@<model>": "<tier>"}):
| Tier | Meaning |
|---|---|
S |
Strongest — code and work that asks for depth or precision |
A |
Default for most work, one step below the flagship |
B |
Mainstream mid tier |
C |
Fast and cheap; calls tools reliably as instructed |
pass |
Never picked by auto routing, subagents or fallback, even when the request names it (since v1.1.0); used only when a session is set to it |
An untiered model follows the built-in naming rules: S = claude-fable, claude-opus, gpt-*-astra; A = gpt-*-sol, grok-4.5+, claude-sonnet, gpt-*-terra, gemini-*-pro, deepseek-pro, glm, kimi; B = claude-haiku, gpt-*-luna, gemini-*-flash, grok below 4.5, deepseek; C = *-mini, *-nano, gemini-*-flash-lite and open-weight models (gemma*, gpt-oss, qwen*, llama*). v1.0.19 moved gpt-*-sol and grok-4.5+ from S to A. Tiers are read per request, so a change applies without a restart.
Dispatcher model
The dispatcher LLM decides which worker model handles each task. It runs in exec.Start through ResolveAgent → SelectAgentNames, before Execute() enters its iteration loop, receiving the registered model list, every registered model grouped by its resolved tier (user-set tier first, then the naming rules), the work-kind table below, the user input, and a hint about any matched skill. Its routing call is issued at ReasoningNone with a 30-second timeout.
Routing is skipped when it cannot matter: a model named explicitly by the caller is used as-is (and fails if unregistered), a session bound to a model other than auto uses that model, and a registry with only one model returns it directly.
The dispatcher answers with the work kind on its first line and a comma-separated list of model names on the second (since v1.1.0), so the work kind also sets the reasoning level for sessions whose reasoning is auto — no TypeSafe call needed. The first name that is registered, not pass and not cooling down becomes the primary model. Since v1.0.20 the rest of that list no longer decides fallback — fallback follows the priority order above. It classifies the request into one of the five work kinds in the table under TypeSafe below and walks that kind's tier order. Since v1.0.21 the LLM dispatcher, the TypeSafe classifier and the subagent planner read one shared work-kind table (internal/session/config/tier.go) and one naming rule, so they rank the same way; before that the LLM dispatcher had no research / work split and sent everything outside code, chat and fetch to S > A > B > C. The dispatcher is told never to return a pass model, and a pass name it returns anyway is dropped (before v1.1.0 a request could still name one). Inside a tier, models on the subscription providers claude-code and then codex rank ahead of every other provider. When the same base model is registered under several providers, the list prefers claude-code, then codex / grok-oauth, then copilot, then the direct API, then openrouter. The subagent planner reads the same table one tier lower (S→A, A→B, B→C), so a leg of code work tries A > B > C > S. If the dispatcher call fails, the next dispatcher candidate is chosen by a fixed provider ranking; if none answers, the priority order alone decides.
Set the dispatcher with /model → dispatch in the TUI, or POST /v1/model with a dispatcher field.