Documentation v1.1.1

Reasoning Levels and Fast Mode

·

The shared reasoning scale, how each session sets its level, and the fast-mode toggle.

Reasoning levels

go-llm-router normalizes reasoning onto one scale — none, low, medium (default), high, xhigh, max — and maps it per provider (Claude thinking budgets, Gemini thinking budgets, OpenAI effort, ...). Aliases minimal, extra, and ultra map to low, xhigh, and max. Levels outside a model's supported range are clamped rather than rejected.

The level is passed explicitly through each Send call; there is no global reasoning setting in config.json. Each session has its own level — auto (the default) or a fixed one — set with /model → reasoning or cycled with Shift+A / Shift+D in the TUI, auto included. With auto, the level comes from the selector's work kind (the LLM dispatcher or TypeSafe, whichever runs), falling back to medium when it returned none; this also applies to a session pinned to a model. A level passed with the request (for example reasoning_effort) wins over both.

Auto reasoning (auto_reasoning, v1.0.18 to v1.0.27) was a global TypeSafe-only toggle; it was removed in v1.1.0 and replaced by the per-session auto level above. The auto_reasoning field of POST /v1/model and GET /v1/session/:id is gone.

Fast mode

Shift+F toggles fast mode, which passes provider.ModeFast through the router so supported backends request a faster service tier. Support is model-specific (core.SupportFast) — for example recent OpenAI generations, Claude Opus 4.8 / Opus 5, most Grok models, and selected Gemini families. Unsupported models silently fall back to the default tier. Fast mode is process-local and not persisted.

中文