Hmbown/CodeWhale · error

maxOutputTokens requires an exact model; Auto routing is…

Error message

maxOutputTokens requires an exact model; Auto routing is unsupported

What it means

The turn handler rejects max_output_tokens together with the "auto" model. Auto routing may pick different providers/protocols with different output-token budget support, so an explicit token limit is only honored for an exactly named model.

Solutions

  1. Name an exact model in req.model when setting max_output_tokens.
  2. Clear max_output_tokens (None) if you want Auto routing.
  3. Verify the chosen model's transport supports output limits via route_supports_output_token_limit.

Example fix

// before
req.max_output_tokens = Some(4096); // model left as "auto"
// after
req.model = "gpt-4.1".into();
req.max_output_tokens = Some(4096);
Defensive patterns

Strategy: validation

Validate before calling

if req.max_output_tokens.is_some() && req.model.as_deref().unwrap_or_default().eq_ignore_ascii_case("auto") {
    return Err("max_output_tokens needs an exact model");
}

Prevention

When it happens

Trigger: Setting ThreadTurnRequest::max_output_tokens to Some(..) while req.model / thread.model resolves to "auto" (case-insensitive).

Common situations: Scripts that set a token budget for cost control but leave model unset/Auto; combining a UI token-limit slider with an Auto model picker; older clients that always set max_output_tokens.

Understand the failure class

Background: Conflicting config options: "cannot be used together" — configuration validation errors across open-source libraries — this error's family across 162 libraries.

Related errors


AI-assisted analysis of Hmbown/CodeWhale@73e0f67d83 (2026-09-22). Data as JSON: /api/errors/b45d7245457e3341. Report an issue: GitHub.

Appendix: source

Thrown at crates/tui/src/runtime_threads.rs:9880

        }
        let cfg_snapshot = self.config.read().clone();
        let configured_reasoning_preference = cfg_snapshot
            .reasoning_effort()
            .map(crate::reasoning_preference::ReasoningEffort::from_setting);
        // Runtime API precedence is explicit and stable: a turn override wins
        // over its persisted thread default, which wins over normal config.
        let reasoning_preference = turn_reasoning_preference
            .or(thread_reasoning_preference)
            .or(configured_reasoning_preference);
        let allow_shell = req.allow_shell.unwrap_or(thread.allow_shell);
        let trust_mode = req.trust_mode.unwrap_or(thread.trust_mode);
        let auto_approve = policy.auto_approve();
        let allowed_tools = req
            .allowed_tools
            .clone()
            .or_else(|| thread.allowed_tools.clone());
        if req.max_output_tokens.is_some() && auto_model {
            bail!("maxOutputTokens requires an exact model; Auto routing is unsupported");
        }
        let operation = if let Some(operation_key) = req.operation_key.as_deref() {
            validate_runtime_turn_operation_key(operation_key)?;
            let request_fingerprint = runtime_turn_request_fingerprint(
                &thread,
                &prompt,
                req.input_summary.as_deref(),
                &requested_model,
                reasoning_preference,
                allowed_tools.as_deref(),
                policy,
                allow_shell,
                trust_mode,
                &req.dynamic_tools,
                req.environment_id.as_deref(),
                &req.images,
                req.max_output_tokens,
            )?;

View on GitHub (pinned to 73e0f67d83)