Hmbown/CodeWhale · error · anyhow::Error

The selected Runtime model does not support maxOutputTokens

Error message

The selected Runtime model does not support maxOutputTokens

What it means

While paging the Runtime model catalog to confirm the selected model supports `maxOutputTokens`, the model entry was found but its `output_token_limit` is not "supported" (e.g. "unsupported" or absent). The requested output-token cap cannot be honored for that model.

Solutions

  1. Remove `maxOutputTokens` from the request for this model.
  2. Select a model whose catalog entry reports `output_token_limit: "supported"`.
  3. Upgrade the Runtime/provider if support for the model was expected.

Example fix

// before
{ "model": "legacy-mini", "maxOutputTokens": 512 }
// after
{ "model": "qwen3-coder" }
Defensive patterns

Strategy: fallback

Validate before calling

// check capability before requesting a token cap
const entry = catalog.models.find(m => m.id === model);
const supportsCap = entry && entry.output_token_limit === 'supported';
const payload = supportsCap ? { model, maxOutputTokens: cap } : { model };

Type guard

function supportsOutputTokenLimit(entry) {
  return !!entry && entry.output_token_limit === 'supported';
}

Try / catch

// fall back to an uncapped request
match send_with_cap(model, cap).await {
    Err(e) if e.to_string().contains("does not support maxOutputTokens") => {
        send_without_cap(model).await
    }
    other => other,
}

Prevention

When it happens

Trigger: The catalog page containing the selected model id has `entry["output_token_limit"] != "supported"` while resolving maxOutputTokens.

Common situations: Using a model that has no output-token-limit support (older or constrained models) and passing maxOutputTokens with the request.

Understand the failure class

Background: UnsupportedOperationException and "is not supported" errors: when a library deliberately refuses a call — this error's family across 30 libraries.

Related errors


AI-assisted analysis of Hmbown/CodeWhale@73e0f67d83 (2026-09-22). Data as JSON: /api/errors/8d966265794a998f. Report an issue: GitHub.

Appendix: source

Thrown at crates/app-server/src/lib.rs:1682

            url.query_pairs_mut().append_pair("limit", "250");
            if let Some(cursor) = cursor.as_deref() {
                url.query_pairs_mut().append_pair("cursor", cursor);
            }
            let catalog = self.request_json(self.authed(self.client.get(url))).await?;
            if let Some(entry) =
                catalog
                    .get("models")
                    .and_then(Value::as_array)
                    .and_then(|models| {
                        models
                            .iter()
                            .find(|entry| entry.get("id").and_then(Value::as_str) == Some(model))
                    })
            {
                if entry.get("output_token_limit").and_then(Value::as_str) == Some("supported") {
                    return Ok(());
                }
                bail!("The selected Runtime model does not support maxOutputTokens");
            }
            let next = catalog
                .get("nextCursor")
                .and_then(Value::as_str)
                .context("Output-limit support is unknown for the selected Runtime model")?;
            if !seen.insert(next.to_string()) {
                bail!("Runtime model catalog cursor repeated");
            }
            cursor = Some(next.to_string());
        }
    }

    /// Resolve `stdio_thread_id` to a runtime thread, minting one only when
    /// `thread_map` has no entry. The map lives on [`AppState`] and outlives
    /// this bridge, so a thread created under a previous child keeps its id
    /// here as long as the store it was persisted to is shared (#6246).
    async fn ensure_runtime_thread(
        &mut self,

View on GitHub (pinned to 73e0f67d83)