BoundaryML/baml · error

OpenAI transcriptions only support audio media parts

Error message

OpenAI transcriptions only support audio media parts

What it means

When walking chat message parts for a transcription request, any Media part whose media_type is not Audio (e.g. Image or Video) is rejected, because the OpenAI transcription endpoint accepts only audio input. Non-audio media in the prompt indicates the wrong function/client pairing.

Solutions

  1. Remove non-audio media parts from the transcription function's prompt.
  2. Point the function at a chat/completions client if you need image inputs.
  3. Check variable types passed into the template to ensure only Audio media is included.

Example fix

// before
prompt #"{{ image }} {{ audio }}"# // image not allowed
// after
prompt #"{{ _.role('user') }} {{ audio }}"#
Defensive patterns

Strategy: validation

Validate before calling

// Validate prompt media types before invoking a transcription function
function assertAudioOnly(messages) {
  for (const m of messages) for (const p of m.parts) {
    if (p.type === "media" && p.media_type !== "audio") {
      throw new Error(`Transcription prompt has non-audio media: ${p.media_type}`);
    }
  }
}

Type guard

const isAudioPart = (p) => p.type === "media" && p.media_type === "audio";

Prevention

When it happens

Trigger: A BAML function targeting a transcription client whose prompt includes an image or other non-audio media part, e.g. {{ image }} rendered in the template.

Common situations: Reusing a prompt template between a vision function and a transcription function; passing the wrong variable type into the template; pointing an image-input function at a whisper/transcription client.

Understand the failure class

Background: UnsupportedOperationException and "is not supported" errors: when a library deliberately refuses a call — this error's family across 30 libraries.

Related errors


AI-assisted analysis of BoundaryML/baml@bd85ce9dee (2026-09-12). Data as JSON: /api/errors/ed2db93eea3d053f. Report an issue: GitHub.

Appendix: source

Thrown at engine/baml-runtime/src/internal/llm_client/primitive/openai/types.rs:141

        fields,
    })
}

fn collect_transcription_prompt_parts(
    part: &ChatMessagePart,
    audio_parts: &mut Vec<BamlMedia>,
    text_parts: &mut Vec<String>,
) -> Result<()> {
    match part {
        ChatMessagePart::Text(text) => {
            let text = text.trim();
            if !text.is_empty() {
                text_parts.push(text.to_string());
            }
        }
        ChatMessagePart::Media(media) => {
            if media.media_type != BamlMediaType::Audio {
                bail!("OpenAI transcriptions only support audio media parts")
            }
            audio_parts.push(media.clone());
        }
        ChatMessagePart::WithMeta(inner, _) => {
            collect_transcription_prompt_parts(inner, audio_parts, text_parts)?;
        }
    }
    Ok(())
}

fn reject_reserved_request_fields(properties: &BamlMap<String, Value>) -> Result<()> {
    for key in ["messages", "stream"] {
        if properties.contains_key(key) {
            bail!("OpenAI transcriptions do not support reserved request field `{key}`")
        }
    }
    Ok(())
}

View on GitHub (pinned to bd85ce9dee)