BoundaryML/baml · error

OpenAI transcriptions require exactly one audio media part…

Error message

OpenAI transcriptions require exactly one audio media part, got {}

What it means

After collecting media parts from all chat messages, build_transcription_parts requires exactly one audio media part because an OpenAI transcription request carries a single audio file. Zero audio parts means nothing to transcribe; more than one is ambiguous, so the count is validated and the actual number reported.

Solutions

  1. Include exactly one audio media part in the prompt (e.g. {{ audio }} in a single user message).
  2. Remove duplicate audio parts if the prompt template references the audio more than once.
  3. Verify the media variable is actually an Audio and rendered as a media part, not plain text.

Example fix

// before (two audio parts)
prompt #"{{ audio }} {{ audio2 }}"#
// after
prompt #"
  {{ _.role('user') }}
  {{ audio }}
  Please transcribe.
"#
Defensive patterns

Strategy: validation

Validate before calling

// Count audio media parts in the rendered prompt before calling the function
function assertSingleAudio(messages) {
  const audioParts = messages.flatMap(m => m.parts).filter(p => p.type === "audio");
  if (audioParts.length !== 1) {
    throw new Error(`Transcription prompt must contain exactly one audio part, got ${audioParts.length}`);
  }
}

Prevention

When it happens

Trigger: A chat prompt with no media parts at all; a prompt with two or more audio parts; multiple user messages each attaching audio, all fed to a transcription-configured client.

Common situations: Forgetting to include the audio variable in the BAML prompt template; concatenating prompts that each embed audio; building messages programmatically and appending audio twice.

Understand the failure class

Background: "Must be a positive integer", "Invalid value", "Unsupported": the invalid-argument-value error family, when a library rejects the value you pass — this error's family across 35 libraries.

Related errors


AI-assisted analysis of BoundaryML/baml@bd85ce9dee (2026-09-12). Data as JSON: /api/errors/53a7d10cd0b94577. Report an issue: GitHub.

Appendix: source

Thrown at engine/baml-runtime/src/internal/llm_client/primitive/openai/types.rs:77

    let messages = match prompt {
        either::Either::Right(messages) => messages,
        either::Either::Left(_) => {
            bail!("OpenAI transcriptions require chat messages with exactly one audio media part")
        }
    };

    reject_reserved_request_fields(properties)?;

    let mut audio_parts = Vec::new();
    let mut text_parts = Vec::new();
    for message in messages {
        for part in &message.parts {
            collect_transcription_prompt_parts(part, &mut audio_parts, &mut text_parts)?;
        }
    }

    if audio_parts.len() != 1 {
        bail!(
            "OpenAI transcriptions require exactly one audio media part, got {}",
            audio_parts.len()
        );
    }

    let audio = audio_parts
        .pop()
        .expect("audio_parts length was already validated");
    let mime = audio.mime_type_as_ok()?;
    let file_bytes = match &audio.content {
        BamlMediaContent::Base64(media_b64) => BASE64_STANDARD
            .decode(&media_b64.base64)
            .context("Failed to decode transcription audio as base64")?,
        BamlMediaContent::Url(_) | BamlMediaContent::File(_) => {
            bail!("OpenAI transcription audio must be resolved to base64 before request building")
        }
    };

View on GitHub (pinned to bd85ce9dee)