{"record":{"id":"ed2db93eea3d053f","repo":"BoundaryML/baml","slug":"openai-transcriptions-only-support-audio-media-parts","errorCode":null,"errorMessage":"OpenAI transcriptions only support audio media parts","messagePattern":"OpenAI transcriptions only support audio media parts","errorType":"validation","errorClass":null,"httpStatus":null,"severity":"error","filePath":"engine/baml-runtime/src/internal/llm_client/primitive/openai/types.rs","lineNumber":141,"sourceCode":"        fields,\n    })\n}\n\nfn collect_transcription_prompt_parts(\n    part: &ChatMessagePart,\n    audio_parts: &mut Vec<BamlMedia>,\n    text_parts: &mut Vec<String>,\n) -> Result<()> {\n    match part {\n        ChatMessagePart::Text(text) => {\n            let text = text.trim();\n            if !text.is_empty() {\n                text_parts.push(text.to_string());\n            }\n        }\n        ChatMessagePart::Media(media) => {\n            if media.media_type != BamlMediaType::Audio {\n                bail!(\"OpenAI transcriptions only support audio media parts\")\n            }\n            audio_parts.push(media.clone());\n        }\n        ChatMessagePart::WithMeta(inner, _) => {\n            collect_transcription_prompt_parts(inner, audio_parts, text_parts)?;\n        }\n    }\n    Ok(())\n}\n\nfn reject_reserved_request_fields(properties: &BamlMap<String, Value>) -> Result<()> {\n    for key in [\"messages\", \"stream\"] {\n        if properties.contains_key(key) {\n            bail!(\"OpenAI transcriptions do not support reserved request field `{key}`\")\n        }\n    }\n    Ok(())\n}","sourceCodeStart":123,"sourceCodeEnd":159,"githubUrl":"https://github.com/BoundaryML/baml/blob/bd85ce9dee1463ff04d27efd20531013a4ff46c1/engine/baml-runtime/src/internal/llm_client/primitive/openai/types.rs#L123-L159","documentation":"When walking chat message parts for a transcription request, any Media part whose media_type is not Audio (e.g. Image or Video) is rejected, because the OpenAI transcription endpoint accepts only audio input. Non-audio media in the prompt indicates the wrong function/client pairing.","triggerScenarios":"A BAML function targeting a transcription client whose prompt includes an image or other non-audio media part, e.g. {{ image }} rendered in the template.","commonSituations":"Reusing a prompt template between a vision function and a transcription function; passing the wrong variable type into the template; pointing an image-input function at a whisper/transcription client.","solutions":["Remove non-audio media parts from the transcription function's prompt.","Point the function at a chat/completions client if you need image inputs.","Check variable types passed into the template to ensure only Audio media is included."],"exampleFix":"// before\nprompt #\"{{ image }} {{ audio }}\"# // image not allowed\n// after\nprompt #\"{{ _.role('user') }} {{ audio }}\"#","handlingStrategy":"validation","validationCode":"// Validate prompt media types before invoking a transcription function\nfunction assertAudioOnly(messages) {\n  for (const m of messages) for (const p of m.parts) {\n    if (p.type === \"media\" && p.media_type !== \"audio\") {\n      throw new Error(`Transcription prompt has non-audio media: ${p.media_type}`);\n    }\n  }\n}","typeGuard":"const isAudioPart = (p) => p.type === \"media\" && p.media_type === \"audio\";","tryCatchPattern":null,"preventionTips":["Use dedicated prompt templates for transcription vs vision functions.","Type-check function signatures so only Audio params reach transcription templates.","Verify the client assignment (baml-cli check) when reusing functions across clients."],"tags":["openai","transcription","media","baml"],"backgroundTag":"unsupported-operation","analyzedSha":"bd85ce9dee1463ff04d27efd20531013a4ff46c1","analyzedAt":"2026-09-12T03:38:25.718Z","contentChangedAt":"2026-09-12T03:38:25.718Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}