moeru-ai/airi · error
Gemini TTS response missing audio data
Error message
Gemini TTS response missing audio data
What it means
The Gemini API responded 200 but the response JSON contained no inlineData audio in candidates[0].content.parts. Gemini can return successful HTTP responses without audio when generation is blocked or empty (safety filters, empty input, prompt feedback).
Solutions
- Log the full response JSON (before this point) to inspect candidates, finishReason, and promptFeedback
- Check finishReason for SAFETY/blocking and adjust input text or safety settings
- Verify the model id is actually a TTS model that returns inline audio data
- Handle the empty-audio case in caller code and surface a user-facing message instead of failing silently
Defensive patterns
Strategy: try-catch
Validate before calling
// no client-side check can detect this; validate input text is non-empty to reduce risk if (!input?.trim()) return null
Type guard
function hasAudio(data: unknown): data is { candidates: Array<{ content: { parts: Array<{ inlineData: { data: string } }> } }> } {
const d = data as any
return Boolean(d?.candidates?.[0]?.content?.parts?.some((p: any) => p?.inlineData?.data))
} Try / catch
try {
const audio = await speech({ input, model })
} catch (e) {
if (e instanceof Error && e.message === 'Gemini TTS response missing audio data') {
// inspect finishReason/safety feedback, adjust input, or inform the user
}
} Prevention
- Avoid input text likely to trip safety filters
- Check finishReason and promptFeedback when synthesis comes back empty
- Pin and track the preview TTS API for schema changes
- Treat empty audio as a distinct UI state, not a generic crash
When it happens
Trigger: The API returned a candidates array with no parts carrying inlineData: blocked by safety settings, empty result for the given text, or a response shape change in the preview API.
Common situations: Input text triggers safety filtering, using a non-TTS Gemini model that returns text-only parts, or Google changing the preview response schema.
Understand the failure class
Background: "empty response", "returned no data", "empty embeddings": what HTTP 200-with-empty-body errors mean across libraries — this error's family across 36 libraries.
Related errors
- Gemini TTS response missing audio data
- Gemini TTS request failed
- Gemini TTS request failed
- MiMo TTS response missing audio data
- MiMo TTS response missing audio data
AI-assisted analysis of moeru-ai/airi@438a067dde (2026-09-08).
Data as JSON: /api/errors/58c8d69dfbb48834.
Report an issue: GitHub.
Appendix: source
Thrown at packages/provider-inference/src/providers/cloud/google-gemini-audio-speech/index.ts:106
contents: [{ parts: [{ text: body.input }] }],
generationConfig: {
responseModalities: ['AUDIO'],
speechConfig: {
voiceConfig: { prebuiltVoiceConfig: { voiceName: body.voice || 'Kore' } },
},
...(body.temperature !== undefined ? { temperature: body.temperature } : {}),
},
}),
})
if (!response.ok)
throw new Error(`Gemini TTS request failed: ${response.status} ${await response.text().catch(() => '')}`)
const data = await response.json() as {
candidates?: Array<{ content?: { parts?: Array<{ inlineData?: { data?: string } }> } }>
}
const audio = data.candidates?.[0]?.content?.parts?.find(part => part.inlineData)?.inlineData?.data
if (!audio)
throw new Error('Gemini TTS response missing audio data')
return new Response(toWavFromPCM16(decodeBase64(audio), 24000), {
status: 200,
headers: { 'Content-Type': 'audio/wav' },
})
}
}
export const providerGoogleGeminiAudioSpeech = defineProvider<GoogleGeminiSpeechConfig, 'google-gemini-audio-speech'>({
id: 'google-gemini-audio-speech',
name: 'Google Gemini',
nameLocalize: ({ t }) => t('settings.pages.providers.provider.google-gemini-audio-speech.title'),
description: 'aistudio.google.com',
descriptionLocalize: ({ t }) => t('settings.pages.providers.provider.google-gemini-audio-speech.description'),
tasks: ['text-to-speech', 'tts'],
icon: 'i-lobe-icons:gemini',
iconColor: 'i-lobe-icons:gemini-color',
createProviderConfig: () => googleGeminiSpeechConfigSchema,View on GitHub (pinned to 438a067dde)