{"record":{"id":"0ebfacef3ca406ea","repo":"moeru-ai/airi","slug":"gemini-tts-response-missing-audio-data","errorCode":null,"errorMessage":"Gemini TTS response missing audio data","messagePattern":"Gemini TTS response missing audio data","errorType":"exception","errorClass":"Error","httpStatus":null,"severity":"error","filePath":"packages/stage-ui/src/libs/providers/providers/google-gemini-audio-speech/index.ts","lineNumber":106,"sourceCode":"        contents: [{ parts: [{ text: body.input }] }],\n        generationConfig: {\n          responseModalities: ['AUDIO'],\n          speechConfig: {\n            voiceConfig: { prebuiltVoiceConfig: { voiceName: body.voice || 'Kore' } },\n          },\n          ...(body.temperature !== undefined ? { temperature: body.temperature } : {}),\n        },\n      }),\n    })\n    if (!response.ok)\n      throw new Error(`Gemini TTS request failed: ${response.status} ${await response.text().catch(() => '')}`)\n\n    const data = await response.json() as {\n      candidates?: Array<{ content?: { parts?: Array<{ inlineData?: { data?: string } }> } }>\n    }\n    const audio = data.candidates?.[0]?.content?.parts?.find(part => part.inlineData)?.inlineData?.data\n    if (!audio)\n      throw new Error('Gemini TTS response missing audio data')\n\n    return new Response(toWavFromPCM16(decodeBase64(audio), 24000), {\n      status: 200,\n      headers: { 'Content-Type': 'audio/wav' },\n    })\n  }\n}\n\nexport const providerGoogleGeminiAudioSpeech = defineProvider<GoogleGeminiSpeechConfig>({\n  id: 'google-gemini-audio-speech',\n  name: 'Google Gemini',\n  nameLocalize: ({ t }) => t('settings.pages.providers.provider.google-gemini-audio-speech.title'),\n  description: 'aistudio.google.com',\n  descriptionLocalize: ({ t }) => t('settings.pages.providers.provider.google-gemini-audio-speech.description'),\n  tasks: ['text-to-speech', 'tts'],\n  icon: 'i-lobe-icons:gemini',\n  iconColor: 'i-lobe-icons:gemini-color',\n  createProviderConfig: () => googleGeminiSpeechConfigSchema,","sourceCodeStart":88,"sourceCodeEnd":124,"githubUrl":"https://github.com/moeru-ai/airi/blob/677329427f32468c74b17f3ec47eeca4e05bec65/packages/stage-ui/src/libs/providers/providers/google-gemini-audio-speech/index.ts#L88-L124","documentation":"The Gemini call succeeded at the HTTP level, but candidates[0].content.parts contains no part with inlineData — the model produced no audio bytes — so the provider refuses rather than returning a broken WAV. A 200 without audio usually means the response was blocked or the model answered in text instead of generating speech.","triggerScenarios":"Safety filters blocked the prompt (promptFeedback.blockReason or empty candidates); a finishReason like MAX_TOKENS or SAFETY cut generation; a model id that does not support the AUDIO response modality; an invalid voice name causing the API to return text only.","commonSituations":"Prompts with content the safety layer flags; the model swapped to a text-only variant while the voice config stayed; long input truncated before any audio part; regional model behavior differences.","solutions":["Log the full response JSON — promptFeedback and finishReason explain the empty audio","Adjust the prompt text that triggers safety blocks","Confirm the model actually supports audio output with responseModalities AUDIO","Retry with the default voice 'Kore' to rule out an invalid voiceName"],"exampleFix":"// before\nconst data = await response.json()\nconst audio = data.candidates?.[0]?.content?.parts?.find(p => p.inlineData)?.inlineData?.data\nif (!audio) throw new Error('Gemini TTS response missing audio data') // opaque\n\n// after\nconst data = await response.json()\nconst reason = data.promptFeedback?.blockReason ?? data.candidates?.[0]?.finishReason\nif (!audio) throw new Error(`Gemini returned no audio: ${reason ?? 'unknown reason'}`)","handlingStrategy":"fallback","validationCode":null,"typeGuard":null,"tryCatchPattern":"try {\n  await speak(text)\n}\ncatch (err) {\n  if (err.message.includes('missing audio data')) {\n    // Retry once with sanitized text and the default voice before giving up.\n    await speak(sanitize(text), { voice: 'Kore' })\n  }\n}","preventionTips":["Inspect finishReason and promptFeedback whenever audio is absent","Keep synthesis prompts short and neutral to avoid safety stops","Validate the model supports AUDIO output before configuring it"],"tags":["gemini","tts","empty-response","safety-filter"],"backgroundTag":"empty-api-response","analyzedSha":"677329427f32468c74b17f3ec47eeca4e05bec65","analyzedAt":"2026-08-18T17:29:58.153Z","contentChangedAt":"2026-08-18T17:29:58.153Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}