moeru-ai/airi · error · Error
Video tool output is not supported by the conversation model
Error message
Video tool output is not supported by the conversation model
What it means
The Responses adapter's readToolResultContent maps function_call_output parts from a stored conversation into portable tool-result segments. When it encounters an 'input_video' part it has no portable representation and throws immediately. Video tool output is intentionally not supported by the conversation model projection, so the adapter fails fast rather than silently dropping the segment.
Solutions
- Remove or replace the input_video part in the tool result content before it is stored in the conversation (e.g. emit input_text describing the video instead).
- Filter out input_video parts when building function_call_output so only input_text/input_image/input_file parts are sent.
- If video support is needed, extend readToolResultContent with a portable mapping for input_video and file an upstream issue, since the conversation model currently has no video segment type.
Example fix
// before
return content.map(part => ({ type: part.type, ... }))
// after
const supported = content.filter(p => p.type !== 'input_video')
return supported.map(part => ({ type: part.type, ... })) Defensive patterns
Strategy: validation
Validate before calling
function hasVideoOutput(output) {
return Array.isArray(output) && output.some(p => p && p.type === 'input_video')
}
if (hasVideoOutput(segment.content)) throw new Error('strip input_video before recording tool output') Type guard
function isSupportedToolOutputPart(part) {
return ['input_text', 'input_image', 'input_file'].includes(part?.type)
} Try / catch
try {
const segments = readOutput(items)
} catch (err) {
if (err.message.includes('Video tool output is not supported')) {
console.error('Tool emitted video output; convert to text summary before recording.', err)
} else throw err
} Prevention
- Never place input_video parts in function_call_output content; represent video results as input_text summaries or input_file references.
- Add a whitelist serializer for tool result content before it enters the conversation.
- Test tool result replay for every content type your tools can emit.
When it happens
Trigger: readOutput() -> readToolResultContent() processes a function_call_output item whose output array contains a part with type 'input_video'. This happens when replaying/re-projecting a conversation that previously recorded a tool result containing video input, or when code constructs a Responses function_call_output with an input_video part.
Common situations: A tool (e.g. a media tool) returned video content that was recorded into the Responses continuation data; the same conversation is later re-projected for the next model turn. Also occurs after a library upgrade where a tool began emitting input_video parts that the adapter never handled.
Understand the failure class
Background: UnsupportedOperationException and "is not supported" errors: when a library deliberately refuses a call — this error's family across 30 libraries.
Related errors
- Only assistant messages can invoke tools
- Responses continuation must contain an item array
- Responses file requires exactly one source
- Responses image output requires a URL
- Tool messages require a correlated tool result
AI-assisted analysis of moeru-ai/airi@438a067dde (2026-09-17).
Data as JSON: /api/errors/e0c8b2742269f917.
Report an issue: GitHub.
Appendix: source
Thrown at packages/core-agent/src/runtime/responses.ts:108
}
function readToolResultContent(content: Extract<ItemParam, { type: 'function_call_output' }>['output']): InputSegment[] {
if (typeof content === 'string')
return [{ type: 'text', text: content }]
return content.map((part) => {
switch (part.type) {
case 'input_text': return { type: 'text', text: part.text }
case 'input_image':
if (!part.image_url)
throw new Error('Responses image output requires a URL')
return { type: 'image', url: part.image_url, detail: part.detail ?? undefined }
case 'input_file':
if (part.file_data != null && part.file_url == null)
return { type: 'file', data: part.file_data, name: part.filename ?? undefined }
if (part.file_url != null && part.file_data == null)
return { type: 'file', url: part.file_url, name: part.filename ?? undefined }
throw new Error('Responses file requires exactly one source')
case 'input_video': throw new Error('Video tool output is not supported by the conversation model')
}
throw new Error('Unsupported Responses tool output')
})
}
type AssistantContent = Exclude<Extract<ItemParam, { role: 'assistant' }>['content'], string>[number]
function readCitations(part: Extract<AssistantContent, { type: 'output_text' }>): Citation[] | undefined {
return part.annotations?.map(entry => ({
url: entry.url,
title: entry.title,
startIndex: entry.start_index,
endIndex: entry.end_index,
}))
}
function readOutput(items: ItemParam[]): ProjectionEntry[] {
return items.flatMap<ProjectionEntry>((item, index) => {View on GitHub (pinned to 438a067dde)