run-llama/llama_index · error · ValueError

Streaming is not enabled. Please use chat() instead.

Error message

Streaming is not enabled. Please use chat() instead.

What it means

Raised by CondenseQuestionChatEngine.stream_chat() when the underlying query_engine returns a non-streaming response. The code checks whether the condensed query produced a response with response_gen; if streaming was not requested from the query engine, there is no token generator to expose and the engine raises ValueError telling you to use chat() instead.

Solutions

  1. Build the query engine with streaming: engine = index.as_query_engine(streaming=True), then pass it to CondenseQuestionChatEngine.from_defaults
  2. Or call chat() instead of stream_chat() if streaming is not required
  3. For retriever-based engines, verify streaming support of the LLM (some custom LLMs lack astream_complete)

Example fix

# before
qe = index.as_query_engine()  # streaming=False by default
chat = CondenseQuestionChatEngine.from_defaults(query_engine=qe)
chat.stream_chat('hello')  # ValueError: Streaming is not enabled

# after
qe = index.as_query_engine(streaming=True)
chat = CondenseQuestionChatEngine.from_defaults(query_engine=qe)
resp = chat.stream_chat('hello')
for token in resp.response_gen:
    print(token, end='')
Defensive patterns

Strategy: validation

Validate before calling

qe = index.as_query_engine(streaming=True)  # set at construction, not later
assert getattr(qe, 'streaming', True) is not False, 'Query engine must be built with streaming=True to use stream_chat()'

Try / catch

try:
    resp = engine.stream_chat(msg)
except ValueError as e:
    if 'Streaming is not enabled' in str(e):
        resp = engine.chat(msg)  # graceful fallback to non-streaming
    else:
        raise

Prevention

When it happens

Trigger: engine = CondenseQuestionChatEngine.from_defaults(query_engine=RetrieverQueryEngine.from_defaults(retriever, streaming=False)); engine.stream_chat('hi'). Any query engine constructed without streaming=True (the default) whose query() path is exercised by stream_chat.

Common situations: Calling stream_chat on an engine built from a default query engine; building the query engine from an index (index.as_query_engine()) which defaults to streaming=False; swapping engines in a streaming UI without updating construction flags.

Related errors


AI-assisted analysis of run-llama/llama_index@afd0fef371 (2026-08-15). Data as JSON: /api/errors/58e14a86a85215e3. Report an issue: GitHub.

Appendix: source

Thrown at llama-index-core/llama_index/core/chat_engine/condense_question.py:273

        # Record response
        if (
            isinstance(query_response, StreamingResponse)
            and query_response.response_gen is not None
        ):
            # override the generator to include writing to chat history
            self._memory.put(ChatMessage(role=MessageRole.USER, content=message))
            response = StreamingAgentChatResponse(
                chat_stream=response_gen_from_query_engine(query_response.response_gen),
                sources=[tool_output],
            )
            thread = Thread(
                target=response.write_response_to_history,
                args=(self._memory,),
            )
            response.write_response_to_history_thread = thread
            thread.start()
        else:
            raise ValueError("Streaming is not enabled. Please use chat() instead.")
        return response

    @trace_method("chat")
    async def achat(
        self, message: str, chat_history: Optional[List[ChatMessage]] = None
    ) -> AgentChatResponse:
        chat_history = chat_history or await self._memory.aget(input=message)

        # Generate standalone question from conversation context and last message
        condensed_question = await self._acondense_question(chat_history, message)

        log_str = f"Querying with: {condensed_question}"
        logger.info(log_str)
        if self._verbose:
            print(log_str)

        # TODO: right now, query engine uses class attribute to configure streaming,
        #       we are moving towards separate streaming and non-streaming methods.

View on GitHub (pinned to afd0fef371)