affaan-m/ECC · error · ContextLengthError

ContextLengthError(msg, provider=ProviderType.OPENAI) from e

Error message

ContextLengthError(msg, provider=ProviderType.OPENAI) from e

What it means

The OpenAI provider raises ContextLengthError when the API error message contains both 'context' and 'length', i.e. the request exceeded the model's maximum context length (prompt tokens + max_tokens). The original SDK exception is chained via 'from e'.

Solutions

  1. Truncate or summarize the prompt/history to fit the model window
  2. Lower input.max_tokens to leave room for the prompt
  3. Switch to a larger-context model (e.g. gpt-4o 128k)
  4. Count tokens before sending with tiktoken and reject oversized requests early

Example fix

// before
messages=[{"role":"user","content":big_doc}]
resp = provider.generate(LLMInput(messages=messages, max_tokens=4096))
// after
messages=[{"role":"user","content":big_doc[:60000]}]
resp = provider.generate(LLMInput(messages=messages, max_tokens=1024))
Defensive patterns

Strategy: validation

Validate before calling

import tiktoken
def fits(prompt: str, model: str, max_tokens_out: int) -> bool:
    enc = tiktoken.encoding_for_model(model)
    return len(enc.encode(prompt)) + max_tokens_out < MODEL_LIMITS[model]

Try / catch

try:
    resp = provider.generate(inp)
except ContextLengthError:
    inp = replace(inp, prompt=summarize(inp.prompt))
    resp = provider.generate(inp)

Prevention

When it happens

Trigger: Calling generate() where the tokenized prompt + max_tokens exceeds the model limit (e.g. 8191 for gpt-3.5-turbo, 128k for gpt-4o); error text like 'This model's maximum context length is ...'.

Common situations: Sending large documents/RAG dumps without truncation; long multi-turn chats without history trimming; requesting large max_tokens on top of a near-limit prompt; downgrading to a smaller-context model.

Related errors


AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16). Data as JSON: /api/errors/ed70d260e5ba588f. Report an issue: GitHub.

Appendix: source

Thrown at src/llm/providers/openai.py:121

                    "completion_tokens": response.usage.completion_tokens,
                    "total_tokens": response.usage.total_tokens,
                }

            return LLMOutput(
                content=choice.message.content or "",
                tool_calls=tool_calls,
                model=response.model,
                usage=usage,
                stop_reason=choice.finish_reason,
            )
        except Exception as e:
            msg = str(e)
            if "401" in msg or "authentication" in msg.lower():
                raise AuthenticationError(msg, provider=ProviderType.OPENAI) from e
            if "429" in msg or "rate_limit" in msg.lower():
                raise RateLimitError(msg, provider=ProviderType.OPENAI) from e
            if "context" in msg.lower() and "length" in msg.lower():
                raise ContextLengthError(msg, provider=ProviderType.OPENAI) from e
            raise

    def list_models(self) -> list[ModelInfo]:
        return self._models.copy()

    def validate_config(self) -> bool:
        return bool(self.client.api_key)

    def get_default_model(self) -> str:
        return "gpt-4o-mini"

View on GitHub (pinned to 8321021c54)