affaan-m/ECC · error · ContextLengthError
ContextLengthError(msg, provider=ProviderType.CLAUDE) from e
Error message
ContextLengthError(msg, provider=ProviderType.CLAUDE) from e
What it means
The Claude LLM provider wraps all exceptions from its generate() call and re-raises them as typed provider errors. When the underlying Claude API error message contains both the words 'context' and 'length', the provider concludes the request exceeded the model's context window and raises ContextLengthError, preserving the original exception with 'from e'.
Solutions
- Reduce the prompt size: truncate or summarize conversation history before sending
- Lower input.max_tokens so prompt + max_tokens fits the model window
- Use a model variant with a larger context window
- Count tokens client-side (e.g. a tokenizer) before calling generate and reject oversized requests early
Example fix
// before resp = provider.generate(LLMInput(prompt=full_transcript)) // after MAX_CHARS = 100_000 resp = provider.generate(LLMInput(prompt=full_transcript[-MAX_CHARS:]))
Defensive patterns
Strategy: try-catch
Validate before calling
def fits_context(text, max_model_tokens, max_output):
approx_tokens = len(text) // 4 # rough estimate; use a tokenizer for accuracy
return approx_tokens + max_output <= max_model_tokens Try / catch
try:
resp = provider.generate(inp)
except ContextLengthError as e:
logger.warning("Prompt too long for model, truncating")
inp = replace(inp, prompt=inp.prompt[-len(inp.prompt)//2:])
resp = provider.generate(inp) Prevention
- Estimate or count tokens before every generate call
- Trim conversation history to a rolling window
- Keep max_tokens well below the model's window
- Pin the model and know its exact context limit
When it happens
Trigger: Calling provider.generate() where the combined prompt + history + max_tokens exceeds the Claude model's context window; the SDK raises an error whose string form contains 'context ... length'.
Common situations: Feeding whole documents or long transcripts into the prompt; accumulating chat history without truncation; passing a large system prompt plus a large max_tokens; using a model with a smaller window than assumed (e.g. switching from a 200k model to a smaller one).
Related errors
- ContextLengthError(msg, provider=ProviderType.OLLAMA) from e
- Claude session not found
- ContextLengthError(msg, provider=ProviderType.OPENAI) from e
- empty response
- LLM returned empty or filtered response
AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16).
Data as JSON: /api/errors/a73256fc049c06c1.
Report an issue: GitHub.
Appendix: source
Thrown at src/llm/providers/claude.py:132
model=response.model,
usage={
"input_tokens": response.usage.input_tokens,
"output_tokens": response.usage.output_tokens,
"cache_creation_input_tokens": getattr(
response.usage, "cache_creation_input_tokens", 0
),
"cache_read_input_tokens": getattr(response.usage, "cache_read_input_tokens", 0),
},
stop_reason=response.stop_reason,
)
except Exception as e:
msg = str(e)
if "401" in msg or "authentication" in msg.lower():
raise AuthenticationError(msg, provider=ProviderType.CLAUDE) from e
if "429" in msg or "rate_limit" in msg.lower():
raise RateLimitError(msg, provider=ProviderType.CLAUDE) from e
if "context" in msg.lower() and "length" in msg.lower():
raise ContextLengthError(msg, provider=ProviderType.CLAUDE) from e
raise
def list_models(self) -> list[ModelInfo]:
return self._models.copy()
def validate_config(self) -> bool:
return bool(self.client.api_key)
def get_default_model(self) -> str:
return _DEFAULT_MODEL
View on GitHub (pinned to 8321021c54)