affaan-m/ECC · error · ContextLengthError

ContextLengthError(msg, provider=ProviderType.CLAUDE) from e

Error message

ContextLengthError(msg, provider=ProviderType.CLAUDE) from e

What it means

The Claude LLM provider wraps all exceptions from its generate() call and re-raises them as typed provider errors. When the underlying Claude API error message contains both the words 'context' and 'length', the provider concludes the request exceeded the model's context window and raises ContextLengthError, preserving the original exception with 'from e'.

Solutions

  1. Reduce the prompt size: truncate or summarize conversation history before sending
  2. Lower input.max_tokens so prompt + max_tokens fits the model window
  3. Use a model variant with a larger context window
  4. Count tokens client-side (e.g. a tokenizer) before calling generate and reject oversized requests early

Example fix

// before
resp = provider.generate(LLMInput(prompt=full_transcript))
// after
MAX_CHARS = 100_000
resp = provider.generate(LLMInput(prompt=full_transcript[-MAX_CHARS:]))
Defensive patterns

Strategy: try-catch

Validate before calling

def fits_context(text, max_model_tokens, max_output):
    approx_tokens = len(text) // 4  # rough estimate; use a tokenizer for accuracy
    return approx_tokens + max_output <= max_model_tokens

Try / catch

try:
    resp = provider.generate(inp)
except ContextLengthError as e:
    logger.warning("Prompt too long for model, truncating")
    inp = replace(inp, prompt=inp.prompt[-len(inp.prompt)//2:])
    resp = provider.generate(inp)

Prevention

When it happens

Trigger: Calling provider.generate() where the combined prompt + history + max_tokens exceeds the Claude model's context window; the SDK raises an error whose string form contains 'context ... length'.

Common situations: Feeding whole documents or long transcripts into the prompt; accumulating chat history without truncation; passing a large system prompt plus a large max_tokens; using a model with a smaller window than assumed (e.g. switching from a 200k model to a smaller one).

Related errors


AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16). Data as JSON: /api/errors/a73256fc049c06c1. Report an issue: GitHub.

Appendix: source

Thrown at src/llm/providers/claude.py:132

                model=response.model,
                usage={
                    "input_tokens": response.usage.input_tokens,
                    "output_tokens": response.usage.output_tokens,
                    "cache_creation_input_tokens": getattr(
                        response.usage, "cache_creation_input_tokens", 0
                    ),
                    "cache_read_input_tokens": getattr(response.usage, "cache_read_input_tokens", 0),
                },
                stop_reason=response.stop_reason,
            )
        except Exception as e:
            msg = str(e)
            if "401" in msg or "authentication" in msg.lower():
                raise AuthenticationError(msg, provider=ProviderType.CLAUDE) from e
            if "429" in msg or "rate_limit" in msg.lower():
                raise RateLimitError(msg, provider=ProviderType.CLAUDE) from e
            if "context" in msg.lower() and "length" in msg.lower():
                raise ContextLengthError(msg, provider=ProviderType.CLAUDE) from e
            raise

    def list_models(self) -> list[ModelInfo]:
        return self._models.copy()

    def validate_config(self) -> bool:
        return bool(self.client.api_key)

    def get_default_model(self) -> str:
        return _DEFAULT_MODEL

View on GitHub (pinned to 8321021c54)