affaan-m/ECC · error · ContextLengthError

ContextLengthError(msg, provider=ProviderType.OLLAMA) from e

Error message

ContextLengthError(msg, provider=ProviderType.OLLAMA) from e

What it means

Ollama's generate() raises ContextLengthError when the underlying error message contains both 'context' and 'length', meaning the prompt (plus num_predict/num_ctx) exceeded Ollama's context size. The original exception is chained via 'from e'.

Solutions

  1. Increase the context size (num_ctx option) when creating/loading the model
  2. Chunk or summarize long inputs before sending
  3. Truncate chat history to fit within num_ctx
  4. Choose a model/variant with a larger supported context window

Example fix

// before
# default num_ctx too small for long prompt
resp = provider.generate(inp)
// after
ollama run llama3 --num-ctx 16384
resp = provider.generate(inp)
Defensive patterns

Strategy: validation

Validate before calling

NUM_CTX = 16384
if len(prompt) // 4 + num_predict >= NUM_CTX:
    raise ValueError("Prompt would exceed Ollama context window")

Try / catch

try:
    resp = provider.generate(inp)
except ContextLengthError:
    inp = replace(inp, prompt=truncate_to_tokens(inp.prompt, NUM_CTX - num_predict))
    resp = provider.generate(inp)

Prevention

When it happens

Trigger: Calling generate() with a prompt larger than the model's configured num_ctx; Ollama returning an error mentioning 'context length'.

Common situations: Loading long documents into prompts without chunking; default num_ctx (often 2048/4096) being far smaller than the model's advertised window; long tool-call conversations accumulating tokens.

Related errors


AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16). Data as JSON: /api/errors/39ed0e3227dcfac3. Report an issue: GitHub.

Appendix: source

Thrown at src/llm/providers/ollama.py:108

                        arguments=tc.get("function", {}).get("arguments", {}),
                    )
                    for tc in result["message"]["tool_calls"]
                ]

            return LLMOutput(
                content=content,
                tool_calls=tool_calls,
                model=model,
                stop_reason=result.get("done_reason"),
            )
        except Exception as e:
            msg = str(e)
            if "401" in msg or "connection" in msg.lower():
                raise AuthenticationError(f"Ollama connection failed: {msg}", provider=ProviderType.OLLAMA) from e
            if "429" in msg or "rate_limit" in msg.lower():
                raise RateLimitError(msg, provider=ProviderType.OLLAMA) from e
            if "context" in msg.lower() and "length" in msg.lower():
                raise ContextLengthError(msg, provider=ProviderType.OLLAMA) from e
            raise

    def list_models(self) -> list[ModelInfo]:
        return self._models.copy()

    def validate_config(self) -> bool:
        return bool(self.base_url)

    def get_default_model(self) -> str:
        return self.default_model

View on GitHub (pinned to 8321021c54)