affaan-m/ECC · error · RateLimitError

RateLimitError(msg, provider=ProviderType.OLLAMA) from e

Error message

RateLimitError(msg, provider=ProviderType.OLLAMA) from e

What it means

Ollama's generate() wraps exceptions whose message contains '429' or 'rate_limit' in a RateLimitError. It signals the Ollama server (or an intermediary) rejected the request due to request throttling, with the provider type OLLAMA attached.

Solutions

  1. Back off and retry with exponential backoff on RateLimitError
  2. Serialize or limit concurrency of generate() calls to the Ollama instance
  3. Increase the rate limit on any intermediary gateway
  4. Check `ollama ps` for GPU/memory contention causing throttling

Example fix

// before
resp = provider.generate(inp)
// after
for attempt in range(5):
    try:
        resp = provider.generate(inp); break
    except RateLimitError:
        time.sleep(2 ** attempt)
Defensive patterns

Strategy: retry

Try / catch

from tenacity import retry, wait_exponential, stop_after_attempt
@retry(retry=retry_if_exception_type(RateLimitError), wait=wait_exponential(min=1, max=30), stop=stop_after_attempt(5))
def safe_generate(inp):
    return provider.generate(inp)

Prevention

When it happens

Trigger: Sending concurrent generate() calls faster than the Ollama server's queue/limiter allows; a gateway in front of Ollama returning HTTP 429.

Common situations: Parallel worker pools hammering a single local Ollama instance; shared Ollama deployments with per-user quotas; API gateways (nginx, Kong) configured with rate limits.

Related errors


AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16). Data as JSON: /api/errors/9d8cb145fc6360fe. Report an issue: GitHub.

Appendix: source

Thrown at src/llm/providers/ollama.py:106

                        id=tc.get("id", ""),
                        name=tc.get("function", {}).get("name", ""),
                        arguments=tc.get("function", {}).get("arguments", {}),
                    )
                    for tc in result["message"]["tool_calls"]
                ]

            return LLMOutput(
                content=content,
                tool_calls=tool_calls,
                model=model,
                stop_reason=result.get("done_reason"),
            )
        except Exception as e:
            msg = str(e)
            if "401" in msg or "connection" in msg.lower():
                raise AuthenticationError(f"Ollama connection failed: {msg}", provider=ProviderType.OLLAMA) from e
            if "429" in msg or "rate_limit" in msg.lower():
                raise RateLimitError(msg, provider=ProviderType.OLLAMA) from e
            if "context" in msg.lower() and "length" in msg.lower():
                raise ContextLengthError(msg, provider=ProviderType.OLLAMA) from e
            raise

    def list_models(self) -> list[ModelInfo]:
        return self._models.copy()

    def validate_config(self) -> bool:
        return bool(self.base_url)

    def get_default_model(self) -> str:
        return self.default_model

View on GitHub (pinned to 8321021c54)