affaan-m/ECC · error · RateLimitError
RateLimitError(msg, provider=ProviderType.OLLAMA) from e
Error message
RateLimitError(msg, provider=ProviderType.OLLAMA) from e
What it means
Ollama's generate() wraps exceptions whose message contains '429' or 'rate_limit' in a RateLimitError. It signals the Ollama server (or an intermediary) rejected the request due to request throttling, with the provider type OLLAMA attached.
Solutions
- Back off and retry with exponential backoff on RateLimitError
- Serialize or limit concurrency of generate() calls to the Ollama instance
- Increase the rate limit on any intermediary gateway
- Check `ollama ps` for GPU/memory contention causing throttling
Example fix
// before
resp = provider.generate(inp)
// after
for attempt in range(5):
try:
resp = provider.generate(inp); break
except RateLimitError:
time.sleep(2 ** attempt) Defensive patterns
Strategy: retry
Try / catch
from tenacity import retry, wait_exponential, stop_after_attempt
@retry(retry=retry_if_exception_type(RateLimitError), wait=wait_exponential(min=1, max=30), stop=stop_after_attempt(5))
def safe_generate(inp):
return provider.generate(inp) Prevention
- Throttle concurrency to the Ollama instance
- Use a queue for bulk generation jobs
- Monitor 429s and back off proactively
- Reserve capacity or scale Ollama replicas under load
When it happens
Trigger: Sending concurrent generate() calls faster than the Ollama server's queue/limiter allows; a gateway in front of Ollama returning HTTP 429.
Common situations: Parallel worker pools hammering a single local Ollama instance; shared Ollama deployments with per-user quotas; API gateways (nginx, Kong) configured with rate limits.
Related errors
- RateLimitError(msg, provider=ProviderType.OPENAI) from e
- ContextLengthError(msg, provider=ProviderType.OLLAMA) from e
- Ollama connection failed
- request retry budget exhausted
AI-assisted analysis of affaan-m/ECC@8321021c54 (2026-09-16).
Data as JSON: /api/errors/9d8cb145fc6360fe.
Report an issue: GitHub.
Appendix: source
Thrown at src/llm/providers/ollama.py:106
id=tc.get("id", ""),
name=tc.get("function", {}).get("name", ""),
arguments=tc.get("function", {}).get("arguments", {}),
)
for tc in result["message"]["tool_calls"]
]
return LLMOutput(
content=content,
tool_calls=tool_calls,
model=model,
stop_reason=result.get("done_reason"),
)
except Exception as e:
msg = str(e)
if "401" in msg or "connection" in msg.lower():
raise AuthenticationError(f"Ollama connection failed: {msg}", provider=ProviderType.OLLAMA) from e
if "429" in msg or "rate_limit" in msg.lower():
raise RateLimitError(msg, provider=ProviderType.OLLAMA) from e
if "context" in msg.lower() and "length" in msg.lower():
raise ContextLengthError(msg, provider=ProviderType.OLLAMA) from e
raise
def list_models(self) -> list[ModelInfo]:
return self._models.copy()
def validate_config(self) -> bool:
return bool(self.base_url)
def get_default_model(self) -> str:
return self.default_model
View on GitHub (pinned to 8321021c54)