thedotmack/claude-mem · error · Error
exceeded the ms per-attempt deadline. Raise…
Error message
${opts.label ?? 'Request'} exceeded the ${opts.perAttemptTimeoutMs}ms per-attempt deadline. Raise CLAUDE_MEM_LLM_TIMEOUT_MS if the backend is simply slow. What it means
withRetry's own per-attempt deadline (perAttemptTimeoutMs, configurable via CLAUDE_MEM_LLM_TIMEOUT_MS) fired and the abort was about to be misclassified as a transient network error and retried. The library deliberately converts it into a non-retryable error: retrying against an already-slow backend turns latency into congestion collapse, so it tells you to raise the deadline instead.
Solutions
- Raise CLAUDE_MEM_LLM_TIMEOUT_MS (e.g. from default to 120000–300000) and retry the operation
- Check backend health/load — if the backend is saturated, fix capacity rather than the timeout
- Reduce prompt size or batch size so attempts fit within the deadline
- If the error persists immediately at the same threshold, confirm CLAUDE_MEM_LLM_TIMEOUT_MS is actually reaching the worker process (env not stripped by your service manager)
Example fix
// before // worker started without timeout env; slow backend hits the 60s default // after // in the worker's environment: CLAUDE_MEM_LLM_TIMEOUT_MS=180000
Defensive patterns
Strategy: retry
Validate before calling
// before wrapping the call, ensure the env knob is set appropriately:
const timeoutMs = Number(process.env.CLAUDE_MEM_LLM_TIMEOUT_MS ?? 60000);
if (!(timeoutMs > 0)) throw new Error('CLAUDE_MEM_LLM_TIMEOUT_MS must be a positive number'); Try / catch
try {
return await withRetry(() => callLlm(prompt), { label: 'Summarize', perAttemptTimeoutMs });
} catch (err) {
if (err instanceof Error && err.message.includes('per-attempt deadline')) {
// non-retryable by design: raise CLAUDE_MEM_LLM_TIMEOUT_MS or shrink the prompt
logger.error('LLM attempt exceeded deadline — not retrying to avoid congestion collapse');
throw err;
}
throw err;
} Prevention
- Set CLAUDE_MEM_LLM_TIMEOUT_MS generously (2–5 min) for large summarization workloads
- Confirm the env var actually reaches the worker process (service managers often strip env)
- Watch backend latency/queue depth and scale capacity before deadlines fire
- Keep prompts/batches small enough that a single attempt fits the deadline with headroom
When it happens
Trigger: Any LLM call wrapped in withRetry (summarization, memory compression) whose single attempt exceeds perAttemptTimeoutMs — the abort error surfaces with no HTTP status, so it would otherwise look transient.
Common situations: Large summarization prompts on an overloaded or rate-limited backend; CLAUDE_MEM_LLM_TIMEOUT_MS left at its default for a slow local/proxy LLM; network path with high latency to the backend; model cold-start latency.
Understand the failure class
Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.
- Timeouts: ETIMEDOUT, deadlines, and hung requests — what actually expires when a request times out.
Related errors
- Auto-reprime failed for corpus
- BullMQ re-enqueue failed (will reconcile on startup)
- cannot process observation generation job after…
- cannot retry observation generation job after max_attempts…
- chroma-mcp connection in backoff
AI-assisted analysis of thedotmack/claude-mem@d8bc9755e7 (2026-09-17).
Data as JSON: /api/errors/483512c30ef3287b.
Report an issue: GitHub.
Appendix: source
Thrown at src/services/worker/retry.ts:146
const attemptController = new AbortController();
let deadlineExpired = false;
const timeoutHandle = setTimeout(() => {
deadlineExpired = true;
attemptController.abort();
}, opts.perAttemptTimeoutMs);
const onExternalAbort = () => attemptController.abort();
options.abortSignal?.addEventListener('abort', onExternalAbort, { once: true });
try {
return await fn(attemptController.signal);
} catch (err: unknown) {
lastError = err;
// Our own deadline, not a network blip. The abort surfaces with no HTTP // status, so it classifies as transient and was retried twice — against a
// backend that is already saturated, those attempts are what turn a
// latency problem into a congestion collapse. Raise the deadline instead.
if (deadlineExpired) {
throw new Error(
`${opts.label ?? 'Request'} exceeded the ${opts.perAttemptTimeoutMs}ms per-attempt deadline. `
+ 'Raise CLAUDE_MEM_LLM_TIMEOUT_MS if the backend is simply slow.',
{ cause: err },
);
}
if (!isRetryableKind(err)) {
throw err;
}
if (attempt === opts.maxRetries) {
throw err;
}
// Honor retryAfterMs from rate_limit errors; otherwise exponential backoff.
let delayMs: number;
if (isClassified(err) && err.kind === 'rate_limit' && err.retryAfterMs !== undefined) {View on GitHub (pinned to d8bc9755e7)