thedotmack/claude-mem · critical · Error

Hard cap exceeded: processes in registry (cap= ). Refusing…

Error message

Hard cap exceeded: ${activeCount} processes in registry (cap=${TOTAL_PROCESS_HARD_CAP}). Refusing to spawn more.

What it means

waitForSlot enforces a global hard cap (TOTAL_PROCESS_HARD_CAP) on active SDK subprocesses in the registry. After pruning dead entries, if the active count is at or above the cap, it throws instead of queueing a waiter — the system refuses to create more processes under any circumstance. This is a last-resort resource guard above the soft per-pool maxConcurrent limit.

Solutions

  1. Reduce concurrent spawn pressure so active processes stay under the hard cap
  2. Ensure every SlotReservation is released (including on error paths) to avoid slot leaks
  3. Increase TOTAL_PROCESS_HARD_CAP only after confirming memory/CPU headroom
  4. Investigate stuck subprocesses that never exit and keep registry entries active

Example fix

// before
const slot = await waitForSlot(maxConcurrent); // throws at cap
// after
if (getActiveSdkCount() >= TOTAL_PROCESS_HARD_CAP) {
  await waitForActiveCountBelow(TOTAL_PROCESS_HARD_CAP);
}
const slot = await waitForSlot(maxConcurrent);
Defensive patterns

Strategy: retry

Validate before calling

const active = getActiveSdkCount();
if (active >= TOTAL_PROCESS_HARD_CAP) await backoffUntil(() => getActiveSdkCount() < TOTAL_PROCESS_HARD_CAP);

Try / catch

try { const slot = await waitForSlot(max); }
catch (e) {
  if (e.message.startsWith('Hard cap exceeded')) { await sleep(backoff); return waitForSlotWithRetry(max); }
  throw e;
}

Prevention

When it happens

Trigger: Requesting a slot reservation when the registry already holds TOTAL_PROCESS_HARD_CAP or more live SDK subprocess entries after pruning.

Common situations: Runaway spawn loops that never release slots; leaked subprocesses whose reservations were never released; load spikes beyond the configured cap; long queues combined with processes that hang and never exit.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of thedotmack/claude-mem@d8bc9755e7 (2026-09-17). Data as JSON: /api/errors/4e055a87a742d560. Report an issue: GitHub.

Appendix: source

Thrown at src/supervisor/process-registry.ts:571

 * that is re-read on every recheck — pass a thunk so a mid-wait settings
 * change (raising CLAUDE_MEM_MAX_CONCURRENT_AGENTS) can release an
 * already-parked waiter without a worker restart (#2756). `sessionId`, when
 * provided, lets callers (SessionRoutes) detect via isSessionParkedForSlot()
 * whether this specific session is currently parked here, to distinguish a
 * provider-switch onto a parked generator from one that is already
 * mid-response.
 */
export async function waitForSlot(
  maxConcurrent: number | (() => number),
  signal?: AbortSignal,
  sessionId?: number | string
): Promise<SlotReservation> {
  const getMax = typeof maxConcurrent === 'function' ? maxConcurrent : () => maxConcurrent;

  getProcessRegistry().pruneDeadEntries();
  const activeCount = getActiveSdkCount();
  if (activeCount >= TOTAL_PROCESS_HARD_CAP) {
    throw new Error(`Hard cap exceeded: ${activeCount} processes in registry (cap=${TOTAL_PROCESS_HARD_CAP}). Refusing to spawn more.`);
  }

  if (activeCount < getMax()) return takeSlotReservation();

  if (signal?.aborted) {
    throw new Error('waitForSlot aborted before queuing');
  }

  logger.info('PROCESS', `Pool limit reached (${activeCount}/${getMax()}), waiting for slot...`);

  return new Promise<SlotReservation>((resolve, reject) => {
    let recheckTimer: ReturnType<typeof setInterval> | null = null;
    let abortHandler: (() => void) | null = null;
    const record: SlotWaiterRecord = {
      sessionId,
      parkedSince: Date.now(),
      warnedParked: false,
      notify: () => {},

View on GitHub (pinned to d8bc9755e7)