abhigyanpatwari/GitNexus · warning · CircuitOpenError

Circuit ' ' is open; retry in s

Error message

Circuit '${key}' is open; retry in ${Math.ceil(retryAfterMs / 1000)}s

What it means

The worker pool's per-job idle timer fired, but the stall tracker shows the main thread was stalled (typically a GC pause near the heap limit) for at least STALL_CREDIT_FRACTION (0.5) of the job's timeout window. The worker's progress postMessages were starved by the blocked event loop, so the worker looked idle while actually healthy. The pool credits this once per job (stallCreditUsed) and re-arms the idle timer instead of requeueing or retiring the worker.

Solutions

  1. Raise the Node heap for the analyze run: NODE_OPTIONS="--max-old-space-size=8192" npx gitnexus analyze
  2. Reduce worker parallelism so the main thread has headroom to drain progress messages (fewer workers, or run on a less loaded machine)
  3. Lower GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES so main-thread reduction batches stay small
  4. If it fires repeatedly on the same job, capture heap snapshots / run with --inspect to find the allocation source on the main thread

Example fix

# before
npx gitnexus analyze

# after (give the main thread heap headroom)
NODE_OPTIONS="--max-old-space-size=8192" npx gitnexus analyze
Defensive patterns

Strategy: retry

Validate before calling

// Before dispatching a heavy ingestion job, confirm main-thread heap headroom
const mu = process.memoryUsage();
if (mu.heapUsed / mu.heapTotal > 0.85) {
  // shed load: fewer workers, smaller sub-batches — GC stalls starve worker progress
  process.env.GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES = String(2 * 1024 * 1024);
}

Prevention

When it happens

Trigger: Idle timeout (subBatchIdleTimeoutMs) fires on a job while stallTracker.read() - stallAtArm >= 0.5 * job.timeoutMs and the one-per-job stall credit is unused — classically a large relationship-batch reduce on the main thread driving GC while N workers stream progress messages that never get processed.

Common situations: Indexing very large repos on memory-constrained CI runners with the default Node heap; heap pressure from concurrent graph writes; machines under heavy CPU contention that stretch GC pauses past half the idle-timeout window. This warn is often the precursor to an OOM crash of the main process.

Related errors


AI-assisted analysis of abhigyanpatwari/GitNexus@aac7515d2a (2026-08-20). Data as JSON: /api/errors/6fd7df14c3dd39a4. Report an issue: GitHub.

Appendix: source

Thrown at gitnexus-shared/src/integrations/circuit-breaker.ts:156

   *      remaining cooldown.
   *   2. Open with cooldown elapsed AND a probe is already in flight
   *      (race: another caller transitioned to half-open and grabbed
   *      the permit on a microtask before us) → throws with
   *      `halfOpenRetryAfterMs`.
   *   3. Half-Open with probe in flight → throws with `halfOpenRetryAfterMs`.
   *
   * **Pairing invariant**: every successful return from `check()` MUST
   * be paired with exactly one `recordSuccess` / `recordFailure` /
   * `recordNeutral` on every code path including thrown exceptions.
   * Failing to pair leaves the probe permit consumed forever and
   * wedges the breaker. See file-header JSDoc for the canonical
   * try/finally pattern.
   */
  check(): void {
    if (this.state === 'open' && this.openedAt !== null) {
      const elapsed = this.now() - this.openedAt;
      if (elapsed < this.cooldownMs) {
        throw new CircuitOpenError(this.cooldownMs - elapsed, this.key);
      }
      // Cooldown elapsed — transition to Half-Open. The very next
      // `probeInFlight` check below decides whether THIS caller gets
      // the permit or hits the gate.
      this.state = 'half-open';
    }

    if (this.state === 'half-open') {
      if (this.probeInFlight) {
        throw new CircuitOpenError(this.halfOpenRetryAfterMs, this.key);
      }
      this.probeInFlight = true;
    }
    // Closed state falls through silently.
  }

  recordSuccess(): void {
    this.probeInFlight = false;

View on GitHub (pinned to aac7515d2a)