abhigyanpatwari/GitNexus · error · ResilientFetchExhaustedError

Request failed after retries

Error message

Request failed after retries (HTTP ${response.status})

What it means

During idle-timeout retries, a slot's respawn counter exceeded poolOptions.maxRespawnsPerSlot, so the pool drops (retires) that slot instead of respawning its worker again. Available parallelism shrinks permanently for the run; if enough slots drop or the breaker also trips, dispatch degrades or fails. This is the budget path that stops infinite respawn loops on a persistently dying slot.

Solutions

  1. Find why workers die rather than raising the budget first: check for worker crash logs (exit codes, OOM killer, segfaults) preceding this warn
  2. If deaths are OOM, raise per-worker memory or shrink sub-batches (GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES)
  3. Raise GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT only to ride through transient environments
  4. Update/rebuild the vendored tree-sitter grammars if a native crash is suspected (prebuild vs platform mismatch)
Defensive patterns

Strategy: try-catch

Type guard

const isWorkerPoolDispatchError = (e: unknown): e is Error & { quarantine?: unknown } =>
  e instanceof Error && /Worker pool/.test(e.message);

Try / catch

try {
  await pool.dispatch(job);
} catch (e) {
  if (isWorkerPoolDispatchError(e)) {
    // slot dropped: parallelism is reduced but the pool may still drain remaining jobs;
    // decide between degrading gracefully and aborting to re-run with a higher
    // GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT after fixing worker deaths
    logger.warn(`slot loss during dispatch: ${e.message}`);
  }
}

Prevention

When it happens

Trigger: respawnCount[workerIndex] > poolOptions.maxRespawnsPerSlot after an idle-timeout retry decision — a slot whose workers keep dying and being respawned within one run; default budget is DEFAULT_MAX_RESPAWNS_PER_SLOT, overridable via options or GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT.

Common situations: Workers OOM-crashing on a huge file faster than the budget refills; native addon (tree-sitter grammar) segfaulting on specific inputs; machines where worker threads are killed by an external supervisor (OOM killer, container limits).

Related errors


AI-assisted analysis of abhigyanpatwari/GitNexus@aac7515d2a (2026-08-20). Data as JSON: /api/errors/e57b5204a003724d. Report an issue: GitHub.

Appendix: source

Thrown at gitnexus-shared/src/integrations/resilient-fetch.ts:248

        // also do NOT call recordSuccess: a 401 sandwiched between
        // 5xx responses would otherwise erase the running outage
        // signal. The breaker's neutral path leaves state untouched.
        breaker.recordNeutral();
        return outcome.resp;

      case 'terminal-network':
        // Either `AbortSignal.timeout()` fired locally OR an external
        // caller cancelled the request via AbortController. The server
        // never had a chance to answer; this reflects the user's
        // network or an explicit cancel, not registry health. Don't
        // punish the breaker AND don't reset its outage signal.
        breaker.recordNeutral();
        throw outcome.err;

      case 'retryable-status':
        if (attempt + 1 >= retryConfig.maxAttempts) {
          breaker.recordFailure();
          throw new ResilientFetchExhaustedError(outcome.resp);
        }
        await sleep(
          computeBackoffMs(
            attempt,
            retryConfig.baseDelayMs,
            retryConfig.capDelayMs,
            outcome.afterMs,
            random,
          ),
        );
        break;

      case 'retryable-network':
        if (attempt + 1 >= retryConfig.maxAttempts) {
          breaker.recordFailure();
          throw outcome.err;
        }
        await sleep(

View on GitHub (pinned to aac7515d2a)