abhigyanpatwari/GitNexus · error · ResilientFetchExhaustedError
Request failed after retries
Error message
Request failed after retries (HTTP ${response.status}) What it means
During idle-timeout retries, a slot's respawn counter exceeded poolOptions.maxRespawnsPerSlot, so the pool drops (retires) that slot instead of respawning its worker again. Available parallelism shrinks permanently for the run; if enough slots drop or the breaker also trips, dispatch degrades or fails. This is the budget path that stops infinite respawn loops on a persistently dying slot.
Solutions
- Find why workers die rather than raising the budget first: check for worker crash logs (exit codes, OOM killer, segfaults) preceding this warn
- If deaths are OOM, raise per-worker memory or shrink sub-batches (GITNEXUS_WORKER_SUB_BATCH_MAX_BYTES)
- Raise GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT only to ride through transient environments
- Update/rebuild the vendored tree-sitter grammars if a native crash is suspected (prebuild vs platform mismatch)
Defensive patterns
Strategy: try-catch
Type guard
const isWorkerPoolDispatchError = (e: unknown): e is Error & { quarantine?: unknown } =>
e instanceof Error && /Worker pool/.test(e.message); Try / catch
try {
await pool.dispatch(job);
} catch (e) {
if (isWorkerPoolDispatchError(e)) {
// slot dropped: parallelism is reduced but the pool may still drain remaining jobs;
// decide between degrading gracefully and aborting to re-run with a higher
// GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT after fixing worker deaths
logger.warn(`slot loss during dispatch: ${e.message}`);
}
} Prevention
- Diagnose why workers die (OOM killer, segfaults) before raising the respawn budget
- Set container/OOM limits above per-worker peak memory
- Rebuild vendored grammars after Node major-version upgrades to avoid native crashes
When it happens
Trigger: respawnCount[workerIndex] > poolOptions.maxRespawnsPerSlot after an idle-timeout retry decision — a slot whose workers keep dying and being respawned within one run; default budget is DEFAULT_MAX_RESPAWNS_PER_SLOT, overridable via options or GITNEXUS_WORKER_MAX_RESPAWNS_PER_SLOT.
Common situations: Workers OOM-crashing on a huge file faster than the budget refills; native addon (tree-sitter grammar) segfaulting on specific inputs; machines where worker threads are killed by an external supervisor (OOM killer, container limits).
Related errors
- Circuit ' ' is open; retry in s
- cpp range-binding: parse timed out, skipping file
- go range-binding: parse timed out, skipping file
- rust range-binding: parse timed out, skipping file
- Worker died; respawning slot (attempt / ).
AI-assisted analysis of abhigyanpatwari/GitNexus@aac7515d2a (2026-08-20).
Data as JSON: /api/errors/e57b5204a003724d.
Report an issue: GitHub.
Appendix: source
Thrown at gitnexus-shared/src/integrations/resilient-fetch.ts:248
// also do NOT call recordSuccess: a 401 sandwiched between
// 5xx responses would otherwise erase the running outage
// signal. The breaker's neutral path leaves state untouched.
breaker.recordNeutral();
return outcome.resp;
case 'terminal-network':
// Either `AbortSignal.timeout()` fired locally OR an external
// caller cancelled the request via AbortController. The server
// never had a chance to answer; this reflects the user's
// network or an explicit cancel, not registry health. Don't
// punish the breaker AND don't reset its outage signal.
breaker.recordNeutral();
throw outcome.err;
case 'retryable-status':
if (attempt + 1 >= retryConfig.maxAttempts) {
breaker.recordFailure();
throw new ResilientFetchExhaustedError(outcome.resp);
}
await sleep(
computeBackoffMs(
attempt,
retryConfig.baseDelayMs,
retryConfig.capDelayMs,
outcome.afterMs,
random,
),
);
break;
case 'retryable-network':
if (attempt + 1 >= retryConfig.maxAttempts) {
breaker.recordFailure();
throw outcome.err;
}
await sleep(View on GitHub (pinned to aac7515d2a)