abhigyanpatwari/GitNexus · critical · AggregateError

Workload-lock cleanup failed

Error message

Workload-lock cleanup failed: ${lockPath}. Acquisition refused; see RUNBOOK.md for quiesced recovery.

What it means

This AggregateError is thrown by acquireViaFile when a non-permission error occurred while trying to claim the workload lock while the guard was held, and the attempt's own token-exact analyze.lock record then failed to unlink during rollback. The acquisition is refused (no handle is returned) but stale lock state may remain under .gitnexus/, so recovery requires quiescing writers and following RUNBOOK.md. Both the original error and the cleanup error are surfaced via AggregateError.errors.

Solutions

  1. Quiesce all writers, inspect both errors in AggregateError.errors, fix the underlying filesystem fault (space, permissions, writability), then follow RUNBOOK.md to clear leftover analyze.lock/guard and retry.
  2. Verify .gitnexus/ is writable by the running user (ls -ld, touch a test file) and that the mount is not read-only (mount | grep ro).
  3. Free disk space if ENOSPC appears in the wrapped errors.
  4. Exclude .gitnexus/ from sync/AV tools and move off NFS if unlink races persist.
  5. As a last resort with no writers active, remove the whole .gitnexus/ slot and re-run analyze to rebuild.

Example fix

// before (mid-run read-only remount leaves the lock)
$ npx gitnexus analyze
# AggregateError: Workload-lock cleanup failed: .gitnexus/analyze.lock ...

// after (check writability first, then retry)
$ mount | grep $(df .gitnexus --output=source | tail -1)   # ensure rw
$ touch .gitnexus/.writetest && rm .gitnexus/.writetest
$ npx gitnexus analyze
Defensive patterns

Strategy: try-catch

Validate before calling

// Pre-flight: writable dir, free space, rw mount
import { accessSync, constants, statfsSync } from 'node:fs';
accessSync('.gitnexus', constants.W_OK);
const { bavail, bsize } = statfsSync('.gitnexus');
if (bavail * bsize < 10 * 1024 * 1024) throw new Error('.gitnexus nearly full');

Type guard

const isWorkloadCleanupAggregate = (e: unknown): e is AggregateError & { errors: [unknown, unknown] } =>
  e instanceof AggregateError &&
  typeof e.message === 'string' &&
  e.message.startsWith('Workload-lock cleanup failed');

Try / catch

try {
  await acquireIndexLock(lockDir);
} catch (err) {
  if (isWorkloadCleanupAggregate(err)) {
    const [error, cleanupError] = err.errors;
    console.error('acquisition refused; quiesce writers, fix fs fault, see RUNBOOK.md', error, cleanupError);
    return; // never write without a verified lock
  }
  throw err;
}

Prevention

When it happens

Trigger: In the catch block of the guarded section: an error other than a retryable EPERM (e.g. guard verification failure, write failure) with createdMain=true, and the rollback readRecord/unlinkSync on analyze.lock throws (EACCES, ENOENT-then-race, EIO, read-only remount).

Common situations: Disk filling up or filesystem going read-only mid-acquisition; permissions on .gitnexus/ changed while analyze ran; concurrent deletion of the lock file by a tool or another user; flaky network mounts dropping unlinks; interrupted mounts or dying disks.

Related errors


AI-assisted analysis of abhigyanpatwari/GitNexus@ac9a4e9abd (2026-09-15). Data as JSON: /api/errors/9fb2a9a1a1eaaf41. Report an issue: GitHub.

Appendix: source

Thrown at gitnexus/src/storage/index-lock.ts:669

      }
      permissionWaitSince = null;
      permissionError = undefined;
    } catch (error) {
      if (!createdMain && (error as NodeJS.ErrnoException).code === 'EPERM') {
        // A releasing owner may leave the main file delete-pending on Windows.
        // Release our guard in finally and retry; never reclaim an unreadable file.
        permissionWaitSince ??= Date.now();
        permissionError = error;
        holder = null;
        if (Date.now() >= Math.min(startedAt + timeoutMs, permissionWaitSince + GUARD_TIMEOUT_MS)) {
          throw error;
        }
      } else {
        if (createdMain) {
          try {
            if (readRecord(lockPath)?.token === me.token) unlinkSync(lockPath);
          } catch (cleanupError) {
            throw new AggregateError(
              [error, cleanupError],
              `Workload-lock cleanup failed: ${lockPath}. Acquisition refused; see RUNBOOK.md for quiesced recovery.`,
            );
          }
        }
        throw error;
      }
    } finally {
      releaseAcquisitionGuard(guardPath, me, createdMain, lockPath);
    }
    const waited = Date.now() - startedAt;

    if (holder) {
      // Live holder → wait.
      lastLiveHolder = holder;
      if (!announcedWait) {
        announcedWait = true;
        opts.onWaitStart?.(holder);

View on GitHub (pinned to ac9a4e9abd)