santifer/career-ops · error · LockTimeoutError

portal-health lock timeout

Error message

portal-health lock timeout: ${lockDir} held > ${timeoutMs}ms

What it means

acquirePortalHealthLock() implements a file-system lock (a lock directory plus an EEXIST guard) for portal-health runs. When the lock cannot be acquired within timeoutMs — because another process still holds the lockDir and is judged 'still wedged' or the retry ceiling is reached — it throws LockTimeoutError with this message. It protects concurrent health checks from corrupting shared state, so it indicates contention or a stale/abandoned lock, not a bug in your call.

Solutions

  1. Wait for the timeoutMs window and retry — the age-out will eventually recover a genuinely stale lock
  2. Check for another running scan/portal-health process (ps / job list) and let it finish or stop it
  3. If no process is running, manually remove the stale lock directory reported in the message (e.g. rm -rf <lockDir>) and re-run
  4. Increase the timeoutMs option or reduce scheduler overlap (widen cron interval, add a jitter) if contention is routine

Example fix

// before
await acquirePortalHealthLock(); // default timeout, overlaps with cron run
// after
const lock = await acquirePortalHealthLock({ timeoutMs: 60000 }); // tolerate slow holders
try { /* ... */ } finally { lock.release(); }
Defensive patterns

Strategy: retry

Validate before calling

// Before starting a run, check the lock is not held
import { existsSync } from 'node:fs';
if (existsSync(lockDir)) console.warn('portal-health lock present; another run may be active — skipping or waiting');

Try / catch

try {
  const lock = await acquirePortalHealthLock({ timeoutMs: 60000 });
  try { await runPortalHealth(); } finally { lock.release(); }
} catch (e) {
  if (String(e.message).startsWith('portal-health lock timeout')) {
    console.warn('Another portal-health run holds the lock; retrying later.');
  } else throw e;
}

Prevention

When it happens

Trigger: Calling acquirePortalHealthLock (directly or via a portal-health run) while another process holds the lockDir longer than timeoutMs and the guard directory is either mid-flight (EPERM/EACCES path) or the non-EEXIST guard branch determines the holder is wedged or the backoff ceiling has been reached.

Common situations: Two scan/health jobs overlapping (e.g. a cron run colliding with a manual run); a previous run was killed (SIGKILL, power loss, crashed container) leaving the lock directory behind while the age-out has not yet elapsed; a hung network fetch inside the lock holder keeping it alive far beyond the timeout; NFS/containers where the age heuristics misjudge liveness.

Understand the failure class

Related errors


AI-assisted analysis of santifer/career-ops@aac998c7ed (2026-09-16). Data as JSON: /api/errors/eb81131efe9a20fb. Report an issue: GitHub.

Appendix: source

Thrown at portal-health-lock.mjs:158

    } catch (err) {
      // Windows answers a mid-flight directory with EPERM/EACCES: contention,
      // not failure. Treating it as fatal kills the writer and loses its write.
      if (!isMkdirContention(err)) throw err;
      noteWaiting();

      // Serialize stale-reclaim behind a second atomic guard so only one
      // caller can be inside the decide-then-delete window at a time.
      let hasRecoverGuard = false;
      try {
        mkdirSync(recoverGuardDir);
        hasRecoverGuard = true;
      } catch (guardErr) {
        if (!isMkdirContention(guardErr)) throw guardErr;
        // Only an EEXIST guard is judged by age: an EPERM/EACCES answer means it
        // is mid-flight, and judging the age of a directory we cannot stat
        // reliably would evict a live guard.
        if (guardErr.code !== 'EEXIST') {
          if (holderStillWedged() || ceilingReached()) throw new LockTimeoutError(lockDir, timeoutMs);
          await sleep(backoffMs());
          continue;
        }
        // A process killed between taking the guard and cleaning it up would
        // otherwise disable stale recovery forever. The guard normally lives
        // for milliseconds, so an old one is judged by the same age rule.
        // STALE only: a guard already gone needs no eviction, and evicting on
        // that answer deletes the guard another caller has just taken.
        if (lockRecoveryVerdict(recoverGuardDir, staleMs) === RECOVER_STALE) {
          rmLockArtifactSync(recoverGuardDir);
        }
      }

      if (hasRecoverGuard) {
        try {
          // STALE only. VANISHED means the lock was absent when we looked, and
          // by the time this line runs another acquirer may have won the mkdir
          // and be partway through writing owner.json — deleting on that answer

View on GitHub (pinned to aac998c7ed)