{"record":{"id":"eb81131efe9a20fb","repo":"santifer/career-ops","slug":"portal-health-lock-timeout-lockdir-held-ti","errorCode":null,"errorMessage":"portal-health lock timeout: ${lockDir} held > ${timeoutMs}ms","messagePattern":"portal-health lock timeout: (.+?) held > (.+?)ms","errorType":"exception","errorClass":"LockTimeoutError","httpStatus":null,"severity":"error","filePath":"portal-health-lock.mjs","lineNumber":158,"sourceCode":"    } catch (err) {\n      // Windows answers a mid-flight directory with EPERM/EACCES: contention,\n      // not failure. Treating it as fatal kills the writer and loses its write.\n      if (!isMkdirContention(err)) throw err;\n      noteWaiting();\n\n      // Serialize stale-reclaim behind a second atomic guard so only one\n      // caller can be inside the decide-then-delete window at a time.\n      let hasRecoverGuard = false;\n      try {\n        mkdirSync(recoverGuardDir);\n        hasRecoverGuard = true;\n      } catch (guardErr) {\n        if (!isMkdirContention(guardErr)) throw guardErr;\n        // Only an EEXIST guard is judged by age: an EPERM/EACCES answer means it\n        // is mid-flight, and judging the age of a directory we cannot stat\n        // reliably would evict a live guard.\n        if (guardErr.code !== 'EEXIST') {\n          if (holderStillWedged() || ceilingReached()) throw new LockTimeoutError(lockDir, timeoutMs);\n          await sleep(backoffMs());\n          continue;\n        }\n        // A process killed between taking the guard and cleaning it up would\n        // otherwise disable stale recovery forever. The guard normally lives\n        // for milliseconds, so an old one is judged by the same age rule.\n        // STALE only: a guard already gone needs no eviction, and evicting on\n        // that answer deletes the guard another caller has just taken.\n        if (lockRecoveryVerdict(recoverGuardDir, staleMs) === RECOVER_STALE) {\n          rmLockArtifactSync(recoverGuardDir);\n        }\n      }\n\n      if (hasRecoverGuard) {\n        try {\n          // STALE only. VANISHED means the lock was absent when we looked, and\n          // by the time this line runs another acquirer may have won the mkdir\n          // and be partway through writing owner.json — deleting on that answer","sourceCodeStart":140,"sourceCodeEnd":176,"githubUrl":"https://github.com/santifer/career-ops/blob/aac998c7ed7248ea853b720ceeb1fdbeb322fc5d/portal-health-lock.mjs#L140-L176","documentation":"acquirePortalHealthLock() implements a file-system lock (a lock directory plus an EEXIST guard) for portal-health runs. When the lock cannot be acquired within timeoutMs — because another process still holds the lockDir and is judged 'still wedged' or the retry ceiling is reached — it throws LockTimeoutError with this message. It protects concurrent health checks from corrupting shared state, so it indicates contention or a stale/abandoned lock, not a bug in your call.","triggerScenarios":"Calling acquirePortalHealthLock (directly or via a portal-health run) while another process holds the lockDir longer than timeoutMs and the guard directory is either mid-flight (EPERM/EACCES path) or the non-EEXIST guard branch determines the holder is wedged or the backoff ceiling has been reached.","commonSituations":"Two scan/health jobs overlapping (e.g. a cron run colliding with a manual run); a previous run was killed (SIGKILL, power loss, crashed container) leaving the lock directory behind while the age-out has not yet elapsed; a hung network fetch inside the lock holder keeping it alive far beyond the timeout; NFS/containers where the age heuristics misjudge liveness.","solutions":["Wait for the timeoutMs window and retry — the age-out will eventually recover a genuinely stale lock","Check for another running scan/portal-health process (ps / job list) and let it finish or stop it","If no process is running, manually remove the stale lock directory reported in the message (e.g. rm -rf <lockDir>) and re-run","Increase the timeoutMs option or reduce scheduler overlap (widen cron interval, add a jitter) if contention is routine"],"exampleFix":"// before\nawait acquirePortalHealthLock(); // default timeout, overlaps with cron run\n// after\nconst lock = await acquirePortalHealthLock({ timeoutMs: 60000 }); // tolerate slow holders\ntry { /* ... */ } finally { lock.release(); }","handlingStrategy":"retry","validationCode":"// Before starting a run, check the lock is not held\nimport { existsSync } from 'node:fs';\nif (existsSync(lockDir)) console.warn('portal-health lock present; another run may be active — skipping or waiting');","typeGuard":null,"tryCatchPattern":"try {\n  const lock = await acquirePortalHealthLock({ timeoutMs: 60000 });\n  try { await runPortalHealth(); } finally { lock.release(); }\n} catch (e) {\n  if (String(e.message).startsWith('portal-health lock timeout')) {\n    console.warn('Another portal-health run holds the lock; retrying later.');\n  } else throw e;\n}","preventionTips":["Schedule health/scan runs serially with enough spacing that one run cannot overlap the next","Always release the lock in a finally block so crashes mid-run do not wedge the next run","Prefer SIGTERM-friendly shutdown so the lock dir is cleaned up instead of abandoned","If timeouts recur without a visible holder, remove the stale lockDir manually and consider a longer timeoutMs"],"tags":["filesystem","lock","concurrency","timeout"],"backgroundTag":"lock-acquisition-timeout","analyzedSha":"aac998c7ed7248ea853b720ceeb1fdbeb322fc5d","analyzedAt":"2026-09-16T06:35:29.214Z","contentChangedAt":"2026-09-16T06:35:29.214Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}