santifer/career-ops · error · LockTimeoutError
portal-health lock timeout
Error message
portal-health lock timeout: ${lockDir} held > ${timeoutMs}ms What it means
acquirePortalHealthLock() implements a file-system lock (a lock directory plus an EEXIST guard) for portal-health runs. When the lock cannot be acquired within timeoutMs — because another process still holds the lockDir and is judged 'still wedged' or the retry ceiling is reached — it throws LockTimeoutError with this message. It protects concurrent health checks from corrupting shared state, so it indicates contention or a stale/abandoned lock, not a bug in your call.
Solutions
- Wait for the timeoutMs window and retry — the age-out will eventually recover a genuinely stale lock
- Check for another running scan/portal-health process (ps / job list) and let it finish or stop it
- If no process is running, manually remove the stale lock directory reported in the message (e.g. rm -rf <lockDir>) and re-run
- Increase the timeoutMs option or reduce scheduler overlap (widen cron interval, add a jitter) if contention is routine
Example fix
// before
await acquirePortalHealthLock(); // default timeout, overlaps with cron run
// after
const lock = await acquirePortalHealthLock({ timeoutMs: 60000 }); // tolerate slow holders
try { /* ... */ } finally { lock.release(); } Defensive patterns
Strategy: retry
Validate before calling
// Before starting a run, check the lock is not held
import { existsSync } from 'node:fs';
if (existsSync(lockDir)) console.warn('portal-health lock present; another run may be active — skipping or waiting'); Try / catch
try {
const lock = await acquirePortalHealthLock({ timeoutMs: 60000 });
try { await runPortalHealth(); } finally { lock.release(); }
} catch (e) {
if (String(e.message).startsWith('portal-health lock timeout')) {
console.warn('Another portal-health run holds the lock; retrying later.');
} else throw e;
} Prevention
- Schedule health/scan runs serially with enough spacing that one run cannot overlap the next
- Always release the lock in a finally block so crashes mid-run do not wedge the next run
- Prefer SIGTERM-friendly shutdown so the lock dir is cleaned up instead of abandoned
- If timeouts recur without a visible holder, remove the stale lockDir manually and consider a longer timeoutMs
When it happens
Trigger: Calling acquirePortalHealthLock (directly or via a portal-health run) while another process holds the lockDir longer than timeoutMs and the guard directory is either mid-flight (EPERM/EACCES path) or the non-EEXIST guard branch determines the holder is wedged or the backoff ceiling has been reached.
Common situations: Two scan/health jobs overlapping (e.g. a cron run colliding with a manual run); a previous run was killed (SIGKILL, power loss, crashed container) leaving the lock directory behind while the age-out has not yet elapsed; a hung network fetch inside the lock holder keeping it alive far beyond the timeout; NFS/containers where the age heuristics misjudge liveness.
Understand the failure class
- Timeouts: ETIMEDOUT, deadlines, and hung requests — what actually expires when a request times out.
Related errors
- LOCK_TIMEOUT
- ⚠️ Could not release report reservation
- pipeline lock timeout
- apify: invalid timeoutMs
- apify: JD cache write failed for
AI-assisted analysis of santifer/career-ops@aac998c7ed (2026-09-16).
Data as JSON: /api/errors/eb81131efe9a20fb.
Report an issue: GitHub.
Appendix: source
Thrown at portal-health-lock.mjs:158
} catch (err) {
// Windows answers a mid-flight directory with EPERM/EACCES: contention,
// not failure. Treating it as fatal kills the writer and loses its write.
if (!isMkdirContention(err)) throw err;
noteWaiting();
// Serialize stale-reclaim behind a second atomic guard so only one
// caller can be inside the decide-then-delete window at a time.
let hasRecoverGuard = false;
try {
mkdirSync(recoverGuardDir);
hasRecoverGuard = true;
} catch (guardErr) {
if (!isMkdirContention(guardErr)) throw guardErr;
// Only an EEXIST guard is judged by age: an EPERM/EACCES answer means it
// is mid-flight, and judging the age of a directory we cannot stat
// reliably would evict a live guard.
if (guardErr.code !== 'EEXIST') {
if (holderStillWedged() || ceilingReached()) throw new LockTimeoutError(lockDir, timeoutMs);
await sleep(backoffMs());
continue;
}
// A process killed between taking the guard and cleaning it up would
// otherwise disable stale recovery forever. The guard normally lives
// for milliseconds, so an old one is judged by the same age rule.
// STALE only: a guard already gone needs no eviction, and evicting on
// that answer deletes the guard another caller has just taken.
if (lockRecoveryVerdict(recoverGuardDir, staleMs) === RECOVER_STALE) {
rmLockArtifactSync(recoverGuardDir);
}
}
if (hasRecoverGuard) {
try {
// STALE only. VANISHED means the lock was absent when we looked, and
// by the time this line runs another acquirer may have won the mkdir
// and be partway through writing owner.json — deleting on that answerView on GitHub (pinned to aac998c7ed)