santifer/career-ops · error · LockTimeoutError

pipeline lock timeout

Error message

pipeline lock timeout: ${lockDir} held > ${timeoutMs}ms

What it means

pipeline-lock.mjs's acquirePipelineLock throws this timeout error when the pipeline lock directory still cannot be acquired after the caller's hard wait ceiling: Date.now() exceeds the hard deadline computed from maxWaitMs, so instead of waiting forever on another process's lock it gives up. The ceiling is checked unconditionally at the top of every retry pass so it always bounds the wait.

Solutions

  1. Find and wait for (or kill) the process holding the lock; the message names the lock dir
  2. If the holder is dead, remove the stale lock directory manually once confirmed orphaned
  3. Increase maxWaitMs if concurrent runs are legitimate and you just need a longer wait
  4. Serialize your runs — don't start a second pipeline command until the first exits

Example fix

// before
await acquirePipelineLock(lockDir, { maxWaitMs: 1000 }); // too short under contention
// after
await acquirePipelineLock(lockDir, { maxWaitMs: 30000 });
Defensive patterns

Strategy: retry

Validate before calling

// before acquiring, check if lock exists and how old it is
try { const st = statSync(lockDir); console.log('lock held, age ms:', Date.now() - st.mtimeMs); } catch {}
await acquirePipelineLock(lockDir, { maxWaitMs: 30000 });

Type guard

const lockLooksStale = (st, maxAgeMs) => Date.now() - st.mtimeMs > maxAgeMs;

Try / catch

try {
  await acquirePipelineLock(lockDir, { maxWaitMs: 30000 });
} catch (e) {
  if (/pipeline lock timeout/.test(e.message)) {
    // inspect holder, wait or clean stale dir, then retry once
  } else throw e;
}

Prevention

When it happens

Trigger: Calling acquirePipelineLock (directly or via any pipeline/batch operation) while another process holds the lock dir and its holder deadline hasn't expired; running two concurrent pipeline workers; a crashed process leaving the lock dir in place within the staleness window.

Common situations: Launching a batch worker while an interactive scan/pipeline run is still going; an earlier killed process's lock not yet considered stale; NFS/filesystem where mkdir contention is slow.

Understand the failure class

Background: Request timed out: what client-side request timeouts mean across libraries (Request timed out, TIMED_OUT, APITimeoutError) — this error's family across 39 libraries.

Related errors


AI-assisted analysis of santifer/career-ops@aac998c7ed (2026-09-16). Data as JSON: /api/errors/0cd6c9adae001ff1. Report an issue: GitHub.

Appendix: source

Thrown at pipeline-lock.mjs:453

        err.message += `, last mkdir error=${lastContentionError.code} on ${lastContentionError.path ?? '?'}`;
      }
    } catch { /* diagnosis must never mask the timeout it describes */ }
    return err;
  };

  // A fresh install may not have data/ yet — plugins.mjs's cmdRun calls
  // appendToPipeline with no directory pre-creation, so create it here rather
  // than letting mkdirSync(lockDir) throw a raw ENOENT.
  mkdirSync(dirname(lockDir), { recursive: true });

  for (;;) {
    // The ceiling is tested FIRST, on its own, on every pass. Nesting it inside
    // the per-holder deadline made it conditional on a branch that may never
    // run: a caller whose maxWaitMs is below timeoutMs never reaches the inner
    // check, a re-arm pushes the next look a whole window away, and the reclaim
    // fast-path continues straight past both. A bound that holds only when
    // another bound happens to fire is not a bound.
    if (Date.now() > hardDeadline) throw buildTimeoutError(maxWaitMs);

    try {
      mkdirSync(lockDir);
    } catch (err) {
      if (!isMkdirContention(err)) throw err;
      lastContentionError = err;
      noteWaiting();

      // Serialize stale-reclaim behind a second atomic guard so only one
      // caller can be inside the decide-then-delete window at a time.
      let hasRecoverGuard = false;
      try {
        mkdirSync(recoverGuardDir);
        hasRecoverGuard = true;
      } catch (guardErr) {
        if (!isMkdirContention(guardErr)) throw guardErr;
        lastContentionError = guardErr;
        // An EPERM/EACCES here says the guard directory is mid-flight, not that

View on GitHub (pinned to aac998c7ed)