santifer/career-ops · warning
concurrent reservation test flaked
Error message
concurrent reservation test flaked (${e.message}). Retrying once... What it means
test-all.mjs stress-tests reserve-report-num.mjs by spawning concurrent reservations and asserting the returned ranges never overlap. Because concurrent file-based allocation can transiently collide, the test wraps the assertion in a retry loop: the first failure logs this warning and retries once; only a second consecutive failure is recorded as a real test failure.
Solutions
- If it flakes once and passes on retry, ignore it — that is the designed behavior for timing-sensitive concurrency.
- If it fails consistently, debug reserve-report-num.mjs's atomic claim path (sentinels, locks, GC of stale sentinels) with the overlapping ranges from the error message.
- Reduce external load or run the concurrency tests on a local filesystem rather than a network mount.
- Increase retries locally while bisecting, but do not paper over a deterministic overlap.
Example fix
// before (diagnosis step) node reserve-report-num.mjs --count 4 & node reserve-report-num.mjs --count 4 & wait // after (compare the two printed ranges; overlap => bug in claim logic, not a flake)
Defensive patterns
Strategy: retry
Validate before calling
// pre-flight: ensure the allocator works serially before stressing it
if (run('node', ['reserve-report-num.mjs', '--count', '1']) === null) fail('allocator broken; skip concurrency stress'); Try / catch
try { assertNoOverlap(rangeX, rangeY); } catch (e) { if (retries-- > 0) warn(`flaked (${e.message}), retrying`); else fail(e.message); } Prevention
- Use local (non-NFS) filesystems for lock-sensitive tests
- Reserve report numbers immediately before spawning parallel workers
- Keep CI runners from sharing one working tree across shards
- Expect one flake-and-retry on loaded machines; investigate only repeat failures
When it happens
Trigger: Two concurrent reservation workers return ranges X and Y sharing an overlap, or another assertion inside the try throws — on the first attempt while reserveRetries > 0, producing this warning; a second failure calls fail().
Common situations: Heavily loaded CI runners slowing atomic claim/lock timing, filesystems with weak locking semantics (some network mounts), or a genuine regression in reserve-report-num.mjs's claim-and-release logic (would fail both attempts).
Related errors
- merge-tracker concurrent write test flaked
- Skipped # : status changed from to during review
- archive render skipped — no Playwright browser in env
- Cannot import generate-pdf.mjs
- Cannot verify tracker lock ownership at
AI-assisted analysis of santifer/career-ops@e7abd431fc (2026-09-16).
Data as JSON: /api/errors/79d54cf5d3f2a7d7.
Report an issue: GitHub.
Appendix: source
Thrown at test-all.mjs:10152
let stdout = '';
child.stdout.on('data', chunk => { stdout += chunk; });
child.on('close', () => resolve(stdout.trim()));
});
const [rangeX, rangeY] = await Promise.all([spawnReserve(), spawnReserve()]);
const toNums = r => {
const [s, e] = r.split('-').map(Number);
return Array.from({ length: e - s + 1 }, (_, i) => s + i);
};
const overlap = toNums(rangeX).filter(n => toNums(rangeY).includes(n));
if (rangeX && rangeY && overlap.length === 0) {
pass(`concurrent --count 4 reservations are disjoint (${rangeX} vs ${rangeY})`);
} else {
throw new Error(`concurrent ranges overlap: ${rangeX} vs ${rangeY} share [${overlap}]`);
}
break;
} catch (e) {
if (reserveRetries > 0) {
warn(`concurrent reservation test flaked (${e.message}). Retrying once...`);
reserveRetries -= 1;
} else {
fail(`concurrent reservation test failed: ${e.message}`);
break;
}
} finally {
rmSync(concTmp, { recursive: true, force: true });
}
}
// --release with a range deletes every sentinel in it.
const reserveRunFail = (args, dir) => {
try {
execFileSync(NODE, [RESERVE, ...args], {
encoding: 'utf-8',
stdio: ['pipe', 'pipe', 'pipe'],
env: { ...process.env, CAREER_OPS_REPORTS_DIR: dir, CAREER_OPS_TRACKER: join(dir, 'applications.md') },
});View on GitHub (pinned to e7abd431fc)