ruvnet/ruflo · warning
swarm state is busy; retry the outcome update
Error message
swarm state is busy; retry the outcome update
What it means
Swarm store mutations that feed APSC pheromone signals (recordSwarmPheromoneSignal, shared with hooks post-task outcome updates) run under an exclusive advisory file lock (swarm-state.lock, opened with 'wx'). The acquisition loop tries 100 times with ~10ms waits (~1s total) and force-breaks locks older than SWARM_LOCK_STALE_MS = 10s. If the lock is still held (continuously refreshed by another live writer, or heavy contention), descriptor stays undefined and this error is thrown — by design, so the caller retries rather than corrupting swarm state.
Solutions
- Retry the outcome update — the message says exactly that; contention usually clears within a second.
- Retry with exponential backoff (e.g. 100ms, 200ms, 400ms) so bursts of parallel post-task hooks don't fight the lock.
- Ensure only ONE MCP server process owns the swarm state directory; kill duplicate daemons (tman list / ps) writing to the same home.
- If it persists >10s, the stale-lock breaker will unlink it on a later attempt; verify no zombie process still refreshes the lock's mtime.
Example fix
// before
const r = recordSwarmPheromoneSignal(signal); // throws 'busy' under load
// after
async function postOutcome(signal, tries = 5) {
for (let i = 0; ; i++) {
try { return recordSwarmPheromoneSignal(signal); }
catch (e) {
if (i >= tries || !String(e?.message).includes('swarm state is busy')) throw e;
await new Promise(r => setTimeout(r, 100 * 2 ** i));
}
}
} Defensive patterns
Strategy: retry
Validate before calling
import { statSync, existsSync } from 'node:fs';
// Pre-flight: report a stale-looking swarm lock before doing work that will need it
function swarmLockAgeMs(lockPath: string): number | null {
try { return Date.now() - statSync(lockPath).mtimeMs; } catch { return null; }
}
// Stale locks (>10_000ms) are self-healing; fresh ones mean live contention — expect to retry. Type guard
function isSwarmBusyError(e: unknown): e is Error {
return e instanceof Error && e.message.includes('swarm state is busy');
} Try / catch
export async function postTaskOutcome(signal: ApscSignal, attempts = 5): Promise<ReturnType<typeof recordSwarmPheromoneSignal>> {
for (let i = 0; ; i++) {
try { return recordSwarmPheromoneSignal(signal); }
catch (e) {
if (i + 1 >= attempts || !isSwarmBusyError(e)) throw e;
await new Promise(r => setTimeout(r, 100 * 2 ** i)); // 100,200,400,800ms — contention window is ~1s
}
}
} Prevention
- Serialize post-task outcome recording through a single queue in your process instead of firing N parallel hooks.
- Run exactly one MCP server/daemon per swarm state directory; duplicates keep locks hot.
- Budget for retry: the lock loop only waits ~1s (100 x 10ms) before throwing, so bursts need caller backoff.
When it happens
Trigger: Many concurrent hooks_post-task calls (parallel agents finishing simultaneously) all trying to record outcome signals at once; a second MCP server process writing the same ~/.claude-flow swarm state; one writer holding the lock for >1s while its lock file mtime is younger than 10s so no stale-break happens.
Common situations: Large pheromone-adaptive swarms with parallel task completion; running two claude-flow instances against the same home directory; a slow disk making the 100x10ms budget insufficient; leftover lock from a process that is still alive but hung (not stale enough to break).
Related errors
- Storage is locked by another process
- timed out acquiring ai-budget lock
- timed out acquiring flywheel attempts lock
- timed out acquiring flywheel transaction lock
- timed out acquiring repo-supervisor lock
AI-assisted analysis of ruvnet/ruflo@fa13ee4ad6 (2026-08-18).
Data as JSON: /api/errors/c8be461867474ef2.
Report an issue: GitHub.
Appendix: source
Thrown at v3/@claude-flow/cli/src/mcp-tools/swarm-tools.ts:188
const lockPath = join(getSwarmDir(), SWARM_STATE_LOCK);
let descriptor: number | undefined;
const waiter = new Int32Array(new SharedArrayBuffer(4));
for (let attempt = 0; attempt < 100; attempt++) {
try {
descriptor = openSync(lockPath, 'wx', 0o600);
break;
} catch (error) {
if ((error as NodeJS.ErrnoException).code !== 'EEXIST') throw error;
try {
if (Date.now() - statSync(lockPath).mtimeMs > SWARM_LOCK_STALE_MS) {
unlinkSync(lockPath);
continue;
}
} catch { /* another writer released it */ }
Atomics.wait(waiter, 0, 0, 10);
}
}
if (descriptor === undefined) throw new Error('swarm state is busy; retry the outcome update');
try {
return operation();
} finally {
try { closeSync(descriptor); } catch { /* best effort */ }
try { unlinkSync(lockPath); } catch { /* best effort */ }
}
}
// Input validation
const VALID_TOPOLOGIES = new Set([
'hierarchical', 'mesh', 'hierarchical-mesh', 'ring', 'star', 'hybrid', 'adaptive', 'pheromone-adaptive',
]);
function latestRunningSwarm(store: SwarmStore): SwarmState | undefined {
return Object.values(store.swarms)
.filter((swarm) => swarm.status === 'running')
.sort((a, b) => new Date(b.updatedAt).getTime() - new Date(a.updatedAt).getTime())[0];
}View on GitHub (pinned to fa13ee4ad6)