ruvnet/ruflo · warning

swarm state is busy; retry the outcome update

Error message

swarm state is busy; retry the outcome update

What it means

Swarm store mutations that feed APSC pheromone signals (recordSwarmPheromoneSignal, shared with hooks post-task outcome updates) run under an exclusive advisory file lock (swarm-state.lock, opened with 'wx'). The acquisition loop tries 100 times with ~10ms waits (~1s total) and force-breaks locks older than SWARM_LOCK_STALE_MS = 10s. If the lock is still held (continuously refreshed by another live writer, or heavy contention), descriptor stays undefined and this error is thrown — by design, so the caller retries rather than corrupting swarm state.

Solutions

  1. Retry the outcome update — the message says exactly that; contention usually clears within a second.
  2. Retry with exponential backoff (e.g. 100ms, 200ms, 400ms) so bursts of parallel post-task hooks don't fight the lock.
  3. Ensure only ONE MCP server process owns the swarm state directory; kill duplicate daemons (tman list / ps) writing to the same home.
  4. If it persists >10s, the stale-lock breaker will unlink it on a later attempt; verify no zombie process still refreshes the lock's mtime.

Example fix

// before
const r = recordSwarmPheromoneSignal(signal); // throws 'busy' under load

// after
async function postOutcome(signal, tries = 5) {
  for (let i = 0; ; i++) {
    try { return recordSwarmPheromoneSignal(signal); }
    catch (e) {
      if (i >= tries || !String(e?.message).includes('swarm state is busy')) throw e;
      await new Promise(r => setTimeout(r, 100 * 2 ** i));
    }
  }
}
Defensive patterns

Strategy: retry

Validate before calling

import { statSync, existsSync } from 'node:fs';
// Pre-flight: report a stale-looking swarm lock before doing work that will need it
function swarmLockAgeMs(lockPath: string): number | null {
  try { return Date.now() - statSync(lockPath).mtimeMs; } catch { return null; }
}
// Stale locks (>10_000ms) are self-healing; fresh ones mean live contention — expect to retry.

Type guard

function isSwarmBusyError(e: unknown): e is Error {
  return e instanceof Error && e.message.includes('swarm state is busy');
}

Try / catch

export async function postTaskOutcome(signal: ApscSignal, attempts = 5): Promise<ReturnType<typeof recordSwarmPheromoneSignal>> {
  for (let i = 0; ; i++) {
    try { return recordSwarmPheromoneSignal(signal); }
    catch (e) {
      if (i + 1 >= attempts || !isSwarmBusyError(e)) throw e;
      await new Promise(r => setTimeout(r, 100 * 2 ** i)); // 100,200,400,800ms — contention window is ~1s
    }
  }
}

Prevention

When it happens

Trigger: Many concurrent hooks_post-task calls (parallel agents finishing simultaneously) all trying to record outcome signals at once; a second MCP server process writing the same ~/.claude-flow swarm state; one writer holding the lock for >1s while its lock file mtime is younger than 10s so no stale-break happens.

Common situations: Large pheromone-adaptive swarms with parallel task completion; running two claude-flow instances against the same home directory; a slow disk making the 100x10ms budget insufficient; leftover lock from a process that is still alive but hung (not stale enough to break).

Related errors


AI-assisted analysis of ruvnet/ruflo@fa13ee4ad6 (2026-08-18). Data as JSON: /api/errors/c8be461867474ef2. Report an issue: GitHub.

Appendix: source

Thrown at v3/@claude-flow/cli/src/mcp-tools/swarm-tools.ts:188

  const lockPath = join(getSwarmDir(), SWARM_STATE_LOCK);
  let descriptor: number | undefined;
  const waiter = new Int32Array(new SharedArrayBuffer(4));
  for (let attempt = 0; attempt < 100; attempt++) {
    try {
      descriptor = openSync(lockPath, 'wx', 0o600);
      break;
    } catch (error) {
      if ((error as NodeJS.ErrnoException).code !== 'EEXIST') throw error;
      try {
        if (Date.now() - statSync(lockPath).mtimeMs > SWARM_LOCK_STALE_MS) {
          unlinkSync(lockPath);
          continue;
        }
      } catch { /* another writer released it */ }
      Atomics.wait(waiter, 0, 0, 10);
    }
  }
  if (descriptor === undefined) throw new Error('swarm state is busy; retry the outcome update');
  try {
    return operation();
  } finally {
    try { closeSync(descriptor); } catch { /* best effort */ }
    try { unlinkSync(lockPath); } catch { /* best effort */ }
  }
}

// Input validation
const VALID_TOPOLOGIES = new Set([
  'hierarchical', 'mesh', 'hierarchical-mesh', 'ring', 'star', 'hybrid', 'adaptive', 'pheromone-adaptive',
]);

function latestRunningSwarm(store: SwarmStore): SwarmState | undefined {
  return Object.values(store.swarms)
    .filter((swarm) => swarm.status === 'running')
    .sort((a, b) => new Date(b.updatedAt).getTime() - new Date(a.updatedAt).getTime())[0];
}

View on GitHub (pinned to fa13ee4ad6)