{"record":{"id":"c8be461867474ef2","repo":"ruvnet/ruflo","slug":"swarm-state-is-busy-retry-the-outcome-update","errorCode":null,"errorMessage":"swarm state is busy; retry the outcome update","messagePattern":"swarm state is busy; retry the outcome update","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"warning","filePath":"v3/@claude-flow/cli/src/mcp-tools/swarm-tools.ts","lineNumber":188,"sourceCode":"  const lockPath = join(getSwarmDir(), SWARM_STATE_LOCK);\n  let descriptor: number | undefined;\n  const waiter = new Int32Array(new SharedArrayBuffer(4));\n  for (let attempt = 0; attempt < 100; attempt++) {\n    try {\n      descriptor = openSync(lockPath, 'wx', 0o600);\n      break;\n    } catch (error) {\n      if ((error as NodeJS.ErrnoException).code !== 'EEXIST') throw error;\n      try {\n        if (Date.now() - statSync(lockPath).mtimeMs > SWARM_LOCK_STALE_MS) {\n          unlinkSync(lockPath);\n          continue;\n        }\n      } catch { /* another writer released it */ }\n      Atomics.wait(waiter, 0, 0, 10);\n    }\n  }\n  if (descriptor === undefined) throw new Error('swarm state is busy; retry the outcome update');\n  try {\n    return operation();\n  } finally {\n    try { closeSync(descriptor); } catch { /* best effort */ }\n    try { unlinkSync(lockPath); } catch { /* best effort */ }\n  }\n}\n\n// Input validation\nconst VALID_TOPOLOGIES = new Set([\n  'hierarchical', 'mesh', 'hierarchical-mesh', 'ring', 'star', 'hybrid', 'adaptive', 'pheromone-adaptive',\n]);\n\nfunction latestRunningSwarm(store: SwarmStore): SwarmState | undefined {\n  return Object.values(store.swarms)\n    .filter((swarm) => swarm.status === 'running')\n    .sort((a, b) => new Date(b.updatedAt).getTime() - new Date(a.updatedAt).getTime())[0];\n}","sourceCodeStart":170,"sourceCodeEnd":206,"githubUrl":"https://github.com/ruvnet/ruflo/blob/fa13ee4ad60ac2090b1480656eb233521790d640/v3/@claude-flow/cli/src/mcp-tools/swarm-tools.ts#L170-L206","documentation":"Swarm store mutations that feed APSC pheromone signals (recordSwarmPheromoneSignal, shared with hooks post-task outcome updates) run under an exclusive advisory file lock (swarm-state.lock, opened with 'wx'). The acquisition loop tries 100 times with ~10ms waits (~1s total) and force-breaks locks older than SWARM_LOCK_STALE_MS = 10s. If the lock is still held (continuously refreshed by another live writer, or heavy contention), descriptor stays undefined and this error is thrown — by design, so the caller retries rather than corrupting swarm state.","triggerScenarios":"Many concurrent hooks_post-task calls (parallel agents finishing simultaneously) all trying to record outcome signals at once; a second MCP server process writing the same ~/.claude-flow swarm state; one writer holding the lock for >1s while its lock file mtime is younger than 10s so no stale-break happens.","commonSituations":"Large pheromone-adaptive swarms with parallel task completion; running two claude-flow instances against the same home directory; a slow disk making the 100x10ms budget insufficient; leftover lock from a process that is still alive but hung (not stale enough to break).","solutions":["Retry the outcome update — the message says exactly that; contention usually clears within a second.","Retry with exponential backoff (e.g. 100ms, 200ms, 400ms) so bursts of parallel post-task hooks don't fight the lock.","Ensure only ONE MCP server process owns the swarm state directory; kill duplicate daemons (tman list / ps) writing to the same home.","If it persists >10s, the stale-lock breaker will unlink it on a later attempt; verify no zombie process still refreshes the lock's mtime."],"exampleFix":"// before\nconst r = recordSwarmPheromoneSignal(signal); // throws 'busy' under load\n\n// after\nasync function postOutcome(signal, tries = 5) {\n  for (let i = 0; ; i++) {\n    try { return recordSwarmPheromoneSignal(signal); }\n    catch (e) {\n      if (i >= tries || !String(e?.message).includes('swarm state is busy')) throw e;\n      await new Promise(r => setTimeout(r, 100 * 2 ** i));\n    }\n  }\n}","handlingStrategy":"retry","validationCode":"import { statSync, existsSync } from 'node:fs';\n// Pre-flight: report a stale-looking swarm lock before doing work that will need it\nfunction swarmLockAgeMs(lockPath: string): number | null {\n  try { return Date.now() - statSync(lockPath).mtimeMs; } catch { return null; }\n}\n// Stale locks (>10_000ms) are self-healing; fresh ones mean live contention — expect to retry.","typeGuard":"function isSwarmBusyError(e: unknown): e is Error {\n  return e instanceof Error && e.message.includes('swarm state is busy');\n}","tryCatchPattern":"export async function postTaskOutcome(signal: ApscSignal, attempts = 5): Promise<ReturnType<typeof recordSwarmPheromoneSignal>> {\n  for (let i = 0; ; i++) {\n    try { return recordSwarmPheromoneSignal(signal); }\n    catch (e) {\n      if (i + 1 >= attempts || !isSwarmBusyError(e)) throw e;\n      await new Promise(r => setTimeout(r, 100 * 2 ** i)); // 100,200,400,800ms — contention window is ~1s\n    }\n  }\n}","preventionTips":["Serialize post-task outcome recording through a single queue in your process instead of firing N parallel hooks.","Run exactly one MCP server/daemon per swarm state directory; duplicates keep locks hot.","Budget for retry: the lock loop only waits ~1s (100 x 10ms) before throwing, so bursts need caller backoff."],"tags":["swarm","file-lock","concurrency","pheromone","retry"],"backgroundTag":"file-lock-contention","analyzedSha":"fa13ee4ad60ac2090b1480656eb233521790d640","analyzedAt":"2026-08-18T21:34:22.708Z","contentChangedAt":"2026-08-18T21:34:22.708Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}