Hmbown/CodeWhale · error

suspended MCP child had no resumable thread

Error message

suspended MCP child had no resumable thread

What it means

On Windows, contain_windows_child suspends all threads of the freshly spawned MCP child via a tool-help snapshot, then resumes them after attaching the job object. If it found no threads it could resume (resumed == 0), it fails — the child would otherwise be left suspended forever or the job containment would be unreliable.

Solutions

  1. Run the server command directly instead of through a shell/cmd wrapper that re-execs.
  2. Verify the command itself works outside the MCP client (missing DLLs or bad interpreter cause instant exit).
  3. Retry the spawn — this can be a startup race with a fast-crashing child.
  4. Check antivirus/EDR logs for interference with the child process.
  5. Report upstream if a legitimately short-lived-but-alive child consistently trips this.

Example fix

// before (config)
command: "cmd /c server.exe"
// after
command: "C:\\path\\server.exe"
Defensive patterns

Strategy: retry

Validate before calling

// Windows: check the executable exists and is a native binary before spawning
use std::path::Path;
fn command_spawnable(cfg: &McpServerConfig) -> bool {
    Path::new(&cfg.command).is_file() || which::which(&cfg.command).is_ok()
}

Try / catch

let client = match spawn_with_timeouts(&config, handshake, request_timeout) {
    Ok(c) => c,
    Err(e) if e.to_string().contains("no resumable thread") && attempt < 2 => {
        tokio::time::sleep(Duration::from_millis(250)).await;
        continue; // retry spawn; likely a startup race
    }
    Err(e) => return Err(e),
};

Prevention

When it happens

Trigger: Windows-only: spawning an MCP server whose process exits or whose threads disappear between snapshot creation and Thread32First iteration, so the snapshot contains zero resumable threads.

Common situations: Child crashing instantly on launch (bad interpreter, missing DLL); race where the process terminated before enumeration; security software interfering with the new process; shell wrappers that re-exec.

Understand the failure class

Background: "This is a bug, please report it": internal invariant violations, unreachable panics, and SNH errors explained — this error's family across 47 libraries.

Related errors


AI-assisted analysis of Hmbown/CodeWhale@73e0f67d83 (2026-09-22). Data as JSON: /api/errors/f7acb5dada5c09cc. Report an issue: GitHub.

Appendix: source

Thrown at crates/mcp/src/stdio_client.rs:835

        let mut entry = THREADENTRY32 {
            dwSize: std::mem::size_of::<THREADENTRY32>() as u32,
            ..Default::default()
        };
        let mut next = Thread32First(HANDLE(snapshot.as_raw_handle()), &mut entry);
        let mut resumed = 0usize;
        while next.is_ok() {
            if entry.th32OwnerProcessID == child.id() {
                let thread = OpenThread(THREAD_SUSPEND_RESUME, false, entry.th32ThreadID)?;
                let thread = OwnedHandle::from_raw_handle(thread.0);
                if ResumeThread(HANDLE(thread.as_raw_handle())) == u32::MAX {
                    return Err(io::Error::last_os_error().into());
                }
                resumed += 1;
            }
            next = Thread32Next(HANDLE(snapshot.as_raw_handle()), &mut entry);
        }
        if resumed == 0 {
            bail!("suspended MCP child had no resumable thread");
        }
        Ok(job)
    }
}

impl Connection {
    fn send(&mut self, message: &Value) -> Result<()> {
        let stdin = self
            .stdin
            .as_ref()
            .context("connection stdin already closed")?;
        let mut stdin = stdin
            .lock()
            .map_err(|_| anyhow!("connection stdin poisoned by an earlier panic"))?;
        write_jsonrpc_line(&mut *stdin, message)
    }

    fn request(

View on GitHub (pinned to 73e0f67d83)