Hmbown/CodeWhale · error

Unaccepted task stage changed during recovery

Error message

Unaccepted task stage changed during recovery

What it means

During recovery of an unaccepted (staged) task, the manager re-reads the staged task file and compares it byte-for-field with the task recovered in memory. If the durable staged record differs from the recovered in-memory record, recovery cannot be trusted, so it bails instead of admitting a mutated task. This protects against concurrent writers or corruption changing a task between the two reads.

Solutions

  1. Retry recovery when no other writer is active (check the store lock)
  2. Delete the stale staged file if the task is known-abandoned and re-admit the task
  3. Ensure only one TUI/manager process writes to this task store at a time

Example fix

// before: recovering while another process still mutates the staged file
recover_task(&task) // -> Unaccepted task stage changed during recovery
// after: take the exclusive store lock first
let _owner = acquire_store_lock(&path).await?;
recover_task(&task)
Defensive patterns

Strategy: try-catch

Validate before calling

let staged: serde_json::Value = serde_json::from_str(&std::fs::read_to_string(&staged_path)?)?;
let live = serde_json::to_value(&task)?;
if staged != live { return Err(anyhow!("staged file differs; resolve before recovery")); }

Type guard

fn staged_matches(task: &TaskRecord, staged_path: &Path) -> bool {
    std::fs::read(staged_path).ok()
        .and_then(|b| serde_json::from_slice::<serde_json::Value>(&b).ok())
        .map(|v| v == serde_json::to_value(task).unwrap()).unwrap_or(false)
}

Try / catch

match manager.recover_task(&task).await {
    Err(e) if e.to_string().contains("changed during recovery") => {
        // re-run once the store is quiescent, or discard the staged file
        manager.recover_task(&task).await?;
    }
    other => other?,
}

Prevention

When it happens

Trigger: Calling task recovery/admission for a task that is still staged (unaccepted) while the staged_task_path file contents differ from the in-memory `task` (JSON value comparison fails).

Common situations: Another process or a previous crashed run rewrote or partially wrote the staged file during recovery; manual editing of the staged JSON; a schema migration or version skew between processes writing different field values.

Understand the failure class

Background: "This is a bug, please report it": internal invariant violations, unreachable panics, and SNH errors explained — this error's family across 47 libraries.

Related errors


AI-assisted analysis of Hmbown/CodeWhale@73e0f67d83 (2026-09-22). Data as JSON: /api/errors/16b72dbedc5b15e6. Report an issue: GitHub.

Appendix: source

Thrown at crates/tui/src/task_manager.rs:1791

            if task.execution_scope.as_deref() != Some(self.execution_scope()) {
                bail!(
                    "Task execution scope is unverified or belongs to another Runtime; refusing adoption"
                );
            }
            let task_path = self.tasks_dir.join(format!("{}.json", task.id));
            // The staged extension is intentionally not `.json`, so startup
            // replay ignores an interrupted create until the queue write has
            // succeeded and this file is atomically promoted.
            let staged_task_path = self.tasks_dir.join(format!(".{}.json.pending", task.id));
            if recover_stage {
                if let Some(accepted) = self.read_bound_task(&task.id)? {
                    validate_bound_task_request(&accepted, &NewTaskRequest::from_task(&task))?;
                    return Ok(accepted);
                }
                let current = read_bound_task_file(&staged_task_path, &task.id)?
                    .context("Unaccepted task stage disappeared during recovery")?;
                if serde_json::to_value(&current)? != serde_json::to_value(&task)? {
                    bail!("Unaccepted task stage changed during recovery");
                }
            }
            if state.tasks.contains_key(&task.id)
                || task_path.exists()
                || (!recover_stage && staged_task_path.exists())
            {
                bail!("Task id already exists: {}", task.id);
            }
            let mut next_queue = state.queue.clone();
            if !next_queue.contains(&task.id) {
                next_queue.push_back(task.id.clone());
            }

            // Stage the owner record, then persist its queue membership, then
            // atomically promote it. A crash before promotion leaves either an
            // ignored staged file or a queue entry with no task (which replay
            // drops); a crash after promotion leaves the complete runnable
            // pair. In-memory scheduling is published only after all three.

View on GitHub (pinned to 73e0f67d83)