{"record":{"id":"16b72dbedc5b15e6","repo":"Hmbown/CodeWhale","slug":"unaccepted-task-stage-changed-during-recovery","errorCode":null,"errorMessage":"Unaccepted task stage changed during recovery","messagePattern":"Unaccepted task stage changed during recovery","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"crates/tui/src/task_manager.rs","lineNumber":1791,"sourceCode":"            if task.execution_scope.as_deref() != Some(self.execution_scope()) {\n                bail!(\n                    \"Task execution scope is unverified or belongs to another Runtime; refusing adoption\"\n                );\n            }\n            let task_path = self.tasks_dir.join(format!(\"{}.json\", task.id));\n            // The staged extension is intentionally not `.json`, so startup\n            // replay ignores an interrupted create until the queue write has\n            // succeeded and this file is atomically promoted.\n            let staged_task_path = self.tasks_dir.join(format!(\".{}.json.pending\", task.id));\n            if recover_stage {\n                if let Some(accepted) = self.read_bound_task(&task.id)? {\n                    validate_bound_task_request(&accepted, &NewTaskRequest::from_task(&task))?;\n                    return Ok(accepted);\n                }\n                let current = read_bound_task_file(&staged_task_path, &task.id)?\n                    .context(\"Unaccepted task stage disappeared during recovery\")?;\n                if serde_json::to_value(&current)? != serde_json::to_value(&task)? {\n                    bail!(\"Unaccepted task stage changed during recovery\");\n                }\n            }\n            if state.tasks.contains_key(&task.id)\n                || task_path.exists()\n                || (!recover_stage && staged_task_path.exists())\n            {\n                bail!(\"Task id already exists: {}\", task.id);\n            }\n            let mut next_queue = state.queue.clone();\n            if !next_queue.contains(&task.id) {\n                next_queue.push_back(task.id.clone());\n            }\n\n            // Stage the owner record, then persist its queue membership, then\n            // atomically promote it. A crash before promotion leaves either an\n            // ignored staged file or a queue entry with no task (which replay\n            // drops); a crash after promotion leaves the complete runnable\n            // pair. In-memory scheduling is published only after all three.","sourceCodeStart":1773,"sourceCodeEnd":1809,"githubUrl":"https://github.com/Hmbown/CodeWhale/blob/73e0f67d83c59909b571efdfc88c4bc28c309cb1/crates/tui/src/task_manager.rs#L1773-L1809","documentation":"During recovery of an unaccepted (staged) task, the manager re-reads the staged task file and compares it byte-for-field with the task recovered in memory. If the durable staged record differs from the recovered in-memory record, recovery cannot be trusted, so it bails instead of admitting a mutated task. This protects against concurrent writers or corruption changing a task between the two reads.","triggerScenarios":"Calling task recovery/admission for a task that is still staged (unaccepted) while the staged_task_path file contents differ from the in-memory `task` (JSON value comparison fails).","commonSituations":"Another process or a previous crashed run rewrote or partially wrote the staged file during recovery; manual editing of the staged JSON; a schema migration or version skew between processes writing different field values.","solutions":["Retry recovery when no other writer is active (check the store lock)","Delete the stale staged file if the task is known-abandoned and re-admit the task","Ensure only one TUI/manager process writes to this task store at a time"],"exampleFix":"// before: recovering while another process still mutates the staged file\nrecover_task(&task) // -> Unaccepted task stage changed during recovery\n// after: take the exclusive store lock first\nlet _owner = acquire_store_lock(&path).await?;\nrecover_task(&task)","handlingStrategy":"try-catch","validationCode":"let staged: serde_json::Value = serde_json::from_str(&std::fs::read_to_string(&staged_path)?)?;\nlet live = serde_json::to_value(&task)?;\nif staged != live { return Err(anyhow!(\"staged file differs; resolve before recovery\")); }","typeGuard":"fn staged_matches(task: &TaskRecord, staged_path: &Path) -> bool {\n    std::fs::read(staged_path).ok()\n        .and_then(|b| serde_json::from_slice::<serde_json::Value>(&b).ok())\n        .map(|v| v == serde_json::to_value(task).unwrap()).unwrap_or(false)\n}","tryCatchPattern":"match manager.recover_task(&task).await {\n    Err(e) if e.to_string().contains(\"changed during recovery\") => {\n        // re-run once the store is quiescent, or discard the staged file\n        manager.recover_task(&task).await?;\n    }\n    other => other?,\n}","preventionTips":["Hold the store lock for the whole recovery","Never edit staged task files by hand","Run only one manager process per task-store directory"],"tags":["recovery","persistence","task-manager","race-condition"],"backgroundTag":"internal-invariant-violation","analyzedSha":"73e0f67d83c59909b571efdfc88c4bc28c309cb1","analyzedAt":"2026-09-22T01:30:00.501Z","contentChangedAt":"2026-09-22T01:30:00.501Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}