{"record":{"id":"82273b33704f96a8","repo":"risingwavelabs/risingwave","slug":"cdc-table-backfill-reschedule-is-unavailable-becau","errorCode":null,"errorMessage":"CDC table backfill reschedule is unavailable because the system is in recovery state","messagePattern":"CDC table backfill reschedule is unavailable because the system is in recovery state","errorType":"exception","errorClass":null,"httpStatus":null,"severity":"error","filePath":"src/meta/src/rpc/ddl_controller.rs","lineNumber":639,"sourceCode":"            tracing::info!(\n                \"alter backfill parallelism is set to deferred mode because the system is in recovery state\"\n            );\n            deferred = true;\n        }\n\n        self.stream_manager\n            .reschedule_streaming_job_backfill_parallelism(job_id, parallelism, deferred)\n            .await\n    }\n\n    pub async fn reschedule_cdc_table_backfill(\n        &self,\n        job_id: JobId,\n        target: ReschedulePolicy,\n    ) -> MetaResult<()> {\n        tracing::info!(\"alter CDC table backfill parallelism\");\n        if self.barrier_manager.check_status_running().is_err() {\n            return Err(anyhow::anyhow!(\"CDC table backfill reschedule is unavailable because the system is in recovery state\").into());\n        }\n        self.stream_manager\n            .reschedule_cdc_table_backfill(job_id, target)\n            .await\n    }\n\n    pub async fn reschedule_fragments(\n        &self,\n        fragment_targets: HashMap<FragmentId, Option<StreamingParallelism>>,\n    ) -> MetaResult<()> {\n        tracing::info!(\n            \"altering parallelism for fragments {:?}\",\n            fragment_targets.keys()\n        );\n        let fragment_targets = fragment_targets\n            .into_iter()\n            .map(|(fragment_id, parallelism)| (fragment_id as CatalogFragmentId, parallelism))\n            .collect();","sourceCodeStart":621,"sourceCodeEnd":657,"githubUrl":"https://github.com/risingwavelabs/risingwave/blob/6469eb736d691e8e9b8a419a57edd6429ca77417/src/meta/src/rpc/ddl_controller.rs#L621-L657","documentation":"CDC table backfill parallelism rescheduling requires the barrier manager to be in a running state. When the system is recovering (e.g. after a leader change or crash recovery), the meta node rejects reschedule requests to avoid conflicting with barrier recovery. This is a guard against mutating the topology while recovery is in flight.","triggerScenarios":"Calling reschedule_cdc_table_backfill (via ALTER CDC TABLE backfill parallelism APIs) while barrier_manager.check_status_running() returns Err, i.e. the barrier manager is pausing/injecting/recovering rather than running.","commonSituations":"Issuing an ALTER to adjust CDC backfill parallelism while the cluster is recovering from a meta node failover or crash; automation scripts racing with cluster recovery; retry storms right after a failover.","solutions":["Wait until the system finishes recovery and the barrier manager returns to running, then retry the reschedule.","Check meta node logs / metrics for barrier manager recovery status before issuing ALTER commands.","Add retry-with-backoff around the reschedule call so transient recovery windows are handled automatically."],"exampleFix":"// before\nmeta.reschedule_cdc_table_backfill(job_id, policy).await?;\n// after\nloop {\n    match meta.reschedule_cdc_table_backfill(job_id, policy).await {\n        Ok(()) => break,\n        Err(e) if e.to_string().contains(\"recovery state\") => tokio::time::sleep(Duration::from_secs(5)).await,\n        Err(e) => return Err(e),\n    }\n}","handlingStrategy":"retry","validationCode":"// Rust caller check\nif !meta.barrier_manager_is_running().await {\n    return Err(\"system in recovery; reschedule unavailable\".into());\n}","typeGuard":null,"tryCatchPattern":"match reschedule.await {\n    Err(e) if e.to_string().contains(\"recovery state\") => retry_with_backoff(e),\n    other => other?,\n}","preventionTips":["Check barrier manager status before ALTER/reschedule calls","Use retry-with-backoff for reschedule operations","Avoid rescheduling during known failover windows"],"tags":["cdc","recovery","concurrency"],"backgroundTag":"invalid-state-transition","analyzedSha":"6469eb736d691e8e9b8a419a57edd6429ca77417","analyzedAt":"2026-09-11T21:06:21.487Z","contentChangedAt":"2026-09-11T21:06:21.487Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}