risingwavelabs/risingwave · error
CDC table backfill reschedule is unavailable because the…
Error message
CDC table backfill reschedule is unavailable because the system is in recovery state
What it means
CDC table backfill parallelism rescheduling requires the barrier manager to be in a running state. When the system is recovering (e.g. after a leader change or crash recovery), the meta node rejects reschedule requests to avoid conflicting with barrier recovery. This is a guard against mutating the topology while recovery is in flight.
Solutions
- Wait until the system finishes recovery and the barrier manager returns to running, then retry the reschedule.
- Check meta node logs / metrics for barrier manager recovery status before issuing ALTER commands.
- Add retry-with-backoff around the reschedule call so transient recovery windows are handled automatically.
Example fix
// before
meta.reschedule_cdc_table_backfill(job_id, policy).await?;
// after
loop {
match meta.reschedule_cdc_table_backfill(job_id, policy).await {
Ok(()) => break,
Err(e) if e.to_string().contains("recovery state") => tokio::time::sleep(Duration::from_secs(5)).await,
Err(e) => return Err(e),
}
} Defensive patterns
Strategy: retry
Validate before calling
// Rust caller check
if !meta.barrier_manager_is_running().await {
return Err("system in recovery; reschedule unavailable".into());
} Try / catch
match reschedule.await {
Err(e) if e.to_string().contains("recovery state") => retry_with_backoff(e),
other => other?,
} Prevention
- Check barrier manager status before ALTER/reschedule calls
- Use retry-with-backoff for reschedule operations
- Avoid rescheduling during known failover windows
When it happens
Trigger: Calling reschedule_cdc_table_backfill (via ALTER CDC TABLE backfill parallelism APIs) while barrier_manager.check_status_running() returns Err, i.e. the barrier manager is pausing/injecting/recovering rather than running.
Common situations: Issuing an ALTER to adjust CDC backfill parallelism while the cluster is recovering from a meta node failover or crash; automation scripts racing with cluster recovery; retry storms right after a failover.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- invalid backfill state: backfill_finished
- invalid backfill state: row_count
- invalid backfill state: unfinished row has null cdc_offset
- a stream has reached the end but some other stream has not…
- adhoc recovery triggered
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/82273b33704f96a8.
Report an issue: GitHub.
Appendix: source
Thrown at src/meta/src/rpc/ddl_controller.rs:639
tracing::info!(
"alter backfill parallelism is set to deferred mode because the system is in recovery state"
);
deferred = true;
}
self.stream_manager
.reschedule_streaming_job_backfill_parallelism(job_id, parallelism, deferred)
.await
}
pub async fn reschedule_cdc_table_backfill(
&self,
job_id: JobId,
target: ReschedulePolicy,
) -> MetaResult<()> {
tracing::info!("alter CDC table backfill parallelism");
if self.barrier_manager.check_status_running().is_err() {
return Err(anyhow::anyhow!("CDC table backfill reschedule is unavailable because the system is in recovery state").into());
}
self.stream_manager
.reschedule_cdc_table_backfill(job_id, target)
.await
}
pub async fn reschedule_fragments(
&self,
fragment_targets: HashMap<FragmentId, Option<StreamingParallelism>>,
) -> MetaResult<()> {
tracing::info!(
"altering parallelism for fragments {:?}",
fragment_targets.keys()
);
let fragment_targets = fragment_targets
.into_iter()
.map(|(fragment_id, parallelism)| (fragment_id as CatalogFragmentId, parallelism))
.collect();View on GitHub (pinned to 6469eb736d)