risingwavelabs/risingwave · error

CDC table backfill reschedule is unavailable because the…

Error message

CDC table backfill reschedule is unavailable because the system is in recovery state

What it means

CDC table backfill parallelism rescheduling requires the barrier manager to be in a running state. When the system is recovering (e.g. after a leader change or crash recovery), the meta node rejects reschedule requests to avoid conflicting with barrier recovery. This is a guard against mutating the topology while recovery is in flight.

Solutions

  1. Wait until the system finishes recovery and the barrier manager returns to running, then retry the reschedule.
  2. Check meta node logs / metrics for barrier manager recovery status before issuing ALTER commands.
  3. Add retry-with-backoff around the reschedule call so transient recovery windows are handled automatically.

Example fix

// before
meta.reschedule_cdc_table_backfill(job_id, policy).await?;
// after
loop {
    match meta.reschedule_cdc_table_backfill(job_id, policy).await {
        Ok(()) => break,
        Err(e) if e.to_string().contains("recovery state") => tokio::time::sleep(Duration::from_secs(5)).await,
        Err(e) => return Err(e),
    }
}
Defensive patterns

Strategy: retry

Validate before calling

// Rust caller check
if !meta.barrier_manager_is_running().await {
    return Err("system in recovery; reschedule unavailable".into());
}

Try / catch

match reschedule.await {
    Err(e) if e.to_string().contains("recovery state") => retry_with_backoff(e),
    other => other?,
}

Prevention

When it happens

Trigger: Calling reschedule_cdc_table_backfill (via ALTER CDC TABLE backfill parallelism APIs) while barrier_manager.check_status_running() returns Err, i.e. the barrier manager is pausing/injecting/recovering rather than running.

Common situations: Issuing an ALTER to adjust CDC backfill parallelism while the cluster is recovering from a meta node failover or crash; automation scripts racing with cluster recovery; retry storms right after a failover.

Understand the failure class

Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/82273b33704f96a8. Report an issue: GitHub.

Appendix: source

Thrown at src/meta/src/rpc/ddl_controller.rs:639

            tracing::info!(
                "alter backfill parallelism is set to deferred mode because the system is in recovery state"
            );
            deferred = true;
        }

        self.stream_manager
            .reschedule_streaming_job_backfill_parallelism(job_id, parallelism, deferred)
            .await
    }

    pub async fn reschedule_cdc_table_backfill(
        &self,
        job_id: JobId,
        target: ReschedulePolicy,
    ) -> MetaResult<()> {
        tracing::info!("alter CDC table backfill parallelism");
        if self.barrier_manager.check_status_running().is_err() {
            return Err(anyhow::anyhow!("CDC table backfill reschedule is unavailable because the system is in recovery state").into());
        }
        self.stream_manager
            .reschedule_cdc_table_backfill(job_id, target)
            .await
    }

    pub async fn reschedule_fragments(
        &self,
        fragment_targets: HashMap<FragmentId, Option<StreamingParallelism>>,
    ) -> MetaResult<()> {
        tracing::info!(
            "altering parallelism for fragments {:?}",
            fragment_targets.keys()
        );
        let fragment_targets = fragment_targets
            .into_iter()
            .map(|(fragment_id, parallelism)| (fragment_id as CatalogFragmentId, parallelism))
            .collect();

View on GitHub (pinned to 6469eb736d)