risingwavelabs/risingwave · error · ConnectorError

too many CDC snapshot splits

Error message

too many CDC snapshot splits

What it means

`try_increase_split_id` increments the running snapshot split id (i64) for each generated split and guards against overflow. Incrementing beyond `i64::MAX` is treated as an unrecoverable 'too many CDC snapshot splits' error. In practice this requires an astronomically large table or a looping id assignment bug.

Solutions

  1. Increase `backfill_num_rows_per_split` so fewer splits are generated.
  2. Verify split-id assignment logic is not looping/reusing ids incorrectly.
  3. Backfill the table in phases (partition upstream) instead of one giant snapshot.
  4. If you maintain the code, consider widening the id space or reusing ids after completion.

Example fix

// before: tiny split size on a huge table
WITH (connector = 'postgres-cdc', backfill_num_rows_per_split = '1');
// after
WITH (connector = 'postgres-cdc', backfill_num_rows_per_split = '100000');
Defensive patterns

Strategy: validation

Validate before calling

// Estimate split count before generating
let est_splits = total_rows / backfill_num_rows_per_split;
if est_splits > i64::MAX as f64 as i64 { /* unreachable, but guard anyway */ }

Try / catch

match try_increase_split_id(&mut split_id) {
    Err(e) if e.to_string().contains("too many CDC snapshot splits") => {
        // stop splitting, use the splits generated so far
        break;
    },
    r => r?,
}

Prevention

When it happens

Trigger: Calling `try_increase_split_id` (from `as_even_splits` or `as_uneven_splits`) when `split_id` is already `i64::MAX` so `checked_add(1)` returns None.

Common situations: Extremely large tables combined with a tiny `backfill_num_rows_per_split`; a bug in split-id assignment causing unbounded splitting; artificially crafted tests with a pre-set max split id.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/63979b21a60377f2. Report an issue: GitHub.

Appendix: source

Thrown at src/connector/src/source/cdc/external/postgres.rs:844

fn to_int_scalar(i: i64, data_type: &DataType) -> ScalarImpl {
    match data_type {
        DataType::Int16 => ScalarImpl::Int16(i.try_into().unwrap()),
        DataType::Int32 => ScalarImpl::Int32(i.try_into().unwrap()),
        DataType::Int64 => ScalarImpl::Int64(i),
        _ => {
            panic!("Can't convert int {} to ScalarImpl::{}", i, data_type)
        }
    }
}

fn try_increase_split_id(split_id: &mut i64) -> ConnectorResult<()> {
    match split_id.checked_add(1) {
        Some(s) => {
            *split_id = s;
            Ok(())
        }
        None => Err(anyhow::anyhow!("too many CDC snapshot splits").into()),
    }
}

/// Use the first column of primary keys to split table.
fn is_supported_even_split_data_type(data_type: &DataType) -> bool {
    matches!(
        data_type,
        DataType::Int16 | DataType::Int32 | DataType::Int64
    )
}

pub fn type_name_to_pg_type(ty_name: &str) -> Option<PgType> {
    let ty_name_lower = ty_name.to_lowercase();
    // Handle array types (prefixed with _)
    if let Some(base_type) = ty_name_lower.strip_prefix('_') {
        match base_type {
            "int2" => Some(PgType::INT2_ARRAY),
            "int4" => Some(PgType::INT4_ARRAY),

View on GitHub (pinned to 6469eb736d)