risingwavelabs/risingwave · error

The schema of iceberg equality delete file must be…

Error message

The schema of iceberg equality delete file must be consistent

What it means

When multiple equality-delete files are collected for a scan, all of them must use the same equality field ids (i.e., the same delete schema). If a second equality delete file has different `equality_ids` than the ones already collected, RisingWave cannot apply them consistently and bails with this error.

Solutions

  1. Rewrite the affected data with a compaction/rewrite job (e.g. Spark `rewrite_data_files` / `rewrite_delete_files`) so only one equality-delete schema remains.
  2. Align writers so all use the same equality fields for deletes.
  3. Check schema evolution history; avoid changing delete key columns, or compact deletes after evolving the schema.
  4. Use copy-on-write mode for deletes if mixed schemas are unavoidable.
Defensive patterns

Strategy: validation

Validate before calling

let ids: HashSet<_> = tasks.iter().flat_map(|t| t.deletes.iter())
    .filter(|d| d.file_type == iceberg::spec::DataContentType::EqualityDeletes)
    .map(|d| d.equality_ids.clone()).collect();
if ids.len() > 1 { /* compact deletes or reject scan */ }

Try / catch

match list_scan_tasks(...) {
    Err(e) if e.to_string().contains("equality delete file must be consistent") => {
        // trigger compaction/rewrite or scan without MoR deletes
    }
    r => r,
}

Prevention

When it happens

Trigger: `list_scan_tasks`/`list_scan_tasks_inner` encountering two or more equality delete files in a scan task's deletes where their `equality_ids` differ — e.g. a table written by different writers with different equality-delete schemas or evolved schema between delete writes.

Common situations: Tables written by multiple engines/writers (Spark, Flink, Trino) with mismatched equality-delete field sets; Iceberg schema evolution changing the delete columns between writes; copy/merge operations that mixed delete schemas.

Understand the failure class

Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.

Related errors


AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11). Data as JSON: /api/errors/c0ebc8662c43cd21. Report an issue: GitHub.

Appendix: source

Thrown at src/connector/src/source/iceberg/mod.rs:397

        let scan = scan_builder.build()?;
        let file_scan_stream = scan.plan_files().await?;

        #[for_await]
        for task in file_scan_stream {
            let task: FileScanTask = task?;

            // Collect delete files for separate scan types, but keep task.deletes intact
            for delete_file in &task.deletes {
                match delete_file.file_type {
                    iceberg::spec::DataContentType::Data => {
                        bail!("Data file should not in task deletes");
                    }
                    iceberg::spec::DataContentType::EqualityDeletes => {
                        if equality_delete_files_set.insert(delete_file.file_path.clone()) {
                            if equality_delete_ids.is_none() {
                                equality_delete_ids = delete_file.equality_ids.clone();
                            } else if equality_delete_ids != delete_file.equality_ids {
                                bail!(
                                    "The schema of iceberg equality delete file must be consistent"
                                );
                            }
                            equality_delete_files.push(delete_file.to_file_scan_task(&task));
                        }
                    }
                    iceberg::spec::DataContentType::PositionDeletes => {
                        if position_delete_files_set.insert(delete_file.file_path.clone()) {
                            position_delete_files.push(delete_file.to_file_scan_task(&task));
                        }
                    }
                }
            }

            // Top-level scan tasks always represent data files. Keep their delete
            // descriptors intact so the SDK reader can apply them when requested.
            data_files.push(task);
        }

View on GitHub (pinned to 6469eb736d)