risingwavelabs/risingwave · error
Iceberg scan limit can only be planned for data scan tasks
Error message
Iceberg scan limit can only be planned for data scan tasks
What it means
`plan_splits` allows an optional row `limit` to be planned only when the task is a plain DataScan. Limit pushdown cannot be applied to equality-delete or position-delete scans because reading delete files requires scanning all applicable deletes; applying a limit would produce incorrect results, so the code bails.
Solutions
- Remove/compaction the delete files first (rewrite/compact the table) so the scan becomes a DataScan and limit pushdown is allowed.
- Avoid LIMIT pushdown for that query (e.g. rewrite the query or run without relying on limit optimization).
- Use copy-on-write delete mode so deletes are merged into data files and scans stay DataScan.
- If it is a planner bug, gate limit pushdown on scan_type before calling plan_splits.
Defensive patterns
Strategy: validation
Validate before calling
let limit = if task.scan_type() == IcebergScanType::DataScan { limit } else { None };
let splits = planner.plan_splits(task, split_num, limit)?; Type guard
fn limit_allowed(t: &IcebergScanType, limit: Option<u64>) -> bool {
limit.is_none() || *t == IcebergScanType::DataScan
} Try / catch
match plan_splits(task, split_num, limit) {
Err(e) if e.to_string().contains("scan limit can only be planned for data scan") => {
// compact deletes or retry without limit pushdown
}
r => r,
} Prevention
- Only push LIMIT into DataScan tasks
- Compact/rewrite MoR delete files before relying on limit optimization
- Prefer copy-on-write tables for limit-heavy workloads
When it happens
Trigger: `plan_splits` (called during Iceberg split planning) receiving `limit: Some(n)` while `task.scan_type() != IcebergScanType::DataScan` — i.e. a LIMIT query planned against a merge-on-read table with equality or position delete files.
Common situations: Running `SELECT ... LIMIT n` on an Iceberg table that has pending delete files; a planner regression that pushes a limit into delete scans; query with filters hitting rows covered by MoR deletes.
Understand the failure class
Background: "Invalid state transition" errors: "status must be X, actually Y", "already rejected/charging/uninstalled", "cannot ... while running" — what they mean when a library rejects your call — this error's family across 31 libraries.
Related errors
- bounded compaction branch
- bounded compaction head sequence
- bounded compaction is not supported for copy-on-write tasks
- bounded compaction requires Iceberg format V2 or V3
- Can't create iceberg dv merger commit result from empty…
AI-assisted analysis of risingwavelabs/risingwave@6469eb736d (2026-09-11).
Data as JSON: /api/errors/b9488132987c8bb4.
Report an issue: GitHub.
Appendix: source
Thrown at src/connector/src/source/iceberg/planner.rs:671
) -> Self {
Self {
chunk_size,
need_seq_num,
need_file_path_and_pos,
handle_delete_files,
}
}
}
impl IcebergScanTaskPlanner {
pub fn plan_splits(
task: IcebergFileScanTask,
split_num: usize,
limit: Option<u64>,
) -> ConnectorResult<Vec<IcebergSplit>> {
let scan_type = task.scan_type();
if limit.is_some() && scan_type != IcebergScanType::DataScan {
bail!("Iceberg scan limit can only be planned for data scan tasks");
}
let task_batches = Self::plan_task_batches(
task,
split_num,
limit,
IcebergScanTaskBatchMode::PreserveParallelism,
);
task_batches
.into_iter()
.enumerate()
.map(|(id, tasks)| {
Ok(IcebergSplit {
split_id: id.try_into().unwrap(),
task: IcebergFileScanTask::from_tasks(scan_type, tasks)?,
limit,
})View on GitHub (pinned to 6469eb736d)