apache/iceberg · warning
Exception processing split
Error message
Exception processing split {} at {} What it means
TableReader.processElement deserializes an IcebergSourceSplit and iterates its data with a DataIterator, forwarding rows downstream. If reading the split throws, the exception is logged as a warning with the split and timestamp, pushed to the DeleteOrphanFiles ERROR_STREAM, and counted — the operator continues with the next split instead of failing the job, so orphan-file detection proceeds with partial data.
Solutions
- Inspect the exception on the DeleteOrphanFiles.ERROR_STREAM side output for the root cause
- Ensure no concurrent snapshot-expiration or orphan-cleanup job deletes files while the maintenance job reads them
- Align Iceberg versions across the Flink job and cluster to avoid split serialization incompatibilities
- Retry the maintenance job; failed splits reduce accuracy of that cycle but do not corrupt state
Defensive patterns
Strategy: retry
Try / catch
// Read the forwarded exception for diagnosis
result.getSideOutput(DeleteOrphanFiles.ERROR_STREAM)
.executeAndCollect().forEachRemaining(e -> log.error("split read failed", e)); Prevention
- Avoid running snapshot expiration/orphan deletion while the reader is active
- Keep Iceberg versions consistent between job and cluster to avoid split deserialization mismatches
- Retry the maintenance cycle after transient object-store read errors
- Verify data files are readable via a small TableScan smoke test
When it happens
Trigger: Raised in processElement when splitSerializer.deserialize or rowDataReaderFunction.createDataIterator(split) / iteration throws — file deleted between planning and reading (expired snapshots/orphan cleanup), schema mismatch against the metadata-table schema, or object-store read errors.
Common situations: Another job expired snapshots and deleted data/metadata files mid-scan; serialization version mismatch after an Iceberg upgrade of one part of the cluster; transient S3/HDFS read failures; corrupt Parquet/metadata files.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Exception listing files for
- Exception listing files for
- Exception listing files for
- Exception planning scan for
- Exception planning scan for
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/db1ee03235421b8d.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/maintenance/operator/TableReader.java:101
.counter(TableMaintenanceMetrics.ERROR_COUNTER);
this.rowDataReaderFunction =
new MetaDataReaderFunction(
new Configuration(),
metaTable.schema(),
projectedSchema,
metaTable.io(),
metaTable.encryption());
this.splitSerializer = new IcebergSourceSplitSerializer(scanContext.caseSensitive());
}
@Override
public void processElement(
MetadataTablePlanner.SplitInfo splitInfo, Context ctx, Collector<R> out) throws Exception {
IcebergSourceSplit split = splitSerializer.deserialize(splitInfo.version(), splitInfo.split());
try (DataIterator<RowData> iterator = rowDataReaderFunction.createDataIterator(split)) {
iterator.forEachRemaining(rowData -> extract(rowData, out));
} catch (Exception e) {
LOG.warn("Exception processing split {} at {}", split, ctx.timestamp(), e);
ctx.output(DeleteOrphanFiles.ERROR_STREAM, e);
errorCounter.inc();
}
}
@Override
public void close() throws Exception {
super.close();
tableLoader.close();
}
/**
* Extracts the desired data from the given RowData.
*
* @param rowData the RowData from which to extract
* @param out the Collector to which to output the extracted data
*/
abstract void extract(RowData rowData, Collector<R> out);View on GitHub (pinned to 86d9c8fc54)