apache/iceberg · warning
Exception processing split
Error message
Exception processing split {} at {} What it means
TableReader processes metadata-table splits (e.g. from the files/manifests metadata table) to feed maintenance actions like orphan-file detection. If deserializing the split or reading its rows fails, the exception is logged with the split and timestamp, forwarded to the ERROR_STREAM side output, and counted instead of failing the job.
Solutions
- Inspect the ERROR_STREAM side output and errorCounter to identify which splits failed and why
- Re-run the maintenance action after fixing the underlying IO/snapshot issue so missed files are re-scanned
- Avoid running expireSnapshots concurrently with orphan-file detection reading the same table
- Check that the job was restored with compatible schema/table state; restart from a fresh savepoint if not
Example fix
// after
DataStream<RowData> errorStream = result.getSideOutput(DeleteOrphanFiles.ERROR_STREAM);
errorStream.addSink(e -> LOG.error("maintenance read failed", e)); Defensive patterns
Strategy: try-catch
Try / catch
result.getSideOutput(DeleteOrphanFiles.ERROR_STREAM)
.addSink(e -> log.error("Split processing failed, re-run maintenance action", e));
// monitor errorCounter and re-run the action once the underlying issue is fixed Prevention
- Do not run expireSnapshots/orphan cleanup concurrently with metadata-table readers
- Check the ERROR_STREAM side output and error counter in production monitoring
- Restore jobs from savepoints created with the same Iceberg version and schema
When it happens
Trigger: processElement receives a SplitInfo whose IcebergSourceSplit cannot be deserialized or whose data iterator throws (corrupt/missing data files, schema mismatch, IO errors) while a DeleteOrphanFiles/ExpireSnapshots pipeline reads the corresponding metadata table.
Common situations: Files deleted or corrupted between planning and reading (e.g. concurrent expireSnapshots); schema evolution between job versions and stored split state; transient S3/HDFS read failures; checkpoint-restored splits referencing removed snapshots.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Can not alter the default database when the iceberg catalog…
- Cannot apply unknown modify-column change:
- Cannot apply unknown modify-column-position change:
- Cannot apply unknown table change:
- Cannot apply unknown unique constraint:
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/8c6d4fb7a754ea2e.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.1/flink/src/main/java/org/apache/iceberg/flink/maintenance/operator/TableReader.java:101
.counter(TableMaintenanceMetrics.ERROR_COUNTER);
this.rowDataReaderFunction =
new MetaDataReaderFunction(
new Configuration(),
metaTable.schema(),
projectedSchema,
metaTable.io(),
metaTable.encryption());
this.splitSerializer = new IcebergSourceSplitSerializer(scanContext.caseSensitive());
}
@Override
public void processElement(
MetadataTablePlanner.SplitInfo splitInfo, Context ctx, Collector<R> out) throws Exception {
IcebergSourceSplit split = splitSerializer.deserialize(splitInfo.version(), splitInfo.split());
try (DataIterator<RowData> iterator = rowDataReaderFunction.createDataIterator(split)) {
iterator.forEachRemaining(rowData -> extract(rowData, out));
} catch (Exception e) {
LOG.warn("Exception processing split {} at {}", split, ctx.timestamp(), e);
ctx.output(DeleteOrphanFiles.ERROR_STREAM, e);
errorCounter.inc();
}
}
@Override
public void close() throws Exception {
super.close();
tableLoader.close();
}
/**
* Extracts the desired data from the given RowData.
*
* @param rowData the RowData from which to extract
* @param out the Collector to which to output the extracted data
*/
abstract void extract(RowData rowData, Collector<R> out);View on GitHub (pinned to 86d9c8fc54)