apache/iceberg · error · UncheckedIOException
Failed to close table scan
Error message
Failed to close table scan: %s
What it means
IcebergInputFormat.planInputSplits closes the TableScan via Tasks and if closing throws IOException it wraps it, with the scan in the message, into an UncheckedIOException. The split planning itself succeeded or partially succeeded; the failure is in releasing scan resources.
Solutions
- Inspect the chained cause (getCause()) for the real IOException from the filesystem/FileIO and fix connectivity/credentials.
- Retry getSplits(); planning is typically retryable.
- Verify Hadoop/Hive conf and access permissions for the table location on the submitting client.
- Check for known FileIO issues with your storage plugin version.
Example fix
// before: failing silently surfaces wrapped scan-close error
List<InputSplit> splits = inputFormat.getSplits(jobContext);
// after: add retry/cause inspection
try {
splits = inputFormat.getSplits(jobContext);
} catch (UncheckedIOException e) {
LOG.error("scan close failed: {}", e.getCause(), e);
throw e;
} Defensive patterns
Strategy: try-catch
Validate before calling
// pre-check storage reachability before submitting the job
fs = new Path(tableLocation).getFileSystem(conf);
if (!fs.exists(new Path(tableLocation))) throw new IOException("table location unreachable"); Try / catch
try {
splits = inputFormat.getSplits(jobContext);
} catch (UncheckedIOException e) {
LOG.error("scan close failed: {}", e.getCause(), e);
throw new IOException(e.getCause());
} Prevention
- Inspect getCause() — the wrap hides the real FileIO/FileSystem error
- Ensure Hadoop credentials/kerberos are valid for the job submitter
- Retry split planning on transient storage errors
- Check NameNode/S3 endpoint reachability from the client node
When it happens
Trigger: TableScan.close() throwing IOException during getSplits() — typically an underlying FileIO/Hadoop filesystem error while releasing resources after listing manifests.
Common situations: HDFS/S3 connectivity blips during job submission; NameNode unreachable; credentials expiring mid-planning when the scan holds open resources.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Error occurred while processing
- Failed reading offset from
- Failed to close changelog scan:
- Failed to close scan: " + scan
- Failed to close task iterable
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/34d82b2b61437785.
Report an issue: GitHub.
Appendix: source
Thrown at mr/src/main/java/org/apache/iceberg/mr/mapreduce/IcebergInputFormat.java:149
}
// TODO add a filter parser to get rid of Serialization
Expression filter =
SerializationUtil.deserializeFromBase64(conf.get(InputFormatConfig.FILTER_EXPRESSION));
if (filter != null) {
scan = scan.filter(filter);
}
List<InputSplit> splits = Lists.newArrayList();
scan = scan.planWith(workerPool);
try (CloseableIterable<CombinedScanTask> tasksIterable = scan.planTasks()) {
Table serializableTable = SerializableTable.copyOf(table);
tasksIterable.forEach(
task -> {
splits.add(new IcebergSplit(serializableTable, conf, task));
});
} catch (IOException e) {
throw new UncheckedIOException(String.format("Failed to close table scan: %s", scan), e);
}
// if enabled, do not serialize FileIO hadoop config to decrease split size
// However, do not skip serialization for metatable queries, because some metadata tasks cache
// the IO object and we
// wouldn't be able to inject the config into these tasks on the deserializer-side, unlike for
// standard queries
if (scan instanceof DataTableScan) {
checkAndSkipIoConfigSerialization(conf, table);
}
return splits;
}
/**
* If enabled, it ensures that the FileIO's hadoop configuration will not be serialized. This
* might be desirable for decreasing the overall size of serialized table objects.
*View on GitHub (pinned to 86d9c8fc54)