apache/iceberg · error · UncheckedIOException
Failed to create iceberg input splits for table: " + table
Error message
Failed to create iceberg input splits for table: " + table
What it means
FlinkSource.build() wraps IOException from FormatModel.createInputSplits(0) — used to count splits when auto-computing parallelism from the number of scan splits — into an UncheckedIOException with the table name. It means the underlying table scan could not produce input splits during job construction.
Solutions
- Check the wrapped IOException (e.getCause()) — it names the file or location that failed to read during split planning.
- Verify FileIO credentials/permissions for the table location (S3 access keys, HDFS delegation tokens).
- Confirm the target snapshot still exists; if concurrent expiration deleted it, point the scan to a valid snapshot or branch/tag.
- Set an explicit scan.parallelism to bypass split-count-based parallelism inference if planning is transiently failing.
Example fix
// before new FlinkSource.ForRowData().tableLoader(loader).build(); // parallelism inferred, triggers split planning // after new FlinkSource.ForRowData().tableLoader(loader).project(schema).configureScanParallelism(4).build();
Defensive patterns
Strategy: validation
Validate before calling
// verify table location readability and snapshot existence before building the source
Table table = tableLoader.loadTable();
Snapshot snap = table.currentSnapshot();
if (snap == null) throw new IllegalStateException("table has no current snapshot");
table.io().newInputFile(snap.manifestListLocation()).getLength(); // proactively read Try / catch
try {
DataStream<RowData> ds = env.fromSource(source, watermark, "iceberg");
} catch (UncheckedIOException e) {
if (e.getMessage().contains("Failed to create iceberg input splits")) {
LOG.error("split planning failed for table; check IO credentials/snapshot", e.getCause());
}
throw e;
} Prevention
- Validate FileIO credentials (S3/HDFS) on the submission node before submitting jobs.
- Avoid running snapshot expiry concurrently with job startup.
- Set explicit scan.parallelism to skip split-count inference when planning is flaky.
When it happens
Trigger: Building a FlinkSource with scan.parallelism unset (so parallelism is inferred from split count) while createInputSplits throws IOException — e.g. unreadable metadata files, missing/invalid manifest list, or underlying file system access failure during scan planning.
Common situations: Table metadata unreadable due to storage permission problems (S3/HDFS credentials); snapshot referenced by the scan no longer exists after concurrent expiry; network interruption to object storage during planning.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Failed to process tasks iterable
- Cannot read data task.
- Exception planning scan for
- Exception planning scan for
- Exception planning scan for
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/06d7d27923db7dc1.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.3/flink/src/main/java/org/apache/iceberg/flink/source/FlinkSource.java:295
public DataStream<RowData> build() {
Preconditions.checkNotNull(env, "StreamExecutionEnvironment should not be null");
FlinkInputFormat format = buildFormat();
ScanContext context = contextBuilder.build();
TypeInformation<RowData> typeInfo =
FlinkCompatibilityUtil.toTypeInfo(FlinkSchemaUtil.convert(context.project()));
if (!context.isStreaming()) {
int parallelism =
SourceUtil.inferParallelism(
readableConfig,
context.limit(),
() -> {
try {
return format.createInputSplits(0).length;
} catch (IOException e) {
throw new UncheckedIOException(
"Failed to create iceberg input splits for table: " + table, e);
}
});
if (env.getMaxParallelism() > 0) {
parallelism = Math.min(parallelism, env.getMaxParallelism());
}
return env.createInput(format, typeInfo).setParallelism(parallelism);
} else {
StreamingMonitorFunction function = new StreamingMonitorFunction(tableLoader, context);
String monitorFunctionName = String.format("Iceberg table (%s) monitor", table);
String readerOperatorName = String.format("Iceberg table (%s) reader", table);
return env.addSource(function, monitorFunctionName)
.transform(readerOperatorName, typeInfo, StreamingReaderOperator.factory(format));
}
}
}View on GitHub (pinned to 86d9c8fc54)