apache/iceberg · error · UncheckedIOException
Failed to create iceberg input splits for table
Error message
Failed to create iceberg input splits for table: {table} What it means
FlinkSource.build() probes the number of input splits (via format.createInputSplits(0).length) to derive a suggested source parallelism. Any IOException from split creation is wrapped in an UncheckedIOException with this message, meaning the table's files could not be listed/opened during split planning.
Solutions
- Verify the catalog/warehouse path and table location exist and are accessible
- Check storage credentials (S3 AK/SK, IAM, HDFS tokens) are valid and not expired
- Confirm the underlying filesystem (HDFS NameNode, S3 endpoint) is reachable
- Test TableScan planning outside Flink with the same Hadoop conf to isolate the IO failure
Example fix
// before DataStream<RowData> stream = FlinkSource.forRowData().project(schema).tableLoader(loader).build(); // after Table table = loader.loadTable(); table.scan().planTasks(); // fails fast with a clearer IO error if storage is unreachable DataStream<RowData> stream = FlinkSource.forRowData().project(schema).tableLoader(loader).build();
Defensive patterns
Strategy: try-catch
Validate before calling
Table table = loader.loadTable(); table.scan().planTasks(); // fail fast with clearer IO error before build()
Try / catch
try {
DataStream<RowData> ds = FlinkSource.forRowData()...build();
} catch (UncheckedIOException e) {
throw new RuntimeException("Check storage credentials/path: " + e.getCause(), e);
} Prevention
- Pre-validate table reachability with a small TableScan before job submission
- Keep cloud credentials (S3/IAM) refreshed for long-running clusters
- Confirm warehouse path and catalog config point to an existing table
When it happens
Trigger: Calling FlinkSource.forRowData()...build() where table.io() fails listing data files or opening FileIO — e.g. missing/invalid warehouse path, deleted table location, expired/insufficient S3/HDFS credentials, or unreachable storage.
Common situations: Wrong catalog/warehouse configuration pointing at a nonexistent location; cloud credential expiry on long-running clusters; HDFS NameNode unavailability; table dropped between planning and build.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Failed to list partitions of table
- Failed to read manifest:
- Failed to read manifest file
- Failed to write manifest
- Location does not exist
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/354db766baf5dbc4.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.1/flink/src/main/java/org/apache/iceberg/flink/source/FlinkSource.java:295
public DataStream<RowData> build() {
Preconditions.checkNotNull(env, "StreamExecutionEnvironment should not be null");
FlinkInputFormat format = buildFormat();
ScanContext context = contextBuilder.build();
TypeInformation<RowData> typeInfo =
FlinkCompatibilityUtil.toTypeInfo(FlinkSchemaUtil.convert(context.project()));
if (!context.isStreaming()) {
int parallelism =
SourceUtil.inferParallelism(
readableConfig,
context.limit(),
() -> {
try {
return format.createInputSplits(0).length;
} catch (IOException e) {
throw new UncheckedIOException(
"Failed to create iceberg input splits for table: " + table, e);
}
});
if (env.getMaxParallelism() > 0) {
parallelism = Math.min(parallelism, env.getMaxParallelism());
}
return env.createInput(format, typeInfo).setParallelism(parallelism);
} else {
StreamingMonitorFunction function = new StreamingMonitorFunction(tableLoader, context);
String monitorFunctionName = String.format("Iceberg table (%s) monitor", table);
String readerOperatorName = String.format("Iceberg table (%s) reader", table);
return env.addSource(function, monitorFunctionName)
.transform(readerOperatorName, typeInfo, StreamingReaderOperator.factory(format));
}
}
}View on GitHub (pinned to 86d9c8fc54)