apache/iceberg · error · UncheckedIOException
Failed to create iceberg input splits for table:
Error message
Failed to create iceberg input splits for table:
What it means
FlinkSource.build throws this UncheckedIOException when estimating input-split count for source parallelism fails: format.createInputSplits(0) raised an IOException while listing/reading table files. The error message embeds the table being scanned. It surfaces during job construction because parallelism auto-computation requires materializing the splits.
Solutions
- Verify the table exists and its location is reachable with the configured FileIO (list the location manually).
- Check credentials/config for the object store or HDFS in the Flink cluster's environment (HADOOP_CONF_DIR, s3 access keys).
- Fix the catalog configuration (type, uri, warehouse) used to load the table.
- Set an explicit parallelism to skip the split-count-based auto-parallelism path, avoiding createInputSplits at planning time.
- Inspect the wrapped IOException cause for the precise filesystem error (not found vs permission).
Example fix
// before: relies on auto parallelism probing splits
FlinkSource.forRowData().env(env).table(table).build();
// after: pin parallelism to avoid split probing, or fix config first
FlinkSource.forRowData()
.env(env)
.table(table)
.parallelism(4)
.build(); Defensive patterns
Strategy: try-catch
Validate before calling
// preflight: ensure the table location is listable before building the source table.io().newInputFile(table.location()).getLength();
Try / catch
try {
DataStream<RowData> stream = FlinkSource.forRowData().env(env).table(table).build();
} catch (UncheckedIOException e) {
LOG.error("Cannot create splits for table {} — check catalog/warehouse/credentials", e);
throw e;
} Prevention
- Verify catalog/warehouse configuration and filesystem credentials on the Flink cluster before submitting jobs.
- Confirm the table location is reachable from jobmanager and taskmanagers.
- Set explicit source parallelism to avoid planning-time split materialization.
- Check the wrapped IOException cause to distinguish not-found vs permission errors.
When it happens
Trigger: Calling FlinkSource.forRowData()...build() (or SQL equivalents) where underlying FileIO cannot read manifest/data file locations — missing/incorrect warehouse path, expired credentials, deleted table location, or Hadoop/S3 misconfiguration.
Common situations: Misconfigured catalog or warehouse path, S3 credentials missing/expired on the jobmanager, table dropped between planning and build, or HDFS NameNode unreachable from the Flink cluster.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Class does not implement DynamicRecordGeneratorSQL
- Class does not implement DynamicRecordGeneratorSQL
- Exception listing files for
- Exception listing files for
- Exception listing files for
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/cd73ffd57aa4dc80.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/source/FlinkSource.java:295
public DataStream<RowData> build() {
Preconditions.checkNotNull(env, "StreamExecutionEnvironment should not be null");
FlinkInputFormat format = buildFormat();
ScanContext context = contextBuilder.build();
TypeInformation<RowData> typeInfo =
FlinkCompatibilityUtil.toTypeInfo(FlinkSchemaUtil.convert(context.project()));
if (!context.isStreaming()) {
int parallelism =
SourceUtil.inferParallelism(
readableConfig,
context.limit(),
() -> {
try {
return format.createInputSplits(0).length;
} catch (IOException e) {
throw new UncheckedIOException(
"Failed to create iceberg input splits for table: " + table, e);
}
});
if (env.getMaxParallelism() > 0) {
parallelism = Math.min(parallelism, env.getMaxParallelism());
}
return env.createInput(format, typeInfo).setParallelism(parallelism);
} else {
StreamingMonitorFunction function = new StreamingMonitorFunction(tableLoader, context);
String monitorFunctionName = String.format("Iceberg table (%s) monitor", table);
String readerOperatorName = String.format("Iceberg table (%s) reader", table);
return env.addSource(function, monitorFunctionName)
.transform(readerOperatorName, typeInfo, StreamingReaderOperator.factory(format));
}
}
}View on GitHub (pinned to 86d9c8fc54)