apache/iceberg · error · UncheckedIOException

Failed to create iceberg input splits for table:

Error message

Failed to create iceberg input splits for table: 

What it means

FlinkSource.build throws this UncheckedIOException when estimating input-split count for source parallelism fails: format.createInputSplits(0) raised an IOException while listing/reading table files. The error message embeds the table being scanned. It surfaces during job construction because parallelism auto-computation requires materializing the splits.

Solutions

  1. Verify the table exists and its location is reachable with the configured FileIO (list the location manually).
  2. Check credentials/config for the object store or HDFS in the Flink cluster's environment (HADOOP_CONF_DIR, s3 access keys).
  3. Fix the catalog configuration (type, uri, warehouse) used to load the table.
  4. Set an explicit parallelism to skip the split-count-based auto-parallelism path, avoiding createInputSplits at planning time.
  5. Inspect the wrapped IOException cause for the precise filesystem error (not found vs permission).

Example fix

// before: relies on auto parallelism probing splits
FlinkSource.forRowData().env(env).table(table).build();
// after: pin parallelism to avoid split probing, or fix config first
FlinkSource.forRowData()
    .env(env)
    .table(table)
    .parallelism(4)
    .build();
Defensive patterns

Strategy: try-catch

Validate before calling

// preflight: ensure the table location is listable before building the source
table.io().newInputFile(table.location()).getLength();

Try / catch

try {
  DataStream<RowData> stream = FlinkSource.forRowData().env(env).table(table).build();
} catch (UncheckedIOException e) {
  LOG.error("Cannot create splits for table {} — check catalog/warehouse/credentials", e);
  throw e;
}

Prevention

When it happens

Trigger: Calling FlinkSource.forRowData()...build() (or SQL equivalents) where underlying FileIO cannot read manifest/data file locations — missing/incorrect warehouse path, expired credentials, deleted table location, or Hadoop/S3 misconfiguration.

Common situations: Misconfigured catalog or warehouse path, S3 credentials missing/expired on the jobmanager, table dropped between planning and build, or HDFS NameNode unreachable from the Flink cluster.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/cd73ffd57aa4dc80. Report an issue: GitHub.

Appendix: source

Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/source/FlinkSource.java:295

    public DataStream<RowData> build() {
      Preconditions.checkNotNull(env, "StreamExecutionEnvironment should not be null");
      FlinkInputFormat format = buildFormat();

      ScanContext context = contextBuilder.build();
      TypeInformation<RowData> typeInfo =
          FlinkCompatibilityUtil.toTypeInfo(FlinkSchemaUtil.convert(context.project()));

      if (!context.isStreaming()) {
        int parallelism =
            SourceUtil.inferParallelism(
                readableConfig,
                context.limit(),
                () -> {
                  try {
                    return format.createInputSplits(0).length;
                  } catch (IOException e) {
                    throw new UncheckedIOException(
                        "Failed to create iceberg input splits for table: " + table, e);
                  }
                });
        if (env.getMaxParallelism() > 0) {
          parallelism = Math.min(parallelism, env.getMaxParallelism());
        }
        return env.createInput(format, typeInfo).setParallelism(parallelism);
      } else {
        StreamingMonitorFunction function = new StreamingMonitorFunction(tableLoader, context);

        String monitorFunctionName = String.format("Iceberg table (%s) monitor", table);
        String readerOperatorName = String.format("Iceberg table (%s) reader", table);

        return env.addSource(function, monitorFunctionName)
            .transform(readerOperatorName, typeInfo, StreamingReaderOperator.factory(format));
      }
    }
  }

View on GitHub (pinned to 86d9c8fc54)