apache/iceberg · error · UncheckedIOException

Failed to process tasks iterable

Error message

Failed to process tasks iterable

What it means

FlinkSplitPlanner.planInputSplits wraps IOException raised while planning legacy FlinkInputSplits (iterating the CombinedScanTask and computing block/host locations via Util.blockLocations) into an UncheckedIOException. It signals that reading table metadata or file locations failed during split planning for the old Flink SourceFunction-style source.

Solutions

  1. Inspect the wrapped IOException cause to identify the unreadable file/location.
  2. Fix FileIO credentials/configuration for the table's storage (Hadoop conf, S3 endpoint/keys).
  3. Re-run planning against a stable snapshot (use a branch/tag or snapshot id) so files are not removed concurrently by maintenance jobs.
  4. Retry submission if the failure was a transient storage/network outage.

Example fix

// before (scan pinned to live HEAD, files expired mid-planning)
ScanContext ctx = ScanContext.builder().build();
// after (pin a stable snapshot/branch)
ScanContext ctx = ScanContext.builder().useBranch("audit-branch").build();
Defensive patterns

Strategy: validation

Validate before calling

// preflight: load the table and confirm all manifest files are readable
Table table = tableLoader.loadTable();
table.snapshots().forEach(s -> table.io().newInputFile(s.manifestListLocation()).getLength());

Try / catch

try {
  FlinkInputSplit[] splits = FlinkSplitPlanner.planInputSplits(table, context);
} catch (UncheckedIOException e) {
  LOG.error("split planning failed; check FileIO access and concurrent maintenance", e.getCause());
  throw e;
}

Prevention

When it happens

Trigger: Calling FlinkSplitPlanner.planInputSplits(table, context) when iterating scan tasks or resolving block locations (Util.blockLocations on table.io()) throws IOException — unreadable manifests, missing data files, or FileIO access errors.

Common situations: Storage credentials misconfigured for the FileIO; files deleted by concurrent snapshot expiry or compaction between scan and locality resolution; HDFS/S3 transient outages during job submission.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/c5464fe432844ba7. Report an issue: GitHub.

Appendix: source

Thrown at flink/v2.3/flink/src/main/java/org/apache/iceberg/flink/source/FlinkSplitPlanner.java:68

      List<CombinedScanTask> tasks = Lists.newArrayList(tasksIterable);
      FlinkInputSplit[] splits = new FlinkInputSplit[tasks.size()];
      boolean exposeLocality = context.exposeLocality();

      Tasks.range(tasks.size())
          .stopOnFailure()
          .executeWith(exposeLocality ? workerPool : null)
          .run(
              index -> {
                CombinedScanTask task = tasks.get(index);
                String[] hostnames = null;
                if (exposeLocality) {
                  hostnames = Util.blockLocations(table.io(), task);
                }
                splits[index] = new FlinkInputSplit(index, task, hostnames);
              });
      return splits;
    } catch (IOException e) {
      throw new UncheckedIOException("Failed to process tasks iterable", e);
    }
  }

  /** This returns splits for the FLIP-27 source */
  public static List<IcebergSourceSplit> planIcebergSourceSplits(
      Table table, ScanContext context, ExecutorService workerPool) {
    try (CloseableIterable<CombinedScanTask> tasksIterable =
        planTasks(table, context, workerPool)) {
      return Lists.newArrayList(
          CloseableIterable.transform(tasksIterable, IcebergSourceSplit::fromCombinedScanTask));
    } catch (IOException e) {
      throw new UncheckedIOException("Failed to process task iterable: ", e);
    }
  }

  static CloseableIterable<CombinedScanTask> planTasks(
      Table table, ScanContext context, ExecutorService workerPool) {
    ScanMode scanMode = checkScanMode(context);

View on GitHub (pinned to 86d9c8fc54)