apache/iceberg · warning

Exception processing split

Error message

Exception processing split {} at {}

What it means

TableReader.processElement deserializes an IcebergSourceSplit and iterates its data with a DataIterator, forwarding rows downstream. If reading the split throws, the exception is logged as a warning with the split and timestamp, pushed to the DeleteOrphanFiles ERROR_STREAM, and counted — the operator continues with the next split instead of failing the job, so orphan-file detection proceeds with partial data.

Solutions

  1. Inspect the exception on the DeleteOrphanFiles.ERROR_STREAM side output for the root cause
  2. Ensure no concurrent snapshot-expiration or orphan-cleanup job deletes files while the maintenance job reads them
  3. Align Iceberg versions across the Flink job and cluster to avoid split serialization incompatibilities
  4. Retry the maintenance job; failed splits reduce accuracy of that cycle but do not corrupt state
Defensive patterns

Strategy: retry

Try / catch

// Read the forwarded exception for diagnosis
result.getSideOutput(DeleteOrphanFiles.ERROR_STREAM)
      .executeAndCollect().forEachRemaining(e -> log.error("split read failed", e));

Prevention

When it happens

Trigger: Raised in processElement when splitSerializer.deserialize or rowDataReaderFunction.createDataIterator(split) / iteration throws — file deleted between planning and reading (expired snapshots/orphan cleanup), schema mismatch against the metadata-table schema, or object-store read errors.

Common situations: Another job expired snapshots and deleted data/metadata files mid-scan; serialization version mismatch after an Iceberg upgrade of one part of the cluster; transient S3/HDFS read failures; corrupt Parquet/metadata files.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/db1ee03235421b8d. Report an issue: GitHub.

Appendix: source

Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/maintenance/operator/TableReader.java:101

            .counter(TableMaintenanceMetrics.ERROR_COUNTER);
    this.rowDataReaderFunction =
        new MetaDataReaderFunction(
            new Configuration(),
            metaTable.schema(),
            projectedSchema,
            metaTable.io(),
            metaTable.encryption());
    this.splitSerializer = new IcebergSourceSplitSerializer(scanContext.caseSensitive());
  }

  @Override
  public void processElement(
      MetadataTablePlanner.SplitInfo splitInfo, Context ctx, Collector<R> out) throws Exception {
    IcebergSourceSplit split = splitSerializer.deserialize(splitInfo.version(), splitInfo.split());
    try (DataIterator<RowData> iterator = rowDataReaderFunction.createDataIterator(split)) {
      iterator.forEachRemaining(rowData -> extract(rowData, out));
    } catch (Exception e) {
      LOG.warn("Exception processing split {} at {}", split, ctx.timestamp(), e);
      ctx.output(DeleteOrphanFiles.ERROR_STREAM, e);
      errorCounter.inc();
    }
  }

  @Override
  public void close() throws Exception {
    super.close();
    tableLoader.close();
  }

  /**
   * Extracts the desired data from the given RowData.
   *
   * @param rowData the RowData from which to extract
   * @param out the Collector to which to output the extracted data
   */
  abstract void extract(RowData rowData, Collector<R> out);

View on GitHub (pinned to 86d9c8fc54)