apache/iceberg · error · RuntimeIOException

Failed to open Parquet file

Error message

Failed to open Parquet file: %s

What it means

ReadConf (the read-planning configuration for a Parquet scan) opens the data file with ParquetFileReader to read footer/metadata; if opening the file throws IOException it is rethrown as RuntimeIOException including the file location. This means the file could not be opened or its footer could not be read at all — most often the file is missing, unreadable, or not a valid Parquet file.

Solutions

  1. Check the file location in the message exists and is reachable with the configured FileIO/credentials
  2. Verify table metadata is not stale — run expire/remove orphan procedures if the file was deleted
  3. Confirm the file is a complete, valid Parquet file (footer present; not truncated by a failed write)
  4. Fix filesystem configuration (endpoint, bucket, HDFS core-site settings) in the read job
  5. If the file is corrupt, re-write the affected data files from a valid snapshot or re-ingest

Example fix

// before: scan fails on a deleted file referenced by old metadata
CloseableIterable<Record> rows = Parquet.read(file).project(schema).build();

// after: check existence and fail with an actionable message before planning
if (!io.newInputFile(file.location()).exists()) {
  throw new FileNotFoundException("Data file missing from storage: " + file.location());
}
CloseableIterable<Record> rows = Parquet.read(file).project(schema).build();
Defensive patterns

Strategy: validation

Validate before calling

if (!io.newInputFile(location).exists()) {
  throw new FileNotFoundException("Data file missing: " + location);
}

Try / catch

try { return Parquet.read(file).project(schema).build(); } catch (RuntimeIOException e) { LOG.error("Cannot open {}: {}", file.location(), e.getCause().getMessage()); throw e; }

Prevention

When it happens

Trigger: Constructing a Parquet reader/ReadConf for a data file whose location cannot be opened by ParquetFileReader.open(ParquetIO.file(file), options) — nonexistent path, missing permissions, truncated/corrupt file, or invalid footer magic.

Common situations: Stale metadata pointing at deleted files (missing snapshot files); incorrect warehouse/filesystem configuration (wrong S3 bucket or HDFS path); partially-written files from a failed job; credentials lacking read access; file written by an incompatible Parquet version.

Understand the failure class

Background: "File not found" and ENOENT errors: why libraries can't find a file that should exist — this error's family across 50 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/ae1344085c9e0dd5. Report an issue: GitHub.

Appendix: source

Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ReadConf.java:196

  }

  Integer batchSize() {
    return batchSize;
  }

  List<Map<ColumnPath, ColumnChunkMetaData>> columnChunkMetadataForRowGroups() {
    return columnChunkMetaDataForRowGroups;
  }

  ReadConf<T> copy() {
    return new ReadConf<>(this);
  }

  private static ParquetFileReader newReader(InputFile file, ParquetReadOptions options) {
    try {
      return ParquetFileReader.open(ParquetIO.file(file), options);
    } catch (IOException e) {
      throw new RuntimeIOException(e, "Failed to open Parquet file: %s", file.location());
    }
  }

  private List<Map<ColumnPath, ColumnChunkMetaData>> getColumnChunkMetadataForRowGroups() {
    Set<ColumnPath> projectedColumns =
        projection.getColumns().stream()
            .map(columnDescriptor -> ColumnPath.get(columnDescriptor.getPath()))
            .collect(Collectors.toSet());
    ImmutableList.Builder<Map<ColumnPath, ColumnChunkMetaData>> listBuilder =
        ImmutableList.builder();
    for (int i = 0; i < rowGroups.size(); i++) {
      if (!shouldSkip[i]) {
        BlockMetaData blockMetaData = rowGroups.get(i);
        ImmutableMap.Builder<ColumnPath, ColumnChunkMetaData> mapBuilder = ImmutableMap.builder();
        blockMetaData.getColumns().stream()
            .filter(columnChunkMetaData -> projectedColumns.contains(columnChunkMetaData.getPath()))
            .forEach(
                columnChunkMetaData ->

View on GitHub (pinned to 86d9c8fc54)