apache/iceberg · error · RuntimeIOException
Failed to open Parquet file
Error message
Failed to open Parquet file: %s
What it means
ReadConf (the read-planning configuration for a Parquet scan) opens the data file with ParquetFileReader to read footer/metadata; if opening the file throws IOException it is rethrown as RuntimeIOException including the file location. This means the file could not be opened or its footer could not be read at all — most often the file is missing, unreadable, or not a valid Parquet file.
Solutions
- Check the file location in the message exists and is reachable with the configured FileIO/credentials
- Verify table metadata is not stale — run expire/remove orphan procedures if the file was deleted
- Confirm the file is a complete, valid Parquet file (footer present; not truncated by a failed write)
- Fix filesystem configuration (endpoint, bucket, HDFS core-site settings) in the read job
- If the file is corrupt, re-write the affected data files from a valid snapshot or re-ingest
Example fix
// before: scan fails on a deleted file referenced by old metadata
CloseableIterable<Record> rows = Parquet.read(file).project(schema).build();
// after: check existence and fail with an actionable message before planning
if (!io.newInputFile(file.location()).exists()) {
throw new FileNotFoundException("Data file missing from storage: " + file.location());
}
CloseableIterable<Record> rows = Parquet.read(file).project(schema).build(); Defensive patterns
Strategy: validation
Validate before calling
if (!io.newInputFile(location).exists()) {
throw new FileNotFoundException("Data file missing: " + location);
} Try / catch
try { return Parquet.read(file).project(schema).build(); } catch (RuntimeIOException e) { LOG.error("Cannot open {}: {}", file.location(), e.getCause().getMessage()); throw e; } Prevention
- Verify data files exist before scanning; run remove_orphan_files/expire_snapshots housekeeping
- Check filesystem configuration (endpoints, buckets, core-site) matches where data was written
- Confirm credentials grant read access to the data location
- Treat this error on a specific file as corruption/truncation — re-ingest or rewrite that file
When it happens
Trigger: Constructing a Parquet reader/ReadConf for a data file whose location cannot be opened by ParquetFileReader.open(ParquetIO.file(file), options) — nonexistent path, missing permissions, truncated/corrupt file, or invalid footer magic.
Common situations: Stale metadata pointing at deleted files (missing snapshot files); incorrect warehouse/filesystem configuration (wrong S3 bucket or HDFS path); partially-written files from a failed job; credentials lacking read access; file written by an incompatible Parquet version.
Understand the failure class
Background: "File not found" and ENOENT errors: why libraries can't find a file that should exist — this error's family across 50 libraries.
Related errors
- could not read page in col " + desc
- could not read page in col
- could not read page " + valueCount + " in col " + desc
- Error reading mini block.
- Failed to close changelog scan
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/ae1344085c9e0dd5.
Report an issue: GitHub.
Appendix: source
Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ReadConf.java:196
}
Integer batchSize() {
return batchSize;
}
List<Map<ColumnPath, ColumnChunkMetaData>> columnChunkMetadataForRowGroups() {
return columnChunkMetaDataForRowGroups;
}
ReadConf<T> copy() {
return new ReadConf<>(this);
}
private static ParquetFileReader newReader(InputFile file, ParquetReadOptions options) {
try {
return ParquetFileReader.open(ParquetIO.file(file), options);
} catch (IOException e) {
throw new RuntimeIOException(e, "Failed to open Parquet file: %s", file.location());
}
}
private List<Map<ColumnPath, ColumnChunkMetaData>> getColumnChunkMetadataForRowGroups() {
Set<ColumnPath> projectedColumns =
projection.getColumns().stream()
.map(columnDescriptor -> ColumnPath.get(columnDescriptor.getPath()))
.collect(Collectors.toSet());
ImmutableList.Builder<Map<ColumnPath, ColumnChunkMetaData>> listBuilder =
ImmutableList.builder();
for (int i = 0; i < rowGroups.size(); i++) {
if (!shouldSkip[i]) {
BlockMetaData blockMetaData = rowGroups.get(i);
ImmutableMap.Builder<ColumnPath, ColumnChunkMetaData> mapBuilder = ImmutableMap.builder();
blockMetaData.getColumns().stream()
.filter(columnChunkMetaData -> projectedColumns.contains(columnChunkMetaData.getPath()))
.forEach(
columnChunkMetaData ->View on GitHub (pinned to 86d9c8fc54)