apache/iceberg · error · ParquetDecodingException
could not read levels in page for col
Error message
could not read levels in page for col %s
What it means
BasePageIterator.newRLEIterator builds an RLE/bit-packing integer iterator for repetition or definition levels. If reading the level bytes throws IOException it is wrapped in ParquetDecodingException. This indicates the level data in the page is unreadable or corrupt.
Solutions
- Run a parquet validation tool on the file to confirm page-level corruption.
- Restore the file from a backup or rewrite the data from the original source.
- If corruption came from a specific writer version, re-encode affected files with a fixed writer.
Example fix
// before: scan fails on levels
// after: rewrite corrupted file from source
spark.read.parquet("good/source").writeTo("db.tbl").overwritePartitions(); Defensive patterns
Strategy: try-catch
Validate before calling
// Pre-check column chunk statistics/size vs expected to catch truncation
if (chunkMeta.getTotalSize() > (fileLength - chunkMeta.getStartOffset())) {
throw new IllegalStateException("Column chunk truncated for " + desc);
} Try / catch
try {
scanData();
} catch (ParquetDecodingException e) {
if (e.getMessage().contains("could not read levels")) {
// page level corruption: quarantine and rewrite file
}
} Prevention
- Detect truncation by comparing footer-declared chunk sizes with file length
- Restore damaged files from upstream before scanning
- Track writer versions and re-encode files from buggy writers
When it happens
Trigger: Called from initRepetitionLevelsReader/initDefinitionLevelsReader when decoding a page whose repetition/definition level stream cannot be read — corrupt level bytes, wrong width (maxLevel), or truncated buffer.
Common situations: Corrupted pages from faulty writers or truncated transfers; files damaged in storage; page-level corruption spotted only when scanning columns with optional/repeated fields.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- could not decode the dictionary for
- could not read page in col
- Failed to read a byte
- AlwaysFalse is a placeholder only
- AlwaysTrue is a placeholder only
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/aa69f042a0c13fcb.
Report an issue: GitHub.
Appendix: source
Thrown at parquet/src/main/java/org/apache/iceberg/parquet/BasePageIterator.java:187
this.delegate = delegate;
}
@Override
int nextInt() {
return delegate.readInteger();
}
}
IntIterator newRLEIterator(int maxLevel, BytesInput bytes) {
try {
if (maxLevel == 0) {
return new NullIntIterator();
}
return new RLEIntIterator(
new RunLengthBitPackingHybridDecoder(
BytesUtils.getWidthFromMaxInt(maxLevel), bytes.toInputStream()));
} catch (IOException e) {
throw new ParquetDecodingException("could not read levels in page for col " + desc, e);
}
}
static class RLEIntIterator extends IntIterator {
private final RunLengthBitPackingHybridDecoder delegate;
RLEIntIterator(RunLengthBitPackingHybridDecoder delegate) {
this.delegate = delegate;
}
@Override
int nextInt() {
try {
return delegate.readInt();
} catch (IOException e) {
throw new ParquetDecodingException(e);
}
}View on GitHub (pinned to 86d9c8fc54)