apache/iceberg · error · ParquetDecodingException
Failed to read bytes
Error message
Failed to read ${length} bytes What it means
ValuesAsBytesReader reads Parquet page values by slicing the underlying values input stream into little-endian ByteBuffers. getBuffer wraps any IOException from slicing into a ParquetDecodingException 'Failed to read N bytes', meaning the page ended before the requested number of bytes could be read — i.e. corrupt or truncated data.
Solutions
- Validate/repair the data file (re-copy from source, verify size and checksums).
- Rewrite the affected data files (e.g. rewrite_data_files / rewrite manifests) from a good source.
- If reproducible, capture the file and writer version and report upstream; work around by excluding the corrupt file from the scan.
Example fix
// before
TableScan scan = table.newScan(); // hits corrupt file
// after
Set<String> skip = Set.of("s3://bucket/path/corrupt-file.parquet");
TableScan scan = table.newScan().planWith(node -> !skip.contains(node.file().path())); Defensive patterns
Strategy: try-catch
Try / catch
try (CloseableIterable<Record> reader = Parquet.read(file).project(schema).build()) {
for (Record r : reader) { process(r); }
} catch (ParquetDecodingException e) {
if (e.getMessage().startsWith("Failed to read")) {
LOG.error("Corrupt/truncated Parquet page in {} — re-copy or rewrite the file", file.location(), e);
quarantine(file);
} else { throw e; }
} Prevention
- Enable checksums/verification on object storage and verify file sizes after transfer.
- Validate files after writing (e.g. footer read) before committing them to the table.
- Keep writer versions consistent across producers to avoid malformed pages.
- Monitor for ParquetDecodingException and route affected files to repair/rewrite pipelines.
When it happens
Trigger: Calling getBuffer via readInteger/readLong/readFloat/readDouble when the values input stream has fewer bytes remaining than requested (4 or 8) — a truncated/corrupt Parquet page or bad page offsets.
Common situations: Corrupt data files (network truncation, partial upload, disk corruption); writer bugs producing wrong page sizes; reading a file written by a buggy/older writer version; checksum-less storage hiding truncation.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- Can not decode bitwidth in block header
- Can not read min delta in current block
- Can't read value in column
- No more values to read. Total values read: " + valuesRead…
- not a valid mode " + this.mode
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/b3bf3ff9490f315a.
Report an issue: GitHub.
Appendix: source
Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ValuesAsBytesReader.java:54
private byte currentByte = 0;
public ValuesAsBytesReader() {}
@Override
public void initFromPage(int valueCount, ByteBufferInputStream in) {
this.valuesInputStream = in;
}
@Override
public void skip() {
throw new UnsupportedOperationException();
}
public ByteBuffer getBuffer(int length) {
try {
return valuesInputStream.slice(length).order(ByteOrder.LITTLE_ENDIAN);
} catch (IOException e) {
throw new ParquetDecodingException("Failed to read " + length + " bytes", e);
}
}
@Override
public final int readInteger() {
return getBuffer(4).getInt();
}
@Override
public final long readLong() {
return getBuffer(8).getLong();
}
@Override
public final float readFloat() {
return getBuffer(4).getFloat();
}
View on GitHub (pinned to 86d9c8fc54)