apache/iceberg · error · ParquetDecodingException

Failed to read bytes

Error message

Failed to read ${length} bytes

What it means

ValuesAsBytesReader reads Parquet page values by slicing the underlying values input stream into little-endian ByteBuffers. getBuffer wraps any IOException from slicing into a ParquetDecodingException 'Failed to read N bytes', meaning the page ended before the requested number of bytes could be read — i.e. corrupt or truncated data.

Solutions

  1. Validate/repair the data file (re-copy from source, verify size and checksums).
  2. Rewrite the affected data files (e.g. rewrite_data_files / rewrite manifests) from a good source.
  3. If reproducible, capture the file and writer version and report upstream; work around by excluding the corrupt file from the scan.

Example fix

// before
TableScan scan = table.newScan(); // hits corrupt file
// after
Set<String> skip = Set.of("s3://bucket/path/corrupt-file.parquet");
TableScan scan = table.newScan().planWith(node -> !skip.contains(node.file().path()));
Defensive patterns

Strategy: try-catch

Try / catch

try (CloseableIterable<Record> reader = Parquet.read(file).project(schema).build()) {
  for (Record r : reader) { process(r); }
} catch (ParquetDecodingException e) {
  if (e.getMessage().startsWith("Failed to read")) {
    LOG.error("Corrupt/truncated Parquet page in {} — re-copy or rewrite the file", file.location(), e);
    quarantine(file);
  } else { throw e; }
}

Prevention

When it happens

Trigger: Calling getBuffer via readInteger/readLong/readFloat/readDouble when the values input stream has fewer bytes remaining than requested (4 or 8) — a truncated/corrupt Parquet page or bad page offsets.

Common situations: Corrupt data files (network truncation, partial upload, disk corruption); writer bugs producing wrong page sizes; reading a file written by a buggy/older writer version; checksum-less storage hiding truncation.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/b3bf3ff9490f315a. Report an issue: GitHub.

Appendix: source

Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ValuesAsBytesReader.java:54

  private byte currentByte = 0;

  public ValuesAsBytesReader() {}

  @Override
  public void initFromPage(int valueCount, ByteBufferInputStream in) {
    this.valuesInputStream = in;
  }

  @Override
  public void skip() {
    throw new UnsupportedOperationException();
  }

  public ByteBuffer getBuffer(int length) {
    try {
      return valuesInputStream.slice(length).order(ByteOrder.LITTLE_ENDIAN);
    } catch (IOException e) {
      throw new ParquetDecodingException("Failed to read " + length + " bytes", e);
    }
  }

  @Override
  public final int readInteger() {
    return getBuffer(4).getInt();
  }

  @Override
  public final long readLong() {
    return getBuffer(8).getLong();
  }

  @Override
  public final float readFloat() {
    return getBuffer(4).getFloat();
  }

View on GitHub (pinned to 86d9c8fc54)