{"record":{"id":"b3bf3ff9490f315a","repo":"apache/iceberg","slug":"failed-to-read-length-bytes","errorCode":null,"errorMessage":"Failed to read ${length} bytes","messagePattern":"Failed to read (.+?) bytes","errorType":"exception","errorClass":"ParquetDecodingException","httpStatus":null,"severity":"error","filePath":"parquet/src/main/java/org/apache/iceberg/parquet/ValuesAsBytesReader.java","lineNumber":54,"sourceCode":"  private byte currentByte = 0;\n\n  public ValuesAsBytesReader() {}\n\n  @Override\n  public void initFromPage(int valueCount, ByteBufferInputStream in) {\n    this.valuesInputStream = in;\n  }\n\n  @Override\n  public void skip() {\n    throw new UnsupportedOperationException();\n  }\n\n  public ByteBuffer getBuffer(int length) {\n    try {\n      return valuesInputStream.slice(length).order(ByteOrder.LITTLE_ENDIAN);\n    } catch (IOException e) {\n      throw new ParquetDecodingException(\"Failed to read \" + length + \" bytes\", e);\n    }\n  }\n\n  @Override\n  public final int readInteger() {\n    return getBuffer(4).getInt();\n  }\n\n  @Override\n  public final long readLong() {\n    return getBuffer(8).getLong();\n  }\n\n  @Override\n  public final float readFloat() {\n    return getBuffer(4).getFloat();\n  }\n","sourceCodeStart":36,"sourceCodeEnd":72,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/parquet/src/main/java/org/apache/iceberg/parquet/ValuesAsBytesReader.java#L36-L72","documentation":"ValuesAsBytesReader reads Parquet page values by slicing the underlying values input stream into little-endian ByteBuffers. getBuffer wraps any IOException from slicing into a ParquetDecodingException 'Failed to read N bytes', meaning the page ended before the requested number of bytes could be read — i.e. corrupt or truncated data.","triggerScenarios":"Calling getBuffer via readInteger/readLong/readFloat/readDouble when the values input stream has fewer bytes remaining than requested (4 or 8) — a truncated/corrupt Parquet page or bad page offsets.","commonSituations":"Corrupt data files (network truncation, partial upload, disk corruption); writer bugs producing wrong page sizes; reading a file written by a buggy/older writer version; checksum-less storage hiding truncation.","solutions":["Validate/repair the data file (re-copy from source, verify size and checksums).","Rewrite the affected data files (e.g. rewrite_data_files / rewrite manifests) from a good source.","If reproducible, capture the file and writer version and report upstream; work around by excluding the corrupt file from the scan."],"exampleFix":"// before\nTableScan scan = table.newScan(); // hits corrupt file\n// after\nSet<String> skip = Set.of(\"s3://bucket/path/corrupt-file.parquet\");\nTableScan scan = table.newScan().planWith(node -> !skip.contains(node.file().path()));","handlingStrategy":"try-catch","validationCode":null,"typeGuard":null,"tryCatchPattern":"try (CloseableIterable<Record> reader = Parquet.read(file).project(schema).build()) {\n  for (Record r : reader) { process(r); }\n} catch (ParquetDecodingException e) {\n  if (e.getMessage().startsWith(\"Failed to read\")) {\n    LOG.error(\"Corrupt/truncated Parquet page in {} — re-copy or rewrite the file\", file.location(), e);\n    quarantine(file);\n  } else { throw e; }\n}","preventionTips":["Enable checksums/verification on object storage and verify file sizes after transfer.","Validate files after writing (e.g. footer read) before committing them to the table.","Keep writer versions consistent across producers to avoid malformed pages.","Monitor for ParquetDecodingException and route affected files to repair/rewrite pipelines."],"tags":["parquet","decoding","corrupt-data"],"backgroundTag":"file-read-failed","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}