apache/iceberg · error · UncheckedIOException

Failed to get file length

Error message

Failed to get file length

What it means

ParquetWriter.length() reports the bytes written so far plus the currently buffered size; it wraps any IOException from querying the underlying OutputFile/PositionOutputStream in an UncheckedIOException. This is called frequently (metrics, split-offset computation, commit-time file sizing), so a failure here means the underlying storage stream is broken or unreadable.

Solutions

  1. Inspect e.getCause() for the root IOException from the output stream
  2. Check network/connectivity to the storage backend (S3, GCS, HDFS NameNode)
  3. Verify the output file/stream has not been closed or removed by another process
  4. Retry the write job if the failure was transient network I/O
  5. Validate FileIO/OutputFile implementation handles position queries correctly

Example fix

// before: transient S3 failure makes length() blow up mid-task
long len = writer.length();

// after: treat as an I/O failure of the stream and fail the task with context
long len;
try {
  len = writer.length();
} catch (UncheckedIOException e) {
  throw new RuntimeException("Output stream for data file is broken: " + e.getCause().getMessage(), e);
}
Defensive patterns

Strategy: try-catch

Try / catch

try { long len = writer.length(); } catch (UncheckedIOException e) { throw new RuntimeException("Broken output stream: " + e.getCause().getMessage(), e); }

Prevention

When it happens

Trigger: Calling ParquetWriter.length() (directly or via checkSize/flushRowGroup/splitOffsets) while the underlying PositionOutputStream backed by the OutputFile throws IOException when queried for position/buffered size.

Common situations: Underlying local file handle closed or deleted mid-write; cloud storage (S3/GCS) stream failing on a position check due to network issues; HDFS lease recovery or client errors when querying the open file's length.

Understand the failure class

Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/da76ba81019cdc6e. Report an issue: GitHub.

Appendix: source

Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ParquetWriter.java:194

  @Override
  public long length() {
    try {
      long length = 0L;

      if (writer != null) {
        length += writer.getPos();
      }

      if (!closed && recordCount > 0) {
        // recordCount > 0 when there are records in the write store that have not been flushed to
        // the Parquet file
        length += writeStore.getBufferedSize();
      }

      return length;

    } catch (IOException e) {
      throw new UncheckedIOException("Failed to get file length", e);
    }
  }

  @Override
  public List<Long> splitOffsets() {
    if (writer != null) {
      return ParquetUtil.getSplitOffsets(writer.getFooter());
    }
    return null;
  }

  private void checkSize() {
    if (trackUncompressedSize) {
      if (rowGroupUncompressedSize >= targetRowGroupSize) {
        flushRowGroup(false);
      } else if (recordCount >= nextCheckRecordCount) {
        evaluateRowGroupSize(rowGroupUncompressedSize, false);
      }

View on GitHub (pinned to 86d9c8fc54)