apache/iceberg · error · UncheckedIOException
Failed to get file length
Error message
Failed to get file length
What it means
ParquetWriter.length() reports the bytes written so far plus the currently buffered size; it wraps any IOException from querying the underlying OutputFile/PositionOutputStream in an UncheckedIOException. This is called frequently (metrics, split-offset computation, commit-time file sizing), so a failure here means the underlying storage stream is broken or unreadable.
Solutions
- Inspect e.getCause() for the root IOException from the output stream
- Check network/connectivity to the storage backend (S3, GCS, HDFS NameNode)
- Verify the output file/stream has not been closed or removed by another process
- Retry the write job if the failure was transient network I/O
- Validate FileIO/OutputFile implementation handles position queries correctly
Example fix
// before: transient S3 failure makes length() blow up mid-task
long len = writer.length();
// after: treat as an I/O failure of the stream and fail the task with context
long len;
try {
len = writer.length();
} catch (UncheckedIOException e) {
throw new RuntimeException("Output stream for data file is broken: " + e.getCause().getMessage(), e);
} Defensive patterns
Strategy: try-catch
Try / catch
try { long len = writer.length(); } catch (UncheckedIOException e) { throw new RuntimeException("Broken output stream: " + e.getCause().getMessage(), e); } Prevention
- Ensure network stability to S3/GCS/HDFS during long write tasks; enable retries in the filesystem client
- Never close or delete the output file from another thread while writing
- Verify your FileIO's PositionOutputStream implements position queries correctly
- Treat UncheckedIOException from length() as a stream-level failure, not a metrics bug
When it happens
Trigger: Calling ParquetWriter.length() (directly or via checkSize/flushRowGroup/splitOffsets) while the underlying PositionOutputStream backed by the OutputFile throws IOException when queried for position/buffered size.
Common situations: Underlying local file handle closed or deleted mid-write; cloud storage (S3/GCS) stream failing on a position check due to network issues; HDFS lease recovery or client errors when querying the open file's length.
Understand the failure class
Background: "failed to read file", EACCES, ENOENT and "could not read <path>" errors: when a program can't read a file from disk — this error's family across 49 libraries.
Related errors
- could not read page in col " + desc
- could not read page in col
- could not read page " + valueCount + " in col " + desc
- Error reading mini block.
- Failed to create Parquet file
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/da76ba81019cdc6e.
Report an issue: GitHub.
Appendix: source
Thrown at parquet/src/main/java/org/apache/iceberg/parquet/ParquetWriter.java:194
@Override
public long length() {
try {
long length = 0L;
if (writer != null) {
length += writer.getPos();
}
if (!closed && recordCount > 0) {
// recordCount > 0 when there are records in the write store that have not been flushed to
// the Parquet file
length += writeStore.getBufferedSize();
}
return length;
} catch (IOException e) {
throw new UncheckedIOException("Failed to get file length", e);
}
}
@Override
public List<Long> splitOffsets() {
if (writer != null) {
return ParquetUtil.getSplitOffsets(writer.getFooter());
}
return null;
}
private void checkSize() {
if (trackUncompressedSize) {
if (rowGroupUncompressedSize >= targetRowGroupSize) {
flushRowGroup(false);
} else if (recordCount >= nextCheckRecordCount) {
evaluateRowGroupSize(rowGroupUncompressedSize, false);
}View on GitHub (pinned to 86d9c8fc54)