apache/iceberg · error · RuntimeIOException
Failed to find sync past position
Error message
Failed to find sync past position %d
What it means
AvroRangeIterator reads only a byte range [start, end) of an Avro file. On construction it calls reader.sync(start) to align to the first block boundary at or after start. An IOException here means the reader could not find a sync marker past the requested start position, so the range cannot be read reliably.
Solutions
- Recompute the split/range offsets from the current manifest (file content offsets in the DataFile), not cached values
- Validate that 0 <= start < file lengthInBytes before building the iterable
- Re-open the file and retry in case of transient read corruption during footer sync lookup
- Re-verify the file's integrity; if offsets no longer match the file, re-plan the scan
Example fix
// before: offsets from stale split AvroIterable<Record> it = Avro.read(in).project(schema).createReaderFunc(...).split(staleStart, length).build(); // after: validate against current file metadata long len = in.getLength(); long start = Math.min(splitStart, len); Preconditions.checkArgument(start >= 0 && start < len, "Invalid range start %s for file of length %s", start, len); AvroIterable<Record> it = Avro.read(in).project(schema).createReaderFunc(...).split(start, len - start).build();
Defensive patterns
Strategy: validation
Validate before calling
long len = inputFile.getLength();
Preconditions.checkArgument(start >= 0 && start < len,
"Range start %s out of bounds for file of length %s", start, len); Try / catch
try (AvroIterable<D> it = buildRangeIterable(start, end)) {
it.forEach(consumer);
} catch (RuntimeIOException e) {
replanAndReread();
} Prevention
- Derive split offsets only from the DataFile offsets recorded in the manifest
- Recompute splits after any table commit rather than caching offsets
- Clamp/validate start offsets against current file length
- Do not append to files managed by Iceberg with external writers
When it happens
Trigger: Constructing an AvroIterable with a start offset beyond or too close to the end of the file, or with a start value computed from stale metadata so that reader.sync() cannot locate a valid sync marker (truncated file, wrong offsets from a rewritten file).
Common situations: Task split offsets computed from an old snapshot after the file was compacted; passing start > file size due to a bug in split computation; reading a range of a file that was appended to by a non-Iceberg writer.
Understand the failure class
Background: "Must be a positive integer", "Invalid value", "Unsupported": the invalid-argument-value error family, when a library rejects the value you pass — this error's family across 35 libraries.
Related errors
- Cannot read manifest list file
- Cannot read manifest list file
- Decoding datum failed
- Failed to check range end
- Failed to close manifest reader
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/0844e3a6e023a4d5.
Report an issue: GitHub.
Appendix: source
Thrown at core/src/main/java/org/apache/iceberg/avro/AvroIterable.java:128
// Ignore close exception
}
}
throw new RuntimeIOException(e, "Failed to open file: %s", file.location());
}
}
private static class AvroRangeIterator<D> implements FileReader<D> {
private final FileReader<D> reader;
private final long end;
AvroRangeIterator(FileReader<D> reader, long start, long end) {
this.reader = reader;
this.end = end;
try {
reader.sync(start);
} catch (IOException e) {
throw new RuntimeIOException(e, "Failed to find sync past position %d", start);
}
}
@Override
public Schema getSchema() {
return reader.getSchema();
}
@Override
public boolean hasNext() {
try {
return reader.hasNext() && !reader.pastSync(end);
} catch (IOException e) {
throw new RuntimeIOException(e, "Failed to check range end: %d", end);
}
}
@OverrideView on GitHub (pinned to 86d9c8fc54)