apache/iceberg · error · RuntimeIOException

Failed to find sync past position

Error message

Failed to find sync past position %d

What it means

AvroRangeIterator reads only a byte range [start, end) of an Avro file. On construction it calls reader.sync(start) to align to the first block boundary at or after start. An IOException here means the reader could not find a sync marker past the requested start position, so the range cannot be read reliably.

Solutions

  1. Recompute the split/range offsets from the current manifest (file content offsets in the DataFile), not cached values
  2. Validate that 0 <= start < file lengthInBytes before building the iterable
  3. Re-open the file and retry in case of transient read corruption during footer sync lookup
  4. Re-verify the file's integrity; if offsets no longer match the file, re-plan the scan

Example fix

// before: offsets from stale split
AvroIterable<Record> it = Avro.read(in).project(schema).createReaderFunc(...).split(staleStart, length).build();
// after: validate against current file metadata
long len = in.getLength();
long start = Math.min(splitStart, len);
Preconditions.checkArgument(start >= 0 && start < len, "Invalid range start %s for file of length %s", start, len);
AvroIterable<Record> it = Avro.read(in).project(schema).createReaderFunc(...).split(start, len - start).build();
Defensive patterns

Strategy: validation

Validate before calling

long len = inputFile.getLength();
Preconditions.checkArgument(start >= 0 && start < len,
    "Range start %s out of bounds for file of length %s", start, len);

Try / catch

try (AvroIterable<D> it = buildRangeIterable(start, end)) {
  it.forEach(consumer);
} catch (RuntimeIOException e) {
  replanAndReread();
}

Prevention

When it happens

Trigger: Constructing an AvroIterable with a start offset beyond or too close to the end of the file, or with a start value computed from stale metadata so that reader.sync() cannot locate a valid sync marker (truncated file, wrong offsets from a rewritten file).

Common situations: Task split offsets computed from an old snapshot after the file was compacted; passing start > file size due to a bug in split computation; reading a range of a file that was appended to by a non-Iceberg writer.

Understand the failure class

Background: "Must be a positive integer", "Invalid value", "Unsupported": the invalid-argument-value error family, when a library rejects the value you pass — this error's family across 35 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/0844e3a6e023a4d5. Report an issue: GitHub.

Appendix: source

Thrown at core/src/main/java/org/apache/iceberg/avro/AvroIterable.java:128

          // Ignore close exception
        }
      }
      throw new RuntimeIOException(e, "Failed to open file: %s", file.location());
    }
  }

  private static class AvroRangeIterator<D> implements FileReader<D> {
    private final FileReader<D> reader;
    private final long end;

    AvroRangeIterator(FileReader<D> reader, long start, long end) {
      this.reader = reader;
      this.end = end;

      try {
        reader.sync(start);
      } catch (IOException e) {
        throw new RuntimeIOException(e, "Failed to find sync past position %d", start);
      }
    }

    @Override
    public Schema getSchema() {
      return reader.getSchema();
    }

    @Override
    public boolean hasNext() {
      try {
        return reader.hasNext() && !reader.pastSync(end);
      } catch (IOException e) {
        throw new RuntimeIOException(e, "Failed to check range end: %d", end);
      }
    }

    @Override

View on GitHub (pinned to 86d9c8fc54)