apache/iceberg · error · InvalidAvroMagicException
Not an Avro file
Error message
Not an Avro file
What it means
During Avro.IO.findStartingRowPos, the file header's magic bytes are read to locate the sync marker; if the magic does not match the expected Avro magic ("Obj\u0001"), an InvalidAvroMagicException('Not an Avro file') is thrown. This indicates the file being position-scanned is not an Avro file, or its header is corrupted.
Solutions
- Verify the file's format in the manifest matches Avro (check file_path and file_format)
- Inspect the first bytes of the file (should be 'Obj\u0001') to confirm it is Avro
- Fix or recompute the manifest entries pointing at the wrong files
- Restore or rewrite the corrupted file if the header is damaged
Example fix
// before // manifest lists a Parquet file with format=AVRO long pos = Avro.IO.findStartingRowPos(...) // InvalidAvroMagicException // after // ensure FileFormat.AVRO files only; or read with the correct format model CloseableIterable<T> rows = Parquet.read(file).project(schema).build();
Defensive patterns
Strategy: try-catch
Validate before calling
try (FSDataInputStream in = io.newInputFile(path).newStream()) { byte[] head = new byte[4]; in.readFully(head); if (!Arrays.equals(head, new byte[]{'O','b','j',1})) { throw new IllegalStateException(path + " is not an Avro file"); } } Try / catch
try { return Avro.IO.findStartingRowPos(...); } catch (InvalidAvroMagicException e) { /* validate manifest format entry and file header; fail the scan with a clear error */ throw e; } Prevention
- Keep file_format entries in manifests accurate
- Validate file headers after external writes or restores
- Never manually swap data files across formats
When it happens
Trigger: Calling findStartingRowPos on a file whose bytes are not Avro (e.g. a Parquet/ORC data file misregistered as Avro in the manifest, or a metadata JSON path pointing at the wrong file), or on a truncated/corrupted Avro file whose header magic was overwritten.
Common situations: Manifests referencing the wrong file format after manual edits or catalog corruption; deleted files replaced by zero-byte or garbage content; cross-format file reuse bugs where delete files and data files were written with different formats.
Understand the failure class
Background: Schema validation failed / invalid input schema: payload rejected because its shape doesn't match the expected schema — this error's family across 28 libraries.
Related errors
- Unrecognized header bytes: 0x%02X 0x%02X
- Invalid sync at
- Avro does not support AAD prefix
- Avro does not support file encryption keys
- Avro does not support LOCAL TIMESTAMP type with precision
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/649179f42cd65e9c.
Report an issue: GitHub.
Appendix: source
Thrown at core/src/main/java/org/apache/iceberg/avro/AvroIO.java:163
static long findStartingRowPos(Supplier<SeekableInputStream> open, long start) {
long totalRows = 0;
try (SeekableInputStream in = open.get()) {
// use a direct decoder that will not buffer so the position of the input stream is accurate
BinaryDecoder decoder = DecoderFactory.get().directBinaryDecoder(in, null);
// an Avro file's layout looks like this:
// header|block|block|...
// the header contains:
// magic|string-map|sync
// each block consists of:
// row-count|compressed-size-in-bytes|block-bytes|sync
// it is necessary to read the header here because this is the only way to get the expected
// file sync bytes
byte[] magic = MAGIC_READER.read(decoder, null);
if (!Arrays.equals(AVRO_MAGIC, magic)) {
throw new InvalidAvroMagicException("Not an Avro file");
}
META_READER.read(decoder, null); // ignore the file metadata, it isn't needed
byte[] fileSync = SYNC_READER.read(decoder, null);
// the while loop reads row counts and seeks past the block bytes until the next sync pos is
// >= start, which
// indicates that the next sync is the start of the split.
byte[] blockSync = new byte[16];
long nextSyncPos = in.getPos();
while (nextSyncPos < start) {
if (nextSyncPos != in.getPos()) {
in.seek(nextSyncPos);
SYNC_READER.read(decoder, blockSync);
if (!Arrays.equals(fileSync, blockSync)) {
throw new RuntimeIOException("Invalid sync at %s", nextSyncPos);View on GitHub (pinned to 86d9c8fc54)