apache/iceberg · error · IllegalStateException
Unrecognized header bytes: 0x%02X 0x%02X
Error message
Unrecognized header bytes: 0x%02X 0x%02X
What it means
AvroEncoderUtil.decode expects the input stream to begin with the Iceberg Avro magic bytes header. When the first two bytes read do not match MAGIC_BYTES, it fails fast with the actual byte values found. This indicates the buffer is not an Iceberg-encoded Avro value, or was produced/corrupted by a different format or version.
Solutions
- Verify the byte[] was produced by AvroEncoderUtil.encode and is being passed unmodified to decode
- Check the first bytes of the payload against the expected magic bytes before decoding
- Check for truncation or an offset error so decoding starts at the true beginning of the buffer
- Regenerate the payload with the same Iceberg version used to read it
Example fix
// before byte[] payload = rawAvroBytes; // not produced by AvroEncoderUtil.encode T value = AvroEncoderUtil.decode(payload); // after byte[] payload = AvroEncoderUtil.encode(datum, schema, icebergSchema); T value = AvroEncoderUtil.decode(payload);
Defensive patterns
Strategy: validation
Validate before calling
byte[] expected = AvroEncoderUtil.MAGIC_BYTES; // or known header
if (payload.length < 2 || payload[0] != expected[0] || payload[1] != expected[1]) {
throw new IllegalArgumentException("payload is not an Iceberg Avro-encoded blob");
} Type guard
boolean isIcebergAvroPayload(byte[] b) { return b != null && b.length >= 2; } // check against magic constants Try / catch
try { return AvroEncoderUtil.decode(payload); } catch (IllegalStateException e) { if (e.getMessage().startsWith("Unrecognized header bytes")) { /* regenerate or re-route payload */ } throw e; } Prevention
- Only pass byte[] produced by AvroEncoderUtil.encode to decode
- Keep the same Iceberg version on writer and reader paths
- Beware byte[] slicing offsets when storing payloads in queues/blobs
When it happens
Trigger: Calling AvroEncoderUtil.decode with a byte[]/InputStream whose first two bytes are not the expected magic bytes — e.g. passing a plain Avro blob, a Parquet/ORC payload, a truncated or corrupted buffer, or data encoded by a different tool version.
Common situations: Passing raw Avro-serialized data instead of Iceberg's AvroEncoderUtil.encode output; version mismatch where an older/newer writer changed the envelope format; byte-offset mistakes that skip or misalign the header; data corruption in transit or storage.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
Related errors
- Not an Avro file
- Invalid sync at
- Avro does not support AAD prefix
- Avro does not support file encryption keys
- Avro does not support LOCAL TIMESTAMP type with precision
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/459b215bb8538c61.
Report an issue: GitHub.
Appendix: source
Thrown at core/src/main/java/org/apache/iceberg/avro/AvroEncoderUtil.java:75
// Encode the datum with avro schema.
BinaryEncoder encoder = EncoderFactory.get().binaryEncoder(out, null);
DatumWriter<T> writer = new GenericAvroWriter<>(avroSchema);
writer.write(datum, encoder);
encoder.flush();
return out.toByteArray();
}
}
public static <T> T decode(byte[] data) throws IOException {
try (ByteArrayInputStream in = new ByteArrayInputStream(data, 0, data.length)) {
DataInputStream dataInput = new DataInputStream(in);
// Read the magic bytes
byte header0 = dataInput.readByte();
byte header1 = dataInput.readByte();
if (header0 != MAGIC_BYTES[0] || header1 != MAGIC_BYTES[1]) {
throw new IllegalStateException(
String.format(
Locale.ROOT, "Unrecognized header bytes: 0x%02X 0x%02X", header0, header1));
}
// Read avro schema
Schema avroSchema = new Schema.Parser().parse(dataInput.readUTF());
// Decode the datum with the parsed avro schema.
BinaryDecoder binaryDecoder = DecoderFactory.get().binaryDecoder(in, null);
DatumReader<T> reader = new GenericAvroReader<>(avroSchema);
reader.setSchema(avroSchema);
return reader.read(null, binaryDecoder);
}
}
}
View on GitHub (pinned to 86d9c8fc54)