apache/iceberg · error · IllegalStateException

Unrecognized header bytes: 0x%02X 0x%02X

Error message

Unrecognized header bytes: 0x%02X 0x%02X

What it means

AvroEncoderUtil.decode expects the input stream to begin with the Iceberg Avro magic bytes header. When the first two bytes read do not match MAGIC_BYTES, it fails fast with the actual byte values found. This indicates the buffer is not an Iceberg-encoded Avro value, or was produced/corrupted by a different format or version.

Solutions

  1. Verify the byte[] was produced by AvroEncoderUtil.encode and is being passed unmodified to decode
  2. Check the first bytes of the payload against the expected magic bytes before decoding
  3. Check for truncation or an offset error so decoding starts at the true beginning of the buffer
  4. Regenerate the payload with the same Iceberg version used to read it

Example fix

// before
byte[] payload = rawAvroBytes; // not produced by AvroEncoderUtil.encode
T value = AvroEncoderUtil.decode(payload);

// after
byte[] payload = AvroEncoderUtil.encode(datum, schema, icebergSchema);
T value = AvroEncoderUtil.decode(payload);
Defensive patterns

Strategy: validation

Validate before calling

byte[] expected = AvroEncoderUtil.MAGIC_BYTES; // or known header
if (payload.length < 2 || payload[0] != expected[0] || payload[1] != expected[1]) {
  throw new IllegalArgumentException("payload is not an Iceberg Avro-encoded blob");
}

Type guard

boolean isIcebergAvroPayload(byte[] b) { return b != null && b.length >= 2; } // check against magic constants

Try / catch

try { return AvroEncoderUtil.decode(payload); } catch (IllegalStateException e) { if (e.getMessage().startsWith("Unrecognized header bytes")) { /* regenerate or re-route payload */ } throw e; }

Prevention

When it happens

Trigger: Calling AvroEncoderUtil.decode with a byte[]/InputStream whose first two bytes are not the expected magic bytes — e.g. passing a plain Avro blob, a Parquet/ORC payload, a truncated or corrupted buffer, or data encoded by a different tool version.

Common situations: Passing raw Avro-serialized data instead of Iceberg's AvroEncoderUtil.encode output; version mismatch where an older/newer writer changed the envelope format; byte-offset mistakes that skip or misalign the header; data corruption in transit or storage.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/459b215bb8538c61. Report an issue: GitHub.

Appendix: source

Thrown at core/src/main/java/org/apache/iceberg/avro/AvroEncoderUtil.java:75

      // Encode the datum with avro schema.
      BinaryEncoder encoder = EncoderFactory.get().binaryEncoder(out, null);
      DatumWriter<T> writer = new GenericAvroWriter<>(avroSchema);
      writer.write(datum, encoder);
      encoder.flush();

      return out.toByteArray();
    }
  }

  public static <T> T decode(byte[] data) throws IOException {
    try (ByteArrayInputStream in = new ByteArrayInputStream(data, 0, data.length)) {
      DataInputStream dataInput = new DataInputStream(in);

      // Read the magic bytes
      byte header0 = dataInput.readByte();
      byte header1 = dataInput.readByte();
      if (header0 != MAGIC_BYTES[0] || header1 != MAGIC_BYTES[1]) {
        throw new IllegalStateException(
            String.format(
                Locale.ROOT, "Unrecognized header bytes: 0x%02X 0x%02X", header0, header1));
      }

      // Read avro schema
      Schema avroSchema = new Schema.Parser().parse(dataInput.readUTF());

      // Decode the datum with the parsed avro schema.
      BinaryDecoder binaryDecoder = DecoderFactory.get().binaryDecoder(in, null);
      DatumReader<T> reader = new GenericAvroReader<>(avroSchema);
      reader.setSchema(avroSchema);
      return reader.read(null, binaryDecoder);
    }
  }
}

View on GitHub (pinned to 86d9c8fc54)