apache/iceberg · error · UTFDataFormatException

malformed input around byte

Error message

malformed input around byte 

What it means

SerializerHelper.readLongUTF decodes a Java-modified-UTF string from serialized bytes. When a multi-byte character starts with a 2-byte lead (110xxxxx) but the following byte does not have the continuation prefix 10xxxxxx, the stream is not valid modified UTF-8 and a UTFDataFormatException with the byte offset is thrown. This protects callers from silently decoding corrupted or misaligned payload data.

Solutions

  1. Verify the bytes were written by SerializerHelper.writeLongUTF (or DataOutput.writeUTF) with matching read/write order and no offset drift.
  2. Dump the byte range around the reported offset and check that continuation bytes (0x80-0xBF) follow each lead byte.
  3. Ensure sender and reader use the same Iceberg/Flink serializer versions for the affected state descriptor.
  4. If data is externally produced, transcode it to Java modified UTF-8 (or use a standard String serializer) before decoding.

Example fix

// before
String value = SerializerHelper.readLongUTF(in); // assumes arbitrary UTF-8 bytes
// after
// write side: ensure symmetric API
SerializerHelper.writeLongUTF(out, value);
// read side: only consume bytes produced by writeLongUTF; validate provenance/version first
String value = SerializerHelper.readLongUTF(in);
Defensive patterns

Strategy: validation

Validate before calling

// verify provenance before decoding: bytes must come from writeLongUTF/writeUTF
if (payload == null || payload.length == 0) throw new IllegalArgumentException("empty payload");
// spot-check: byte at any lead position must be a valid lead (not 0x80-0xBF, not 0xF0-0xFF)
for (int i = 0; i < payload.length; ) {
  int b = payload[i] & 0xFF;
  if ((b & 0xC0) == 0x80 || (b & 0xF8) == 0xF0) throw new IllegalArgumentException("invalid UTF lead at " + i);
  i += (b & 0x80) == 0 ? 1 : (b & 0xE0) == 0xC0 ? 2 : (b & 0xF0) == 0xE0 ? 3 : 1;
}

Try / catch

try {
  String s = SerializerHelper.readLongUTF(in);
} catch (UTFDataFormatException e) {
  throw new IOException("corrupt string state at offset " + e.getMessage(), e);
}

Prevention

When it happens

Trigger: Calling readLongUTF on bytes where a 2-byte UTF lead byte is followed by a non-continuation byte; typically the byte array was produced by a different serializer, truncated, or the length prefix was read at the wrong offset.

Common situations: Flink job state deserialized across incompatible serializer versions; hand-rolled binary formats mixing standard UTF-8 with Java modified UTF; corrupted Kafka/checkpoint payloads; reading past a field boundary so the decoder starts mid-character.

Understand the failure class

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/cab0c02ff7b49f47. Report an issue: GitHub.

Appendix: source

Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:138

        case 3:
        case 4:
        case 5:
        case 6:
        case 7:
          /* 0xxxxxxx */
          count++;
          chararr[chararrCount++] = (char) ch;
          break;
        case 12:
        case 13:
          /* 110x xxxx 10xx xxxx */
          count += 2;
          if (count > utflen) {
            throw new UTFDataFormatException("malformed input: partial character at end");
          }
          char2 = bytearr[count - 1];
          if ((char2 & 0xC0) != 0x80) {
            throw new UTFDataFormatException("malformed input around byte " + count);
          }
          chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));
          break;
        case 14:
          /* 1110 xxxx 10xx xxxx 10xx xxxx */
          count += 3;
          if (count > utflen) {
            throw new UTFDataFormatException("malformed input: partial character at end");
          }
          char2 = bytearr[count - 2];
          char3 = bytearr[count - 1];
          if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {
            throw new UTFDataFormatException("malformed input around byte " + (count - 1));
          }
          chararr[chararrCount++] =
              (char) (((ch & 0x0F) << 12) | ((char2 & 0x3F) << 6) | (char3 & 0x3F));
          break;
        default:

View on GitHub (pinned to 86d9c8fc54)