apache/iceberg · error · UTFDataFormatException
malformed input around byte
Error message
malformed input around byte
What it means
SerializerHelper.readLongUTF decodes a Java-modified-UTF string from serialized bytes. When a multi-byte character starts with a 2-byte lead (110xxxxx) but the following byte does not have the continuation prefix 10xxxxxx, the stream is not valid modified UTF-8 and a UTFDataFormatException with the byte offset is thrown. This protects callers from silently decoding corrupted or misaligned payload data.
Solutions
- Verify the bytes were written by SerializerHelper.writeLongUTF (or DataOutput.writeUTF) with matching read/write order and no offset drift.
- Dump the byte range around the reported offset and check that continuation bytes (0x80-0xBF) follow each lead byte.
- Ensure sender and reader use the same Iceberg/Flink serializer versions for the affected state descriptor.
- If data is externally produced, transcode it to Java modified UTF-8 (or use a standard String serializer) before decoding.
Example fix
// before String value = SerializerHelper.readLongUTF(in); // assumes arbitrary UTF-8 bytes // after // write side: ensure symmetric API SerializerHelper.writeLongUTF(out, value); // read side: only consume bytes produced by writeLongUTF; validate provenance/version first String value = SerializerHelper.readLongUTF(in);
Defensive patterns
Strategy: validation
Validate before calling
// verify provenance before decoding: bytes must come from writeLongUTF/writeUTF
if (payload == null || payload.length == 0) throw new IllegalArgumentException("empty payload");
// spot-check: byte at any lead position must be a valid lead (not 0x80-0xBF, not 0xF0-0xFF)
for (int i = 0; i < payload.length; ) {
int b = payload[i] & 0xFF;
if ((b & 0xC0) == 0x80 || (b & 0xF8) == 0xF0) throw new IllegalArgumentException("invalid UTF lead at " + i);
i += (b & 0x80) == 0 ? 1 : (b & 0xE0) == 0xC0 ? 2 : (b & 0xF0) == 0xE0 ? 3 : 1;
} Try / catch
try {
String s = SerializerHelper.readLongUTF(in);
} catch (UTFDataFormatException e) {
throw new IOException("corrupt string state at offset " + e.getMessage(), e);
} Prevention
- Always pair writeLongUTF with readLongUTF from the same serializer version
- Never decode externally produced UTF-8 bytes with the modified-UTF reader
- Log the serializer version alongside serialized state for drift diagnosis
When it happens
Trigger: Calling readLongUTF on bytes where a 2-byte UTF lead byte is followed by a non-continuation byte; typically the byte array was produced by a different serializer, truncated, or the length prefix was read at the wrong offset.
Common situations: Flink job state deserialized across incompatible serializer versions; hand-rolled binary formats mixing standard UTF-8 with Java modified UTF; corrupted Kafka/checkpoint payloads; reading past a field boundary so the decoder starts mid-character.
Understand the failure class
- Parsing and encoding errors: unexpected token, malformed input — why parsers reject input and how to find the real culprit.
Related errors
- Could not deserialize the WriteResult object
- Could not deserialize the WriteResult object
- Could not deserialize the WriteResult object
- Could not deserialize the WriteResult object
- Encoded string is too long:
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/cab0c02ff7b49f47.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:138
case 3:
case 4:
case 5:
case 6:
case 7:
/* 0xxxxxxx */
count++;
chararr[chararrCount++] = (char) ch;
break;
case 12:
case 13:
/* 110x xxxx 10xx xxxx */
count += 2;
if (count > utflen) {
throw new UTFDataFormatException("malformed input: partial character at end");
}
char2 = bytearr[count - 1];
if ((char2 & 0xC0) != 0x80) {
throw new UTFDataFormatException("malformed input around byte " + count);
}
chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));
break;
case 14:
/* 1110 xxxx 10xx xxxx 10xx xxxx */
count += 3;
if (count > utflen) {
throw new UTFDataFormatException("malformed input: partial character at end");
}
char2 = bytearr[count - 2];
char3 = bytearr[count - 1];
if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {
throw new UTFDataFormatException("malformed input around byte " + (count - 1));
}
chararr[chararrCount++] =
(char) (((ch & 0x0F) << 12) | ((char2 & 0x3F) << 6) | (char3 & 0x3F));
break;
default:View on GitHub (pinned to 86d9c8fc54)