apache/iceberg · error · UTFDataFormatException
malformed input around byte
Error message
malformed input around byte
What it means
While decoding a 2-byte modified-UTF-8 sequence, readLongUTF checks that the continuation byte has the 10xx xxxx bit pattern. If bytearr[count-1] fails the (char2 & 0xC0) != 0x80 test, the stream contains an invalid UTF-8 continuation byte and it throws UTFDataFormatException reporting the byte offset.
Solutions
- Ensure both writer and reader use the same modified-UTF-8 encoding (writeLongUTF on the write side)
- Do not feed raw UTF-8 or platform-encoded bytes into readLongUTF; re-encode via String.getBytes with a compatible scheme or rewrite with writeLongUTF
- Validate/repair the payload bytes if coming from an external source
- Add checksums at write time to catch corruption early
Example fix
// before out.write(str.getBytes(StandardCharsets.UTF_8)); // plain UTF-8, wrong framing String s = SerializerHelper.readLongUTF(in); // fails on continuation check // after SerializerHelper.writeLongUTF(out, str); // matched write/read pair String s = SerializerHelper.readLongUTF(in);
Defensive patterns
Strategy: try-catch
Validate before calling
// validate payload is well-formed modified UTF-8 before decoding
boolean validModifiedUtf8(byte[] b) { /* scan lead/continuation byte patterns */ } Type guard
null
Try / catch
try {
return SerializerHelper.readLongUTF(in);
} catch (UTFDataFormatException e) {
if (e.getMessage().startsWith("malformed input around byte")) {
throw new CorruptPayloadException("Non-modified-UTF-8 bytes in stream; check writer encoding", e);
}
throw e;
} Prevention
- Encode strings only via writeLongUTF (modified UTF-8), never raw platform charsets
- Keep reader aligned: read the length prefix first, then exactly that many bytes
- Checksum payloads written across process/network boundaries
When it happens
Trigger: Reading a stream whose bytes are not valid modified UTF-8: e.g. raw 8-bit characters, Latin-1 encoded bytes, or a UTF-8 variant (CESU-8 / different surrogate handling) written by another tool.
Common situations: Strings written by non-Java encoders, data crossing system boundaries with different default charsets, or manually crafted/corrupted serialized payloads.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
- Parsing and encoding errors: unexpected token, malformed input — why parsers reject input and how to find the real culprit.
Related errors
- malformed input: partial character at end
- Encoded string is too long:
- Encoded string reached maximum length:
- Could not deserialize the WriteResult object
- Could not deserialize the WriteResult object
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/27640b97bc5cb5ef.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:138
case 3:
case 4:
case 5:
case 6:
case 7:
/* 0xxxxxxx */
count++;
chararr[chararrCount++] = (char) ch;
break;
case 12:
case 13:
/* 110x xxxx 10xx xxxx */
count += 2;
if (count > utflen) {
throw new UTFDataFormatException("malformed input: partial character at end");
}
char2 = bytearr[count - 1];
if ((char2 & 0xC0) != 0x80) {
throw new UTFDataFormatException("malformed input around byte " + count);
}
chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));
break;
case 14:
/* 1110 xxxx 10xx xxxx 10xx xxxx */
count += 3;
if (count > utflen) {
throw new UTFDataFormatException("malformed input: partial character at end");
}
char2 = bytearr[count - 2];
char3 = bytearr[count - 1];
if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {
throw new UTFDataFormatException("malformed input around byte " + (count - 1));
}
chararr[chararrCount++] =
(char) (((ch & 0x0F) << 12) | ((char2 & 0x3F) << 6) | (char3 & 0x3F));
break;
default:View on GitHub (pinned to 86d9c8fc54)