apache/iceberg · error · UTFDataFormatException
malformed input: partial character at end
Error message
malformed input: partial character at end
What it means
readLongUTF decodes modified UTF-8 and, when it reads a 2-byte lead byte (110x xxxx), expects exactly one continuation byte within the declared length. If consuming it would pass the encoded length (count > utflen), the stream ended mid-character, so it throws UTFDataFormatException("malformed input: partial character at end").
Solutions
- Verify data was written with the matching writeLongUTF/SerializerHelper version — mixed serializer versions are the usual cause
- Check the stream source for truncation (incomplete read, closed stream early)
- Recompute/cross-check the length prefix before decoding; add a checksum to serialized payloads
- Regenerate the corrupted state/data if the bytes are simply bad
Example fix
// before
DataInputStream in = new DataInputStream(fis); // file previously truncated
String s = SerializerHelper.readLongUTF(in);
// after
// verify file completeness/checksum before decoding
if (!checksumMatches(file)) { throw new CorruptDataException("truncated payload"); }
String s = SerializerHelper.readLongUTF(in); Defensive patterns
Strategy: try-catch
Validate before calling
// after obtaining bytes but before readLongUTF
if (declaredLength > availableBytes) {
throw new IOException("Truncated payload: declared " + declaredLength + " bytes, have " + availableBytes);
} Type guard
null
Try / catch
try {
return SerializerHelper.readLongUTF(in);
} catch (UTFDataFormatException e) {
if (e.getMessage().contains("partial character at end")) {
throw new CorruptPayloadException("Serialized string truncated mid-character; regenerate or re-fetch data", e);
}
throw e;
} Prevention
- Always pair writeLongUTF/readLongUTF from the same SerializerHelper version
- Verify file/stream completeness (checksums, size checks) before decoding
- Avoid manual length-prefix writes — let writeLongUTF compute encoded length
When it happens
Trigger: Reading a byte stream where a 2-byte UTF-8 sequence's continuation byte falls past the declared utflen — i.e., the written length prefix disagrees with the actual bytes or the data is truncated/corrupt.
Common situations: Reading data written by a different serializer version; truncated checkpoint/state files; byte-order or framing bugs in custom serialization; network corruption in the payload.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
- Parsing and encoding errors: unexpected token, malformed input — why parsers reject input and how to find the real culprit.
Related errors
- malformed input around byte
- Encoded string is too long:
- Encoded string reached maximum length:
- Could not deserialize the WriteResult object
- Could not deserialize the WriteResult object
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/5010f3f468c804bd.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:134
switch (ch >> 4) {
case 0:
case 1:
case 2:
case 3:
case 4:
case 5:
case 6:
case 7:
/* 0xxxxxxx */
count++;
chararr[chararrCount++] = (char) ch;
break;
case 12:
case 13:
/* 110x xxxx 10xx xxxx */
count += 2;
if (count > utflen) {
throw new UTFDataFormatException("malformed input: partial character at end");
}
char2 = bytearr[count - 1];
if ((char2 & 0xC0) != 0x80) {
throw new UTFDataFormatException("malformed input around byte " + count);
}
chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));
break;
case 14:
/* 1110 xxxx 10xx xxxx 10xx xxxx */
count += 3;
if (count > utflen) {
throw new UTFDataFormatException("malformed input: partial character at end");
}
char2 = bytearr[count - 2];
char3 = bytearr[count - 1];
if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {
throw new UTFDataFormatException("malformed input around byte " + (count - 1));
}View on GitHub (pinned to 86d9c8fc54)