apache/iceberg · error · UTFDataFormatException
Encoded string reached maximum length:
Error message
Encoded string reached maximum length:
What it means
SerializerHelper.writeLongUTF encodes a String as modified UTF-8 with an explicit length prefix. While accumulating encoded length, if the running total exceeds Integer.MAX_VALUE it throws UTFDataFormatException, because the length cannot be represented in the following 4-byte int prefix. It is a hard size ceiling on any single string serialized this way.
Solutions
- Chunk the string into smaller segments before serializing and reassemble on read
- Do not pass unbounded strings through this serializer; enforce an application-level maximum size upstream
- If legitimate data approaches 2GB, switch to a length-prefixed long/segmented representation instead of one writeLongUTF call
Example fix
// before
helper.writeLongUTF(hugeString); // > Integer.MAX_VALUE encoded bytes
// after
for (String chunk : splitIntoChunks(hugeString, 1 << 20)) {
helper.writeLongUTF(chunk);
} Defensive patterns
Strategy: validation
Validate before calling
static final int MAX_ENCODED = Integer.MAX_VALUE;
if (estimateModifiedUtf8Length(s) > MAX_ENCODED) {
throw new IllegalArgumentException("String too large to serialize via writeLongUTF");
}
serializerHelper.writeLongUTF(s); Type guard
boolean canSerialize(String s) { return s != null && estimateModifiedUtf8Length(s) <= Integer.MAX_VALUE; } Try / catch
try {
helper.writeLongUTF(str);
} catch (UTFDataFormatException e) {
throw new IOException("String exceeds serializer limit; chunk or compress the payload", e);
} Prevention
- Cap application string sizes well below 2GB
- Chunk large text fields into segments before serializing
- Never stream unbounded user input through a single writeLongUTF call
When it happens
Trigger: Serializing a single String field via writeLongUTF whose modified-UTF-8 encoding exceeds Integer.MAX_VALUE (~2.1 billion) bytes — effectively only reachable with enormous strings or a length-accumulation overflow.
Common situations: Serializing huge VARCHAR payloads or corrupted/attacker-controlled length bookkeeping; streaming unbounded string data through a Flink serializer without chunking.
Understand the failure class
Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.
Related errors
- Encoded string is too long:
- malformed input around byte
- malformed input: partial character at end
- Could not deserialize the WriteResult object
- Could not deserialize the WriteResult object
AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12).
Data as JSON: /api/errors/22b1d9804eaaa704.
Report an issue: GitHub.
Appendix: source
Thrown at flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:65
*
* <p>See * <a href="https://issues.apache.org/jira/browse/FLINK-34228">FLINK-34228</a> * <a
* href="https://github.com/apache/flink/pull/24191">https://github.com/apache/flink/pull/24191</a>
*
* @param out the output stream to write the string to.
* @param str the string value to be written.
*/
public static void writeLongUTF(DataOutputView out, String str) throws IOException {
int strlen = str.length();
long utflen = 0;
int ch;
/* use charAt instead of copying String to char array */
for (int i = 0; i < strlen; i++) {
ch = str.charAt(i);
utflen += getUTFBytesSize(ch);
if (utflen > Integer.MAX_VALUE) {
throw new UTFDataFormatException("Encoded string reached maximum length: " + utflen);
}
}
if (utflen > Integer.MAX_VALUE - 4) {
throw new UTFDataFormatException("Encoded string is too long: " + utflen);
}
out.writeInt((int) utflen);
writeUTFBytes(out, str, (int) utflen);
}
/**
* Similar to {@link DataInputDeserializer#readUTF()}. Except this supports larger payloads which
* is up to max integer value.
*
* <p>Note: This method can be removed when the method which does similar thing within the {@link
* DataOutputSerializer} already which does the same thing, so use that one instead once that is
* released on Flink version 1.20.View on GitHub (pinned to 86d9c8fc54)