apache/iceberg · error · UTFDataFormatException

Encoded string reached maximum length:

Error message

Encoded string reached maximum length: 

What it means

SerializerHelper.writeLongUTF encodes a String as modified UTF-8 with an explicit length prefix. While accumulating encoded length, if the running total exceeds Integer.MAX_VALUE it throws UTFDataFormatException, because the length cannot be represented in the following 4-byte int prefix. It is a hard size ceiling on any single string serialized this way.

Solutions

  1. Chunk the string into smaller segments before serializing and reassemble on read
  2. Do not pass unbounded strings through this serializer; enforce an application-level maximum size upstream
  3. If legitimate data approaches 2GB, switch to a length-prefixed long/segmented representation instead of one writeLongUTF call

Example fix

// before
helper.writeLongUTF(hugeString); // > Integer.MAX_VALUE encoded bytes
// after
for (String chunk : splitIntoChunks(hugeString, 1 << 20)) {
  helper.writeLongUTF(chunk);
}
Defensive patterns

Strategy: validation

Validate before calling

static final int MAX_ENCODED = Integer.MAX_VALUE;
if (estimateModifiedUtf8Length(s) > MAX_ENCODED) {
  throw new IllegalArgumentException("String too large to serialize via writeLongUTF");
}
serializerHelper.writeLongUTF(s);

Type guard

boolean canSerialize(String s) { return s != null && estimateModifiedUtf8Length(s) <= Integer.MAX_VALUE; }

Try / catch

try {
  helper.writeLongUTF(str);
} catch (UTFDataFormatException e) {
  throw new IOException("String exceeds serializer limit; chunk or compress the payload", e);
}

Prevention

When it happens

Trigger: Serializing a single String field via writeLongUTF whose modified-UTF-8 encoding exceeds Integer.MAX_VALUE (~2.1 billion) bytes — effectively only reachable with enormous strings or a length-accumulation overflow.

Common situations: Serializing huge VARCHAR payloads or corrupted/attacker-controlled length bookkeeping; streaming unbounded string data through a Flink serializer without chunking.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/22b1d9804eaaa704. Report an issue: GitHub.

Appendix: source

Thrown at flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:65

   *
   * <p>See * <a href="https://issues.apache.org/jira/browse/FLINK-34228">FLINK-34228</a> * <a
   * href="https://github.com/apache/flink/pull/24191">https://github.com/apache/flink/pull/24191</a>
   *
   * @param out the output stream to write the string to.
   * @param str the string value to be written.
   */
  public static void writeLongUTF(DataOutputView out, String str) throws IOException {
    int strlen = str.length();
    long utflen = 0;
    int ch;

    /* use charAt instead of copying String to char array */
    for (int i = 0; i < strlen; i++) {
      ch = str.charAt(i);
      utflen += getUTFBytesSize(ch);

      if (utflen > Integer.MAX_VALUE) {
        throw new UTFDataFormatException("Encoded string reached maximum length: " + utflen);
      }
    }

    if (utflen > Integer.MAX_VALUE - 4) {
      throw new UTFDataFormatException("Encoded string is too long: " + utflen);
    }

    out.writeInt((int) utflen);
    writeUTFBytes(out, str, (int) utflen);
  }

  /**
   * Similar to {@link DataInputDeserializer#readUTF()}. Except this supports larger payloads which
   * is up to max integer value.
   *
   * <p>Note: This method can be removed when the method which does similar thing within the {@link
   * DataOutputSerializer} already which does the same thing, so use that one instead once that is
   * released on Flink version 1.20.

View on GitHub (pinned to 86d9c8fc54)