apache/iceberg · error · UTFDataFormatException

Encoded string is too long:

Error message

Encoded string is too long: 

What it means

writeLongUTF reserves 4 bytes of headroom: after encoding, if the total length exceeds Integer.MAX_VALUE - 4 it throws UTFDataFormatException. This check guarantees the encoded length plus any small framing the caller adds still fits in an int when written via out.writeInt((int) utflen).

Solutions

  1. Reduce the payload size (chunk, truncate, or compress) before serializing
  2. Validate string length before calling writeLongUTF
  3. Split into multiple records/parts so each encoded segment stays under the limit

Example fix

// before
helper.writeLongUTF(value);
// after
if (estimateUTFLength(value) <= Integer.MAX_VALUE - 4) {
  helper.writeLongUTF(value);
} else {
  writeChunked(helper, value);
}
Defensive patterns

Strategy: validation

Validate before calling

static final int MAX_ENCODED = Integer.MAX_VALUE - 4;
if (estimateModifiedUtf8Length(s) > MAX_ENCODED) {
  throw new IllegalArgumentException("Encoded string exceeds " + MAX_ENCODED + " bytes");
}

Type guard

boolean withinLimit(String s) {
  return s != null && estimateModifiedUtf8Length(s) <= Integer.MAX_VALUE - 4;
}

Try / catch

try {
  helper.writeLongUTF(str);
} catch (UTFDataFormatException e) {
  log.error("Payload too large for long-UTF serializer");
  writeChunked(helper, str);
}

Prevention

When it happens

Trigger: Calling SerializerHelper.writeLongUTF with a String whose modified-UTF-8 encoded size is greater than Integer.MAX_VALUE - 4 bytes (~2147483643).

Common situations: Extreme single-record sizes in Flink serialization; passing unbounded user input through a serialized state or checkpoint path.

Understand the failure class

Background: "value must be between 0 and 1" / "out of range" / "must not be negative" errors: fixing range-validation failures across open-source libraries — this error's family across 42 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/bed4b29a69bb00a8. Report an issue: GitHub.

Appendix: source

Thrown at flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:70

   * @param str the string value to be written.
   */
  public static void writeLongUTF(DataOutputView out, String str) throws IOException {
    int strlen = str.length();
    long utflen = 0;
    int ch;

    /* use charAt instead of copying String to char array */
    for (int i = 0; i < strlen; i++) {
      ch = str.charAt(i);
      utflen += getUTFBytesSize(ch);

      if (utflen > Integer.MAX_VALUE) {
        throw new UTFDataFormatException("Encoded string reached maximum length: " + utflen);
      }
    }

    if (utflen > Integer.MAX_VALUE - 4) {
      throw new UTFDataFormatException("Encoded string is too long: " + utflen);
    }

    out.writeInt((int) utflen);
    writeUTFBytes(out, str, (int) utflen);
  }

  /**
   * Similar to {@link DataInputDeserializer#readUTF()}. Except this supports larger payloads which
   * is up to max integer value.
   *
   * <p>Note: This method can be removed when the method which does similar thing within the {@link
   * DataOutputSerializer} already which does the same thing, so use that one instead once that is
   * released on Flink version 1.20.
   *
   * <p>See * <a href="https://issues.apache.org/jira/browse/FLINK-34228">FLINK-34228</a> * <a
   * href="https://github.com/apache/flink/pull/24191">https://github.com/apache/flink/pull/24191</a>
   *
   * @param in the input stream to read the string from.

View on GitHub (pinned to 86d9c8fc54)