apache/iceberg · error · UTFDataFormatException

malformed input: partial character at end

Error message

malformed input: partial character at end

What it means

readLongUTF decodes modified UTF-8 and, when it reads a 2-byte lead byte (110x xxxx), expects exactly one continuation byte within the declared length. If consuming it would pass the encoded length (count > utflen), the stream ended mid-character, so it throws UTFDataFormatException("malformed input: partial character at end").

Solutions

  1. Verify data was written with the matching writeLongUTF/SerializerHelper version — mixed serializer versions are the usual cause
  2. Check the stream source for truncation (incomplete read, closed stream early)
  3. Recompute/cross-check the length prefix before decoding; add a checksum to serialized payloads
  4. Regenerate the corrupted state/data if the bytes are simply bad

Example fix

// before
DataInputStream in = new DataInputStream(fis); // file previously truncated
String s = SerializerHelper.readLongUTF(in);
// after
// verify file completeness/checksum before decoding
if (!checksumMatches(file)) { throw new CorruptDataException("truncated payload"); }
String s = SerializerHelper.readLongUTF(in);
Defensive patterns

Strategy: try-catch

Validate before calling

// after obtaining bytes but before readLongUTF
if (declaredLength > availableBytes) {
  throw new IOException("Truncated payload: declared " + declaredLength + " bytes, have " + availableBytes);
}

Type guard

null

Try / catch

try {
  return SerializerHelper.readLongUTF(in);
} catch (UTFDataFormatException e) {
  if (e.getMessage().contains("partial character at end")) {
    throw new CorruptPayloadException("Serialized string truncated mid-character; regenerate or re-fetch data", e);
  }
  throw e;
}

Prevention

When it happens

Trigger: Reading a byte stream where a 2-byte UTF-8 sequence's continuation byte falls past the declared utflen — i.e., the written length prefix disagrees with the actual bytes or the data is truncated/corrupt.

Common situations: Reading data written by a different serializer version; truncated checkpoint/state files; byte-order or framing bugs in custom serialization; network corruption in the payload.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/5010f3f468c804bd. Report an issue: GitHub.

Appendix: source

Thrown at flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:134

      switch (ch >> 4) {
        case 0:
        case 1:
        case 2:
        case 3:
        case 4:
        case 5:
        case 6:
        case 7:
          /* 0xxxxxxx */
          count++;
          chararr[chararrCount++] = (char) ch;
          break;
        case 12:
        case 13:
          /* 110x xxxx 10xx xxxx */
          count += 2;
          if (count > utflen) {
            throw new UTFDataFormatException("malformed input: partial character at end");
          }
          char2 = bytearr[count - 1];
          if ((char2 & 0xC0) != 0x80) {
            throw new UTFDataFormatException("malformed input around byte " + count);
          }
          chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));
          break;
        case 14:
          /* 1110 xxxx 10xx xxxx 10xx xxxx */
          count += 3;
          if (count > utflen) {
            throw new UTFDataFormatException("malformed input: partial character at end");
          }
          char2 = bytearr[count - 2];
          char3 = bytearr[count - 1];
          if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {
            throw new UTFDataFormatException("malformed input around byte " + (count - 1));
          }

View on GitHub (pinned to 86d9c8fc54)