apache/iceberg · error · UTFDataFormatException

malformed input around byte

Error message

malformed input around byte 

What it means

While decoding a 2-byte modified-UTF-8 sequence, readLongUTF checks that the continuation byte has the 10xx xxxx bit pattern. If bytearr[count-1] fails the (char2 & 0xC0) != 0x80 test, the stream contains an invalid UTF-8 continuation byte and it throws UTFDataFormatException reporting the byte offset.

Solutions

  1. Ensure both writer and reader use the same modified-UTF-8 encoding (writeLongUTF on the write side)
  2. Do not feed raw UTF-8 or platform-encoded bytes into readLongUTF; re-encode via String.getBytes with a compatible scheme or rewrite with writeLongUTF
  3. Validate/repair the payload bytes if coming from an external source
  4. Add checksums at write time to catch corruption early

Example fix

// before
out.write(str.getBytes(StandardCharsets.UTF_8)); // plain UTF-8, wrong framing
String s = SerializerHelper.readLongUTF(in); // fails on continuation check
// after
SerializerHelper.writeLongUTF(out, str); // matched write/read pair
String s = SerializerHelper.readLongUTF(in);
Defensive patterns

Strategy: try-catch

Validate before calling

// validate payload is well-formed modified UTF-8 before decoding
boolean validModifiedUtf8(byte[] b) { /* scan lead/continuation byte patterns */ }

Type guard

null

Try / catch

try {
  return SerializerHelper.readLongUTF(in);
} catch (UTFDataFormatException e) {
  if (e.getMessage().startsWith("malformed input around byte")) {
    throw new CorruptPayloadException("Non-modified-UTF-8 bytes in stream; check writer encoding", e);
  }
  throw e;
}

Prevention

When it happens

Trigger: Reading a stream whose bytes are not valid modified UTF-8: e.g. raw 8-bit characters, Latin-1 encoded bytes, or a UTF-8 variant (CESU-8 / different surrogate handling) written by another tool.

Common situations: Strings written by non-Java encoders, data crossing system boundaries with different default charsets, or manually crafted/corrupted serialized payloads.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of apache/iceberg@86d9c8fc54 (2026-09-12). Data as JSON: /api/errors/27640b97bc5cb5ef. Report an issue: GitHub.

Appendix: source

Thrown at flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java:138

        case 3:
        case 4:
        case 5:
        case 6:
        case 7:
          /* 0xxxxxxx */
          count++;
          chararr[chararrCount++] = (char) ch;
          break;
        case 12:
        case 13:
          /* 110x xxxx 10xx xxxx */
          count += 2;
          if (count > utflen) {
            throw new UTFDataFormatException("malformed input: partial character at end");
          }
          char2 = bytearr[count - 1];
          if ((char2 & 0xC0) != 0x80) {
            throw new UTFDataFormatException("malformed input around byte " + count);
          }
          chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));
          break;
        case 14:
          /* 1110 xxxx 10xx xxxx 10xx xxxx */
          count += 3;
          if (count > utflen) {
            throw new UTFDataFormatException("malformed input: partial character at end");
          }
          char2 = bytearr[count - 2];
          char3 = bytearr[count - 1];
          if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {
            throw new UTFDataFormatException("malformed input around byte " + (count - 1));
          }
          chararr[chararrCount++] =
              (char) (((ch & 0x0F) << 12) | ((char2 & 0x3F) << 6) | (char3 & 0x3F));
          break;
        default:

View on GitHub (pinned to 86d9c8fc54)