{"record":{"id":"5010f3f468c804bd","repo":"apache/iceberg","slug":"malformed-input-partial-character-at-end","errorCode":null,"errorMessage":"malformed input: partial character at end","messagePattern":"malformed input: partial character at end","errorType":"exception","errorClass":"UTFDataFormatException","httpStatus":null,"severity":"error","filePath":"flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java","lineNumber":134,"sourceCode":"      switch (ch >> 4) {\n        case 0:\n        case 1:\n        case 2:\n        case 3:\n        case 4:\n        case 5:\n        case 6:\n        case 7:\n          /* 0xxxxxxx */\n          count++;\n          chararr[chararrCount++] = (char) ch;\n          break;\n        case 12:\n        case 13:\n          /* 110x xxxx 10xx xxxx */\n          count += 2;\n          if (count > utflen) {\n            throw new UTFDataFormatException(\"malformed input: partial character at end\");\n          }\n          char2 = bytearr[count - 1];\n          if ((char2 & 0xC0) != 0x80) {\n            throw new UTFDataFormatException(\"malformed input around byte \" + count);\n          }\n          chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));\n          break;\n        case 14:\n          /* 1110 xxxx 10xx xxxx 10xx xxxx */\n          count += 3;\n          if (count > utflen) {\n            throw new UTFDataFormatException(\"malformed input: partial character at end\");\n          }\n          char2 = bytearr[count - 2];\n          char3 = bytearr[count - 1];\n          if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {\n            throw new UTFDataFormatException(\"malformed input around byte \" + (count - 1));\n          }","sourceCodeStart":116,"sourceCodeEnd":152,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java#L116-L152","documentation":"readLongUTF decodes modified UTF-8 and, when it reads a 2-byte lead byte (110x xxxx), expects exactly one continuation byte within the declared length. If consuming it would pass the encoded length (count > utflen), the stream ended mid-character, so it throws UTFDataFormatException(\"malformed input: partial character at end\").","triggerScenarios":"Reading a byte stream where a 2-byte UTF-8 sequence's continuation byte falls past the declared utflen — i.e., the written length prefix disagrees with the actual bytes or the data is truncated/corrupt.","commonSituations":"Reading data written by a different serializer version; truncated checkpoint/state files; byte-order or framing bugs in custom serialization; network corruption in the payload.","solutions":["Verify data was written with the matching writeLongUTF/SerializerHelper version — mixed serializer versions are the usual cause","Check the stream source for truncation (incomplete read, closed stream early)","Recompute/cross-check the length prefix before decoding; add a checksum to serialized payloads","Regenerate the corrupted state/data if the bytes are simply bad"],"exampleFix":"// before\nDataInputStream in = new DataInputStream(fis); // file previously truncated\nString s = SerializerHelper.readLongUTF(in);\n// after\n// verify file completeness/checksum before decoding\nif (!checksumMatches(file)) { throw new CorruptDataException(\"truncated payload\"); }\nString s = SerializerHelper.readLongUTF(in);","handlingStrategy":"try-catch","validationCode":"// after obtaining bytes but before readLongUTF\nif (declaredLength > availableBytes) {\n  throw new IOException(\"Truncated payload: declared \" + declaredLength + \" bytes, have \" + availableBytes);\n}","typeGuard":"null","tryCatchPattern":"try {\n  return SerializerHelper.readLongUTF(in);\n} catch (UTFDataFormatException e) {\n  if (e.getMessage().contains(\"partial character at end\")) {\n    throw new CorruptPayloadException(\"Serialized string truncated mid-character; regenerate or re-fetch data\", e);\n  }\n  throw e;\n}","preventionTips":["Always pair writeLongUTF/readLongUTF from the same SerializerHelper version","Verify file/stream completeness (checksums, size checks) before decoding","Avoid manual length-prefix writes — let writeLongUTF compute encoded length"],"tags":["flink","serialization","utf-8","corrupt-data"],"backgroundTag":"invalid-argument-format","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}