{"record":{"id":"27640b97bc5cb5ef","repo":"apache/iceberg","slug":"malformed-input-around-byte","errorCode":null,"errorMessage":"malformed input around byte ","messagePattern":"malformed input around byte ","errorType":"exception","errorClass":"UTFDataFormatException","httpStatus":null,"severity":"error","filePath":"flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java","lineNumber":138,"sourceCode":"        case 3:\n        case 4:\n        case 5:\n        case 6:\n        case 7:\n          /* 0xxxxxxx */\n          count++;\n          chararr[chararrCount++] = (char) ch;\n          break;\n        case 12:\n        case 13:\n          /* 110x xxxx 10xx xxxx */\n          count += 2;\n          if (count > utflen) {\n            throw new UTFDataFormatException(\"malformed input: partial character at end\");\n          }\n          char2 = bytearr[count - 1];\n          if ((char2 & 0xC0) != 0x80) {\n            throw new UTFDataFormatException(\"malformed input around byte \" + count);\n          }\n          chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));\n          break;\n        case 14:\n          /* 1110 xxxx 10xx xxxx 10xx xxxx */\n          count += 3;\n          if (count > utflen) {\n            throw new UTFDataFormatException(\"malformed input: partial character at end\");\n          }\n          char2 = bytearr[count - 2];\n          char3 = bytearr[count - 1];\n          if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {\n            throw new UTFDataFormatException(\"malformed input around byte \" + (count - 1));\n          }\n          chararr[chararrCount++] =\n              (char) (((ch & 0x0F) << 12) | ((char2 & 0x3F) << 6) | (char3 & 0x3F));\n          break;\n        default:","sourceCodeStart":120,"sourceCodeEnd":156,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/flink/v1.20/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java#L120-L156","documentation":"While decoding a 2-byte modified-UTF-8 sequence, readLongUTF checks that the continuation byte has the 10xx xxxx bit pattern. If bytearr[count-1] fails the (char2 & 0xC0) != 0x80 test, the stream contains an invalid UTF-8 continuation byte and it throws UTFDataFormatException reporting the byte offset.","triggerScenarios":"Reading a stream whose bytes are not valid modified UTF-8: e.g. raw 8-bit characters, Latin-1 encoded bytes, or a UTF-8 variant (CESU-8 / different surrogate handling) written by another tool.","commonSituations":"Strings written by non-Java encoders, data crossing system boundaries with different default charsets, or manually crafted/corrupted serialized payloads.","solutions":["Ensure both writer and reader use the same modified-UTF-8 encoding (writeLongUTF on the write side)","Do not feed raw UTF-8 or platform-encoded bytes into readLongUTF; re-encode via String.getBytes with a compatible scheme or rewrite with writeLongUTF","Validate/repair the payload bytes if coming from an external source","Add checksums at write time to catch corruption early"],"exampleFix":"// before\nout.write(str.getBytes(StandardCharsets.UTF_8)); // plain UTF-8, wrong framing\nString s = SerializerHelper.readLongUTF(in); // fails on continuation check\n// after\nSerializerHelper.writeLongUTF(out, str); // matched write/read pair\nString s = SerializerHelper.readLongUTF(in);","handlingStrategy":"try-catch","validationCode":"// validate payload is well-formed modified UTF-8 before decoding\nboolean validModifiedUtf8(byte[] b) { /* scan lead/continuation byte patterns */ }","typeGuard":"null","tryCatchPattern":"try {\n  return SerializerHelper.readLongUTF(in);\n} catch (UTFDataFormatException e) {\n  if (e.getMessage().startsWith(\"malformed input around byte\")) {\n    throw new CorruptPayloadException(\"Non-modified-UTF-8 bytes in stream; check writer encoding\", e);\n  }\n  throw e;\n}","preventionTips":["Encode strings only via writeLongUTF (modified UTF-8), never raw platform charsets","Keep reader aligned: read the length prefix first, then exactly that many bytes","Checksum payloads written across process/network boundaries"],"tags":["flink","serialization","utf-8","corrupt-data"],"backgroundTag":"invalid-argument-format","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}