{"record":{"id":"cab0c02ff7b49f47","repo":"apache/iceberg","slug":"malformed-input-around-byte-cab0c0","errorCode":null,"errorMessage":"malformed input around byte ","messagePattern":"malformed input around byte ","errorType":"exception","errorClass":"UTFDataFormatException","httpStatus":null,"severity":"error","filePath":"flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java","lineNumber":138,"sourceCode":"        case 3:\n        case 4:\n        case 5:\n        case 6:\n        case 7:\n          /* 0xxxxxxx */\n          count++;\n          chararr[chararrCount++] = (char) ch;\n          break;\n        case 12:\n        case 13:\n          /* 110x xxxx 10xx xxxx */\n          count += 2;\n          if (count > utflen) {\n            throw new UTFDataFormatException(\"malformed input: partial character at end\");\n          }\n          char2 = bytearr[count - 1];\n          if ((char2 & 0xC0) != 0x80) {\n            throw new UTFDataFormatException(\"malformed input around byte \" + count);\n          }\n          chararr[chararrCount++] = (char) (((ch & 0x1F) << 6) | (char2 & 0x3F));\n          break;\n        case 14:\n          /* 1110 xxxx 10xx xxxx 10xx xxxx */\n          count += 3;\n          if (count > utflen) {\n            throw new UTFDataFormatException(\"malformed input: partial character at end\");\n          }\n          char2 = bytearr[count - 2];\n          char3 = bytearr[count - 1];\n          if (((char2 & 0xC0) != 0x80) || ((char3 & 0xC0) != 0x80)) {\n            throw new UTFDataFormatException(\"malformed input around byte \" + (count - 1));\n          }\n          chararr[chararrCount++] =\n              (char) (((ch & 0x0F) << 12) | ((char2 & 0x3F) << 6) | (char3 & 0x3F));\n          break;\n        default:","sourceCodeStart":120,"sourceCodeEnd":156,"githubUrl":"https://github.com/apache/iceberg/blob/86d9c8fc543e7c56c9f624eb725f76c9baff9570/flink/v2.2/flink/src/main/java/org/apache/iceberg/flink/util/SerializerHelper.java#L120-L156","documentation":"SerializerHelper.readLongUTF decodes a Java-modified-UTF string from serialized bytes. When a multi-byte character starts with a 2-byte lead (110xxxxx) but the following byte does not have the continuation prefix 10xxxxxx, the stream is not valid modified UTF-8 and a UTFDataFormatException with the byte offset is thrown. This protects callers from silently decoding corrupted or misaligned payload data.","triggerScenarios":"Calling readLongUTF on bytes where a 2-byte UTF lead byte is followed by a non-continuation byte; typically the byte array was produced by a different serializer, truncated, or the length prefix was read at the wrong offset.","commonSituations":"Flink job state deserialized across incompatible serializer versions; hand-rolled binary formats mixing standard UTF-8 with Java modified UTF; corrupted Kafka/checkpoint payloads; reading past a field boundary so the decoder starts mid-character.","solutions":["Verify the bytes were written by SerializerHelper.writeLongUTF (or DataOutput.writeUTF) with matching read/write order and no offset drift.","Dump the byte range around the reported offset and check that continuation bytes (0x80-0xBF) follow each lead byte.","Ensure sender and reader use the same Iceberg/Flink serializer versions for the affected state descriptor.","If data is externally produced, transcode it to Java modified UTF-8 (or use a standard String serializer) before decoding."],"exampleFix":"// before\nString value = SerializerHelper.readLongUTF(in); // assumes arbitrary UTF-8 bytes\n// after\n// write side: ensure symmetric API\nSerializerHelper.writeLongUTF(out, value);\n// read side: only consume bytes produced by writeLongUTF; validate provenance/version first\nString value = SerializerHelper.readLongUTF(in);","handlingStrategy":"validation","validationCode":"// verify provenance before decoding: bytes must come from writeLongUTF/writeUTF\nif (payload == null || payload.length == 0) throw new IllegalArgumentException(\"empty payload\");\n// spot-check: byte at any lead position must be a valid lead (not 0x80-0xBF, not 0xF0-0xFF)\nfor (int i = 0; i < payload.length; ) {\n  int b = payload[i] & 0xFF;\n  if ((b & 0xC0) == 0x80 || (b & 0xF8) == 0xF0) throw new IllegalArgumentException(\"invalid UTF lead at \" + i);\n  i += (b & 0x80) == 0 ? 1 : (b & 0xE0) == 0xC0 ? 2 : (b & 0xF0) == 0xE0 ? 3 : 1;\n}","typeGuard":null,"tryCatchPattern":"try {\n  String s = SerializerHelper.readLongUTF(in);\n} catch (UTFDataFormatException e) {\n  throw new IOException(\"corrupt string state at offset \" + e.getMessage(), e);\n}","preventionTips":["Always pair writeLongUTF with readLongUTF from the same serializer version","Never decode externally produced UTF-8 bytes with the modified-UTF reader","Log the serializer version alongside serialized state for drift diagnosis"],"tags":["flink","serialization","utf-data-format"],"backgroundTag":"malformed-utf-input","analyzedSha":"86d9c8fc543e7c56c9f624eb725f76c9baff9570","analyzedAt":"2026-09-12T00:46:39.097Z","contentChangedAt":"2026-09-12T00:46:39.097Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}