{"record":{"id":"9d415927e50d1ead","repo":"spring-projects/spring-security","slug":"unexpected-low-surrogate-character-value-ch","errorCode":null,"errorMessage":"Unexpected low surrogate character, value = <ch>","messagePattern":"Unexpected low surrogate character, value = <ch>","errorType":"exception","errorClass":"IllegalArgumentException","httpStatus":null,"severity":"error","filePath":"web/src/main/java/org/springframework/security/web/util/TextEscapeUtils.java","lineNumber":71,"sourceCode":"\t\t\t}\n\t\t\telse if (Character.isHighSurrogate(ch)) {\n\t\t\t\tif (i + 1 >= s.length()) {\n\t\t\t\t\t// Unexpected end\n\t\t\t\t\tthrow new IllegalArgumentException(\"Missing low surrogate character at end of string\");\n\t\t\t\t}\n\t\t\t\tchar low = s.charAt(i + 1);\n\t\t\t\tif (!Character.isLowSurrogate(low)) {\n\t\t\t\t\tthrow new IllegalArgumentException(\n\t\t\t\t\t\t\t\"Expected low surrogate character but found value = \" + (int) low);\n\t\t\t\t}\n\t\t\t\tint codePoint = Character.toCodePoint(ch, low);\n\t\t\t\tif (Character.isDefined(codePoint)) {\n\t\t\t\t\tsb.append(\"&#\").append(codePoint).append(\";\");\n\t\t\t\t}\n\t\t\t\ti++; // skip the next character as we have already dealt with it\n\t\t\t}\n\t\t\telse if (Character.isLowSurrogate(ch)) {\n\t\t\t\tthrow new IllegalArgumentException(\"Unexpected low surrogate character, value = \" + (int) ch);\n\t\t\t}\n\t\t\telse if (Character.isDefined(ch)) {\n\t\t\t\tsb.append(\"&#\").append((int) ch).append(\";\");\n\t\t\t}\n\t\t\t// Ignore anything else\n\t\t}\n\t\treturn sb.toString();\n\t}\n\n}\n","sourceCodeStart":53,"sourceCodeEnd":82,"githubUrl":"https://github.com/spring-projects/spring-security/blob/96852e8860138a482cb13d1479573f24ff6443c6/web/src/main/java/org/springframework/security/web/util/TextEscapeUtils.java#L53-L82","documentation":"TextEscapeUtils.escapeEntities() throws this when a low surrogate (U+DC00–U+DFFF) appears without a preceding high surrogate. A lone low surrogate is not a valid standalone character, so the method fails fast instead of escaping an invalid value.","triggerScenarios":"Calling escapeEntities(s) where s starts with or contains a low surrogate not preceded by a high surrogate — typically from slicing a string starting mid-pair (e.g. s.substring(1) where s starts with an emoji) or corrupt decoding.","commonSituations":"substring/indexOf offsets off by one relative to a supplementary character; concatenating string fragments split inside a code point; malformed data read from network or files.","solutions":["Correct the offset/split logic so string slices start and end on code-point boundaries (use offsetByCodePoints or check Character.isHighSurrogate at cut points).","Strip lone surrogates from input before escaping (regex or codePoints().filter).","Fix the decoder that produced invalid UTF-16 data at the source."],"exampleFix":"// before\nString escaped = TextEscapeUtils.escapeEntities(input.substring(1));\n// after\nint start = 1;\nif (Character.isLowSurrogate(input.charAt(start))) start--;\nString escaped = TextEscapeUtils.escapeEntities(input.substring(start));","handlingStrategy":"validation","validationCode":"public static boolean startsOnCodePointBoundary(String s) {\n    return s.isEmpty() || !Character.isLowSurrogate(s.charAt(0));\n}\n// apply to every substring/slice fed into escapeEntities\n","typeGuard":"public static String stripLoneLowSurrogates(String s) {\n    return s.codePoints().filter(cp -> !(cp >= 0xDC00 && cp <= 0xDFFF))\n        .collect(StringBuilder::new, StringBuilder::appendCodePoint, StringBuilder::append).toString();\n}\n","tryCatchPattern":"try {\n    escaped = TextEscapeUtils.escapeEntities(slice);\n} catch (IllegalArgumentException e) {\n    if (!e.getMessage().startsWith(\"Unexpected low surrogate\")) throw e;\n    escaped = TextEscapeUtils.escapeEntities(stripLoneLowSurrogates(slice));\n}\n","preventionTips":["When calling substring with an offset derived by counting chars, verify the boundary with Character.isLowSurrogate","Use s.offsetByCodePoints() to compute valid cut positions","Treat any low surrogate appearing at a string start or after a non-surrogate char as data corruption and fix the producer","Test string-processing code with emoji and other supplementary characters"],"tags":["java","string-encoding","unicode","xml-escaping"],"backgroundTag":"invalid-argument-format","analyzedSha":"96852e8860138a482cb13d1479573f24ff6443c6","analyzedAt":"2026-09-10T23:25:23.477Z","contentChangedAt":"2026-09-10T23:25:23.477Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}