{"record":{"id":"9c4bf032f4cf7aab","repo":"spring-projects/spring-security","slug":"missing-low-surrogate-character-at-end-of-string","errorCode":null,"errorMessage":"Missing low surrogate character at end of string","messagePattern":"Missing low surrogate character at end of string","errorType":"exception","errorClass":"IllegalArgumentException","httpStatus":null,"severity":"error","filePath":"web/src/main/java/org/springframework/security/web/util/TextEscapeUtils.java","lineNumber":57,"sourceCode":"\t\t\telse if (ch == '<') {\n\t\t\t\tsb.append(\"&lt;\");\n\t\t\t}\n\t\t\telse if (ch == '>') {\n\t\t\t\tsb.append(\"&gt;\");\n\t\t\t}\n\t\t\telse if (ch == '&') {\n\t\t\t\tsb.append(\"&amp;\");\n\t\t\t}\n\t\t\telse if (Character.isWhitespace(ch)) {\n\t\t\t\tsb.append(\"&#\").append((int) ch).append(\";\");\n\t\t\t}\n\t\t\telse if (Character.isISOControl(ch)) {\n\t\t\t\t// ignore control chars\n\t\t\t}\n\t\t\telse if (Character.isHighSurrogate(ch)) {\n\t\t\t\tif (i + 1 >= s.length()) {\n\t\t\t\t\t// Unexpected end\n\t\t\t\t\tthrow new IllegalArgumentException(\"Missing low surrogate character at end of string\");\n\t\t\t\t}\n\t\t\t\tchar low = s.charAt(i + 1);\n\t\t\t\tif (!Character.isLowSurrogate(low)) {\n\t\t\t\t\tthrow new IllegalArgumentException(\n\t\t\t\t\t\t\t\"Expected low surrogate character but found value = \" + (int) low);\n\t\t\t\t}\n\t\t\t\tint codePoint = Character.toCodePoint(ch, low);\n\t\t\t\tif (Character.isDefined(codePoint)) {\n\t\t\t\t\tsb.append(\"&#\").append(codePoint).append(\";\");\n\t\t\t\t}\n\t\t\t\ti++; // skip the next character as we have already dealt with it\n\t\t\t}\n\t\t\telse if (Character.isLowSurrogate(ch)) {\n\t\t\t\tthrow new IllegalArgumentException(\"Unexpected low surrogate character, value = \" + (int) ch);\n\t\t\t}\n\t\t\telse if (Character.isDefined(ch)) {\n\t\t\t\tsb.append(\"&#\").append((int) ch).append(\";\");\n\t\t\t}","sourceCodeStart":39,"sourceCodeEnd":75,"githubUrl":"https://github.com/spring-projects/spring-security/blob/96852e8860138a482cb13d1479573f24ff6443c6/web/src/main/java/org/springframework/security/web/util/TextEscapeUtils.java#L39-L75","documentation":"TextEscapeUtils.escapeEntities() escapes non-ASCII characters of a string into XML numeric entities. When a high surrogate (the first half of a Unicode code point above U+FFFF) appears as the very last character of the string, there is no following low surrogate to complete the code point, so the method throws this IllegalArgumentException. The library refuses to emit a malformed half of a character.","triggerScenarios":"Calling TextEscapeUtils.escapeEntities(s) where s ends with a character in the range U+D800–U+DBFF (a high surrogate) with no trailing low surrogate, e.g. a string truncated mid-code-point by a fixed-length cut or an encoding/conversion bug.","commonSituations":"Strings truncated with substring(0, maxLen) slicing a supplementary character (emoji, CJK ext-B, rare symbols) in half; decoding malformed byte input with an error-tolerant charset; manual string surgery on UTF-16 data.","solutions":["Fix the truncation/processing logic so it never splits a surrogate pair (e.g. use an offset that lands on a code-point boundary, check Character.isHighSurrogate before the cut).","Sanitize/replace lone surrogates before calling escapeEntities (e.g. s.replaceAll(\"[\\\\uD800-\\\\uDBFF](?![\\\\uDC00-\\\\uDFFF])\", \"\")) or via CharsetEncoder with CodingErrorAction.REPLACE.","If the input comes from a decoder, fix the source encoding step so code points are decoded completely (UTF-8 → UTF-16 conversion done correctly)."],"exampleFix":"// before\nString escaped = TextEscapeUtils.escapeEntities(input.substring(0, 255));\n// after\nint end = Math.min(255, input.length());\nif (end > 0 && Character.isHighSurrogate(input.charAt(end - 1))) end--;\nString escaped = TextEscapeUtils.escapeEntities(input.substring(0, end));","handlingStrategy":"validation","validationCode":"public static boolean endsOnCompleteCodePoint(String s) {\n    return s.isEmpty() || !Character.isHighSurrogate(s.charAt(s.length() - 1));\n}\n// call: if (!endsOnCompleteCodePoint(input)) trim/repair before escapeEntities\n","typeGuard":"public static boolean isWellFormedUtf16(String s) {\n    int i = 0;\n    while (i < s.length()) {\n        char c = s.charAt(i);\n        if (Character.isHighSurrogate(c)) {\n            if (i + 1 >= s.length() || !Character.isLowSurrogate(s.charAt(i + 1))) return false;\n            i += 2;\n        } else if (Character.isLowSurrogate(c)) {\n            return false;\n        } else {\n            i++;\n        }\n    }\n    return true;\n}\n","tryCatchPattern":"try {\n    String escaped = TextEscapeUtils.escapeEntities(input);\n} catch (IllegalArgumentException e) {\n    // malformed surrogate — sanitize and retry\n    String clean = input.codePoints().filter(cp -> !Character.isSurrogate(cp))\n        .collect(StringBuilder::new, StringBuilder::appendCodePoint, StringBuilder::append).toString();\n    escaped = TextEscapeUtils.escapeEntities(clean);\n}\n","preventionTips":["Truncate strings by code points, not chars: use offsetByCodePoints rather than raw char indexes","Never split strings at boundaries adjacent to a surrogate char without checking Character.isHighSurrogate/isLowSurrogate","Prefer codePoints().stream() processing over char-by-char iteration for user data","Run isWellFormedUtf16 validation on externally sourced strings before escaping"],"tags":["java","string-encoding","unicode","xml-escaping"],"backgroundTag":"invalid-argument-format","analyzedSha":"96852e8860138a482cb13d1479573f24ff6443c6","analyzedAt":"2026-09-10T23:25:23.477Z","contentChangedAt":"2026-09-10T23:25:23.477Z","schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}