spring-projects/spring-security · error · IllegalArgumentException
Missing low surrogate character at end of string
Error message
Missing low surrogate character at end of string
What it means
TextEscapeUtils.escapeEntities() escapes non-ASCII characters of a string into XML numeric entities. When a high surrogate (the first half of a Unicode code point above U+FFFF) appears as the very last character of the string, there is no following low surrogate to complete the code point, so the method throws this IllegalArgumentException. The library refuses to emit a malformed half of a character.
Solutions
- Fix the truncation/processing logic so it never splits a surrogate pair (e.g. use an offset that lands on a code-point boundary, check Character.isHighSurrogate before the cut).
- Sanitize/replace lone surrogates before calling escapeEntities (e.g. s.replaceAll("[\\uD800-\\uDBFF](?![\\uDC00-\\uDFFF])", "")) or via CharsetEncoder with CodingErrorAction.REPLACE.
- If the input comes from a decoder, fix the source encoding step so code points are decoded completely (UTF-8 → UTF-16 conversion done correctly).
Example fix
// before String escaped = TextEscapeUtils.escapeEntities(input.substring(0, 255)); // after int end = Math.min(255, input.length()); if (end > 0 && Character.isHighSurrogate(input.charAt(end - 1))) end--; String escaped = TextEscapeUtils.escapeEntities(input.substring(0, end));
Defensive patterns
Strategy: validation
Validate before calling
public static boolean endsOnCompleteCodePoint(String s) {
return s.isEmpty() || !Character.isHighSurrogate(s.charAt(s.length() - 1));
}
// call: if (!endsOnCompleteCodePoint(input)) trim/repair before escapeEntities
Type guard
public static boolean isWellFormedUtf16(String s) {
int i = 0;
while (i < s.length()) {
char c = s.charAt(i);
if (Character.isHighSurrogate(c)) {
if (i + 1 >= s.length() || !Character.isLowSurrogate(s.charAt(i + 1))) return false;
i += 2;
} else if (Character.isLowSurrogate(c)) {
return false;
} else {
i++;
}
}
return true;
}
Try / catch
try {
String escaped = TextEscapeUtils.escapeEntities(input);
} catch (IllegalArgumentException e) {
// malformed surrogate — sanitize and retry
String clean = input.codePoints().filter(cp -> !Character.isSurrogate(cp))
.collect(StringBuilder::new, StringBuilder::appendCodePoint, StringBuilder::append).toString();
escaped = TextEscapeUtils.escapeEntities(clean);
}
Prevention
- Truncate strings by code points, not chars: use offsetByCodePoints rather than raw char indexes
- Never split strings at boundaries adjacent to a surrogate char without checking Character.isHighSurrogate/isLowSurrogate
- Prefer codePoints().stream() processing over char-by-char iteration for user data
- Run isWellFormedUtf16 validation on externally sourced strings before escaping
When it happens
Trigger: Calling TextEscapeUtils.escapeEntities(s) where s ends with a character in the range U+D800–U+DBFF (a high surrogate) with no trailing low surrogate, e.g. a string truncated mid-code-point by a fixed-length cut or an encoding/conversion bug.
Common situations: Strings truncated with substring(0, maxLen) slicing a supplementary character (emoji, CJK ext-B, rare symbols) in half; decoding malformed byte input with an error-tolerant charset; manual string surgery on UTF-16 data.
Understand the failure class
Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.
Related errors
- Expected low surrogate character but found value =
- Unexpected low surrogate character, value =
- Amount of performance parameters invalid
- authorizationManagerFactory must be an instance of…
- Bad number of rounds
AI-assisted analysis of spring-projects/spring-security@96852e8860 (2026-09-10).
Data as JSON: /api/errors/9c4bf032f4cf7aab.
Report an issue: GitHub.
Appendix: source
Thrown at web/src/main/java/org/springframework/security/web/util/TextEscapeUtils.java:57
else if (ch == '<') {
sb.append("<");
}
else if (ch == '>') {
sb.append(">");
}
else if (ch == '&') {
sb.append("&");
}
else if (Character.isWhitespace(ch)) {
sb.append("&#").append((int) ch).append(";");
}
else if (Character.isISOControl(ch)) {
// ignore control chars
}
else if (Character.isHighSurrogate(ch)) {
if (i + 1 >= s.length()) {
// Unexpected end
throw new IllegalArgumentException("Missing low surrogate character at end of string");
}
char low = s.charAt(i + 1);
if (!Character.isLowSurrogate(low)) {
throw new IllegalArgumentException(
"Expected low surrogate character but found value = " + (int) low);
}
int codePoint = Character.toCodePoint(ch, low);
if (Character.isDefined(codePoint)) {
sb.append("&#").append(codePoint).append(";");
}
i++; // skip the next character as we have already dealt with it
}
else if (Character.isLowSurrogate(ch)) {
throw new IllegalArgumentException("Unexpected low surrogate character, value = " + (int) ch);
}
else if (Character.isDefined(ch)) {
sb.append("&#").append((int) ch).append(";");
}View on GitHub (pinned to 96852e8860)