spring-projects/spring-security · error · IllegalArgumentException

Missing low surrogate character at end of string

Error message

Missing low surrogate character at end of string

What it means

TextEscapeUtils.escapeEntities() escapes non-ASCII characters of a string into XML numeric entities. When a high surrogate (the first half of a Unicode code point above U+FFFF) appears as the very last character of the string, there is no following low surrogate to complete the code point, so the method throws this IllegalArgumentException. The library refuses to emit a malformed half of a character.

Solutions

  1. Fix the truncation/processing logic so it never splits a surrogate pair (e.g. use an offset that lands on a code-point boundary, check Character.isHighSurrogate before the cut).
  2. Sanitize/replace lone surrogates before calling escapeEntities (e.g. s.replaceAll("[\\uD800-\\uDBFF](?![\\uDC00-\\uDFFF])", "")) or via CharsetEncoder with CodingErrorAction.REPLACE.
  3. If the input comes from a decoder, fix the source encoding step so code points are decoded completely (UTF-8 → UTF-16 conversion done correctly).

Example fix

// before
String escaped = TextEscapeUtils.escapeEntities(input.substring(0, 255));
// after
int end = Math.min(255, input.length());
if (end > 0 && Character.isHighSurrogate(input.charAt(end - 1))) end--;
String escaped = TextEscapeUtils.escapeEntities(input.substring(0, end));
Defensive patterns

Strategy: validation

Validate before calling

public static boolean endsOnCompleteCodePoint(String s) {
    return s.isEmpty() || !Character.isHighSurrogate(s.charAt(s.length() - 1));
}
// call: if (!endsOnCompleteCodePoint(input)) trim/repair before escapeEntities

Type guard

public static boolean isWellFormedUtf16(String s) {
    int i = 0;
    while (i < s.length()) {
        char c = s.charAt(i);
        if (Character.isHighSurrogate(c)) {
            if (i + 1 >= s.length() || !Character.isLowSurrogate(s.charAt(i + 1))) return false;
            i += 2;
        } else if (Character.isLowSurrogate(c)) {
            return false;
        } else {
            i++;
        }
    }
    return true;
}

Try / catch

try {
    String escaped = TextEscapeUtils.escapeEntities(input);
} catch (IllegalArgumentException e) {
    // malformed surrogate — sanitize and retry
    String clean = input.codePoints().filter(cp -> !Character.isSurrogate(cp))
        .collect(StringBuilder::new, StringBuilder::appendCodePoint, StringBuilder::append).toString();
    escaped = TextEscapeUtils.escapeEntities(clean);
}

Prevention

When it happens

Trigger: Calling TextEscapeUtils.escapeEntities(s) where s ends with a character in the range U+D800–U+DBFF (a high surrogate) with no trailing low surrogate, e.g. a string truncated mid-code-point by a fixed-length cut or an encoding/conversion bug.

Common situations: Strings truncated with substring(0, maxLen) slicing a supplementary character (emoji, CJK ext-B, rare symbols) in half; decoding malformed byte input with an error-tolerant charset; manual string surgery on UTF-16 data.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of spring-projects/spring-security@96852e8860 (2026-09-10). Data as JSON: /api/errors/9c4bf032f4cf7aab. Report an issue: GitHub.

Appendix: source

Thrown at web/src/main/java/org/springframework/security/web/util/TextEscapeUtils.java:57

			else if (ch == '<') {
				sb.append("&lt;");
			}
			else if (ch == '>') {
				sb.append("&gt;");
			}
			else if (ch == '&') {
				sb.append("&amp;");
			}
			else if (Character.isWhitespace(ch)) {
				sb.append("&#").append((int) ch).append(";");
			}
			else if (Character.isISOControl(ch)) {
				// ignore control chars
			}
			else if (Character.isHighSurrogate(ch)) {
				if (i + 1 >= s.length()) {
					// Unexpected end
					throw new IllegalArgumentException("Missing low surrogate character at end of string");
				}
				char low = s.charAt(i + 1);
				if (!Character.isLowSurrogate(low)) {
					throw new IllegalArgumentException(
							"Expected low surrogate character but found value = " + (int) low);
				}
				int codePoint = Character.toCodePoint(ch, low);
				if (Character.isDefined(codePoint)) {
					sb.append("&#").append(codePoint).append(";");
				}
				i++; // skip the next character as we have already dealt with it
			}
			else if (Character.isLowSurrogate(ch)) {
				throw new IllegalArgumentException("Unexpected low surrogate character, value = " + (int) ch);
			}
			else if (Character.isDefined(ch)) {
				sb.append("&#").append((int) ch).append(";");
			}

View on GitHub (pinned to 96852e8860)