spring-projects/spring-security · error · IllegalArgumentException

Expected low surrogate character but found value =

Error message

Expected low surrogate character but found value = <low>

What it means

TextEscapeUtils.escapeEntities() throws this when it encounters a high surrogate whose immediately following character exists but is not a valid low surrogate — the pair cannot be combined into a Unicode code point. The numeric char value of the offending character is included in the message.

Solutions

  1. Fix the upstream code that corrupts surrogate pairs — never split or reassemble strings between the two halves of a pair.
  2. Sanitize the input first: remove or replace lone surrogates (e.g. with a regex or CharsetEncoder REPLACE) before calling escapeEntities.
  3. If the char is expected, use Character.toCodePoint / codePointAt-based processing so pairs travel together.

Example fix

// before
String escaped = TextEscapeUtils.escapeEntities(corruptedInput);
// after
String clean = corruptedInput.codePoints()
    .filter(cp -> !Character.isSurrogate(cp))
    .collect(StringBuilder::new, StringBuilder::appendCodePoint, StringBuilder::append)
    .toString();
String escaped = TextEscapeUtils.escapeEntities(clean);
Defensive patterns

Strategy: validation

Validate before calling

public static boolean hasNoLoneSurrogates(String s) {
    for (int i = 0; i < s.length(); i++) {
        char c = s.charAt(i);
        if (Character.isHighSurrogate(c)) {
            if (i + 1 >= s.length() || !Character.isLowSurrogate(s.charAt(i + 1))) return false;
            i++;
        } else if (Character.isLowSurrogate(c)) return false;
    }
    return true;
}

Type guard

public static String dropMalformedPairs(String s) {
    return s.codePoints().filter(cp -> !Character.isSurrogate(cp))
        .collect(StringBuilder::new, StringBuilder::appendCodePoint, StringBuilder::append).toString();
}

Try / catch

try {
    escaped = TextEscapeUtils.escapeEntities(input);
} catch (IllegalArgumentException e) {
    if (!e.getMessage().startsWith("Expected low surrogate") && !e.getMessage().startsWith("Missing low surrogate")
        && !e.getMessage().startsWith("Unexpected low surrogate")) throw e;
    escaped = TextEscapeUtils.escapeEntities(dropMalformedPairs(input));
}

Prevention

When it happens

Trigger: Calling escapeEntities(s) where s contains a high surrogate (U+D800–U+DBFF) followed by a non-surrogate char (e.g. '\uD83D' followed by 'a'), produced by corrupted UTF-16 data, buggy character-level edits, or concatenated fragments of different strings.

Common situations: Manually assembled strings where an emoji was copied without its second char; regex/matcher results cut at surrogate boundaries; strings mutated by tools unaware of UTF-16 pairs.

Understand the failure class

Background: "Invalid ... format", "must be in format X", "does not look like a ..." — invalid argument format errors across CLI tools and libraries — this error's family across 17 libraries.

Related errors


AI-assisted analysis of spring-projects/spring-security@96852e8860 (2026-09-10). Data as JSON: /api/errors/187358f40dd01173. Report an issue: GitHub.

Appendix: source

Thrown at web/src/main/java/org/springframework/security/web/util/TextEscapeUtils.java:61

				sb.append("&gt;");
			}
			else if (ch == '&') {
				sb.append("&amp;");
			}
			else if (Character.isWhitespace(ch)) {
				sb.append("&#").append((int) ch).append(";");
			}
			else if (Character.isISOControl(ch)) {
				// ignore control chars
			}
			else if (Character.isHighSurrogate(ch)) {
				if (i + 1 >= s.length()) {
					// Unexpected end
					throw new IllegalArgumentException("Missing low surrogate character at end of string");
				}
				char low = s.charAt(i + 1);
				if (!Character.isLowSurrogate(low)) {
					throw new IllegalArgumentException(
							"Expected low surrogate character but found value = " + (int) low);
				}
				int codePoint = Character.toCodePoint(ch, low);
				if (Character.isDefined(codePoint)) {
					sb.append("&#").append(codePoint).append(";");
				}
				i++; // skip the next character as we have already dealt with it
			}
			else if (Character.isLowSurrogate(ch)) {
				throw new IllegalArgumentException("Unexpected low surrogate character, value = " + (int) ch);
			}
			else if (Character.isDefined(ch)) {
				sb.append("&#").append((int) ch).append(";");
			}
			// Ignore anything else
		}
		return sb.toString();
	}

View on GitHub (pinned to 96852e8860)