antlr/antlr4 · critical · RuntimeException
Invalid UTF-16 (dangling high surrogate
Error message
Invalid UTF-16 (dangling high surrogate
What it means
UnbufferedCharStream read a high surrogate but the following character was not a low surrogate (and not EOF), so the pair is dangling. The stream throws because it cannot form a code point and will not store unpaired surrogates. As with the related errors, the usual root cause is an encoding mismatch between the Reader and the actual bytes.
Solutions
- Open the Reader with the explicit correct charset (StandardCharsets.UTF_8 for typical text files)
- Validate/sanitize the input for unpaired surrogates before lexing if the source is untrusted
- Re-encode the file to UTF-8 once, then always read it as UTF-8
Example fix
// before new UnbufferedCharStream(new FileReader(f)); // platform-default charset // after new UnbufferedCharStream(Files.newBufferedReader(f.toPath(), StandardCharsets.UTF_8));
Defensive patterns
Strategy: try-catch
Validate before calling
static boolean hasUnpairedSurrogates(String s) {
for (int i = 0; i < s.length(); i++) {
char c = s.charAt(i);
if (Character.isHighSurrogate(c)
&& (i + 1 >= s.length() || !Character.isLowSurrogate(s.charAt(i + 1)))) return true;
if (Character.isLowSurrogate(c)
&& (i == 0 || !Character.isHighSurrogate(s.charAt(i - 1)))) return true;
}
return false;
} Try / catch
try { parse(stream); }
catch (RuntimeException e) {
if (e.getMessage() != null && e.getMessage().contains("UTF-16"))
return Result.badEncoding();
throw e;
} Prevention
- Read files via Files.newBufferedReader with an explicit charset
- Never build Readers with the platform default charset
- Sanitize untrusted text for unpaired surrogates before lexing
When it happens
Trigger: UTF-8 or single-byte content opened through a UTF-16 Reader so random adjacent characters land after a 0xD800..0xDBFF value; content containing literal escaped surrogates; transcode pipelines that split pairs.
Common situations: Default-charset Readers on JVMs where the platform encoding differs from the file's; concatenating strings that each hold half of a surrogate pair.
Understand the failure class
- Parsing and encoding errors: unexpected token, malformed input — why parsers reject input and how to find the real culprit.
Related errors
- Invalid UTF-16 (low surrogate with no preceding high…
- Invalid UTF-16 (dangling high surrogate at end of file)
- Invalid UTF-16 (high surrogate followed by code point >…
- cannot consume EOF
- cannot consume EOF
AI-assisted analysis of antlr/antlr4@7d5770395b (2026-08-14).
Data as JSON: /api/errors/1daa38031773049f.
Report an issue: GitHub.
Appendix: source
Thrown at runtime/Java/src/org/antlr/v4/runtime/UnbufferedCharStream.java:188
char ch = (char) c;
if (Character.isLowSurrogate(ch)) {
throw new RuntimeException("Invalid UTF-16 (low surrogate with no preceding high surrogate)");
}
else if (Character.isHighSurrogate(ch)) {
int lowSurrogate = nextChar();
if (lowSurrogate > Character.MAX_VALUE) {
throw new RuntimeException("Invalid UTF-16 (high surrogate followed by code point > U+FFFF");
}
else if (lowSurrogate == IntStream.EOF) {
throw new RuntimeException("Invalid UTF-16 (dangling high surrogate at end of file)");
}
else {
char lowSurrogateChar = (char) lowSurrogate;
if (Character.isLowSurrogate(lowSurrogateChar)) {
add(Character.toCodePoint(ch, lowSurrogateChar));
}
else {
throw new RuntimeException("Invalid UTF-16 (dangling high surrogate");
}
}
}
else {
add(c);
}
}
}
catch (IOException ioe) {
throw new RuntimeException(ioe);
}
}
return n;
}
/**
* Override to provide different source of characters thanView on GitHub (pinned to 7d5770395b)