pypa/pip · error · ValueError
uncompilable regex in state of
Error message
uncompilable regex {tdef[0]!r} in state {state!r} of {cls!r}: {err} What it means
Raised during RegexLexer class construction (_process_state) when a token rule's regex string fails to compile. The exception from re.compile is wrapped in a ValueError naming the offending regex, the state, and the lexer class. This almost always indicates a bug in a custom lexer definition's tokens dict.
Solutions
- Test each regex in your tokens dict with re.compile() in isolation to find the bad one.
- Fix the regex syntax (balance parens, correct escapes, valid character classes).
- Use raw strings (r'...') for token regexes to avoid escape interpretation issues.
Example fix
# before
class MyLexer(RegexLexer):
tokens = {'root': [(r'(unclosed', Text)]}
# after
class MyLexer(RegexLexer):
tokens = {'root': [(r'\(unclosed', Text)]} Defensive patterns
Strategy: validation
Validate before calling
import re
for state, rules in MyLexer.tokens.items():
for rule in rules:
if isinstance(rule, tuple):
re.compile(rule[0]) # raises early if invalid Type guard
def all_rules_compile(tokens_dict) -> bool:
import re
for rules in tokens_dict.values():
for r in rules:
if isinstance(r, tuple):
try:
re.compile(r[0])
except re.error:
return False
return True Try / catch
try:
class MyLexer(RegexLexer):
tokens = {...}
except ValueError as e:
if 'uncompilable regex' in str(e):
# fix the offending regex then retry
raise
raise Prevention
- Use raw strings for token regexes.
- Compile each regex standalone during lexer development.
- Run a unit test that imports the lexer to surface compile errors at test time.
When it happens
Trigger: Defining a custom RegexLexer subclass whose 'tokens' dict contains a tuple whose first element is an invalid regex (unbalanced parentheses, invalid escape, bad character class). Compilation happens at lexer.py:583 and is caught at :584-585.
Common situations: Typos in regex strings during custom lexer development; copy-paste of regexes that worked in another engine but not Python re; trailing backslash; unbalanced grouping.
Related errors
- No such group
- lex() argument must be a lexer instance, not a class
- no lexer for alias found
- To enable chardet encoding guessing, please install the…
- cannot read
AI-assisted analysis of pypa/pip@f399c37189 (2026-08-08).
Data as JSON: /api/errors/333e653fc14a53d8.
Report an issue: GitHub.
Appendix: source
Thrown at src/pip/_vendor/pygments/lexer.py:585
tokens.extend(cls._process_state(unprocessed, processed,
str(tdef)))
continue
if isinstance(tdef, _inherit):
# should be processed already, but may not in the case of:
# 1. the state has no counterpart in any parent
# 2. the state includes more than one 'inherit'
continue
if isinstance(tdef, default):
new_state = cls._process_new_state(tdef.state, unprocessed, processed)
tokens.append((re.compile('').match, None, new_state))
continue
assert type(tdef) is tuple, f"wrong rule def {tdef!r}"
try:
rex = cls._process_regex(tdef[0], rflags, state)
except Exception as err:
raise ValueError(f"uncompilable regex {tdef[0]!r} in state {state!r} of {cls!r}: {err}") from err
token = cls._process_token(tdef[1])
if len(tdef) == 2:
new_state = None
else:
new_state = cls._process_new_state(tdef[2],
unprocessed, processed)
tokens.append((rex, token, new_state))
return tokens
def process_tokendef(cls, name, tokendefs=None):
"""Preprocess a dictionary of token definitions."""
processed = cls._all_tokens[name] = {}
tokendefs = tokendefs or cls.tokens[name]
for state in list(tokendefs):
cls._process_state(tokendefs, processed, state)View on GitHub (pinned to f399c37189)