pypa/pip · error · ValueError

uncompilable regex in state of

Error message

uncompilable regex {tdef[0]!r} in state {state!r} of {cls!r}: {err}

What it means

Raised during RegexLexer class construction (_process_state) when a token rule's regex string fails to compile. The exception from re.compile is wrapped in a ValueError naming the offending regex, the state, and the lexer class. This almost always indicates a bug in a custom lexer definition's tokens dict.

Solutions

  1. Test each regex in your tokens dict with re.compile() in isolation to find the bad one.
  2. Fix the regex syntax (balance parens, correct escapes, valid character classes).
  3. Use raw strings (r'...') for token regexes to avoid escape interpretation issues.

Example fix

# before
class MyLexer(RegexLexer):
    tokens = {'root': [(r'(unclosed', Text)]}
# after
class MyLexer(RegexLexer):
    tokens = {'root': [(r'\(unclosed', Text)]}
Defensive patterns

Strategy: validation

Validate before calling

import re
for state, rules in MyLexer.tokens.items():
    for rule in rules:
        if isinstance(rule, tuple):
            re.compile(rule[0])  # raises early if invalid

Type guard

def all_rules_compile(tokens_dict) -> bool:
    import re
    for rules in tokens_dict.values():
        for r in rules:
            if isinstance(r, tuple):
                try:
                    re.compile(r[0])
                except re.error:
                    return False
    return True

Try / catch

try:
    class MyLexer(RegexLexer):
        tokens = {...}
except ValueError as e:
    if 'uncompilable regex' in str(e):
        # fix the offending regex then retry
        raise
    raise

Prevention

When it happens

Trigger: Defining a custom RegexLexer subclass whose 'tokens' dict contains a tuple whose first element is an invalid regex (unbalanced parentheses, invalid escape, bad character class). Compilation happens at lexer.py:583 and is caught at :584-585.

Common situations: Typos in regex strings during custom lexer development; copy-paste of regexes that worked in another engine but not Python re; trailing backslash; unbalanced grouping.

Related errors


AI-assisted analysis of pypa/pip@f399c37189 (2026-08-08). Data as JSON: /api/errors/333e653fc14a53d8. Report an issue: GitHub.

Appendix: source

Thrown at src/pip/_vendor/pygments/lexer.py:585

                tokens.extend(cls._process_state(unprocessed, processed,
                                                 str(tdef)))
                continue
            if isinstance(tdef, _inherit):
                # should be processed already, but may not in the case of:
                # 1. the state has no counterpart in any parent
                # 2. the state includes more than one 'inherit'
                continue
            if isinstance(tdef, default):
                new_state = cls._process_new_state(tdef.state, unprocessed, processed)
                tokens.append((re.compile('').match, None, new_state))
                continue

            assert type(tdef) is tuple, f"wrong rule def {tdef!r}"

            try:
                rex = cls._process_regex(tdef[0], rflags, state)
            except Exception as err:
                raise ValueError(f"uncompilable regex {tdef[0]!r} in state {state!r} of {cls!r}: {err}") from err

            token = cls._process_token(tdef[1])

            if len(tdef) == 2:
                new_state = None
            else:
                new_state = cls._process_new_state(tdef[2],
                                                   unprocessed, processed)

            tokens.append((rex, token, new_state))
        return tokens

    def process_tokendef(cls, name, tokendefs=None):
        """Preprocess a dictionary of token definitions."""
        processed = cls._all_tokens[name] = {}
        tokendefs = tokendefs or cls.tokens[name]
        for state in list(tokendefs):
            cls._process_state(tokendefs, processed, state)

View on GitHub (pinned to f399c37189)