{"record":{"id":"333e653fc14a53d8","repo":"pypa/pip","slug":"uncompilable-regex-tdef-0-r-in-state-state-r","errorCode":null,"errorMessage":"uncompilable regex {tdef[0]!r} in state {state!r} of {cls!r}: {err}","messagePattern":"uncompilable regex (.+?) in state (.+?) of (.+?): (.+?)","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"src/pip/_vendor/pygments/lexer.py","lineNumber":585,"sourceCode":"                tokens.extend(cls._process_state(unprocessed, processed,\n                                                 str(tdef)))\n                continue\n            if isinstance(tdef, _inherit):\n                # should be processed already, but may not in the case of:\n                # 1. the state has no counterpart in any parent\n                # 2. the state includes more than one 'inherit'\n                continue\n            if isinstance(tdef, default):\n                new_state = cls._process_new_state(tdef.state, unprocessed, processed)\n                tokens.append((re.compile('').match, None, new_state))\n                continue\n\n            assert type(tdef) is tuple, f\"wrong rule def {tdef!r}\"\n\n            try:\n                rex = cls._process_regex(tdef[0], rflags, state)\n            except Exception as err:\n                raise ValueError(f\"uncompilable regex {tdef[0]!r} in state {state!r} of {cls!r}: {err}\") from err\n\n            token = cls._process_token(tdef[1])\n\n            if len(tdef) == 2:\n                new_state = None\n            else:\n                new_state = cls._process_new_state(tdef[2],\n                                                   unprocessed, processed)\n\n            tokens.append((rex, token, new_state))\n        return tokens\n\n    def process_tokendef(cls, name, tokendefs=None):\n        \"\"\"Preprocess a dictionary of token definitions.\"\"\"\n        processed = cls._all_tokens[name] = {}\n        tokendefs = tokendefs or cls.tokens[name]\n        for state in list(tokendefs):\n            cls._process_state(tokendefs, processed, state)","sourceCodeStart":567,"sourceCodeEnd":603,"githubUrl":"https://github.com/pypa/pip/blob/f399c3718970b1b0e2478dac5296eb62679a9b86/src/pip/_vendor/pygments/lexer.py#L567-L603","documentation":"Raised during RegexLexer class construction (_process_state) when a token rule's regex string fails to compile. The exception from re.compile is wrapped in a ValueError naming the offending regex, the state, and the lexer class. This almost always indicates a bug in a custom lexer definition's tokens dict.","triggerScenarios":"Defining a custom RegexLexer subclass whose 'tokens' dict contains a tuple whose first element is an invalid regex (unbalanced parentheses, invalid escape, bad character class). Compilation happens at lexer.py:583 and is caught at :584-585.","commonSituations":"Typos in regex strings during custom lexer development; copy-paste of regexes that worked in another engine but not Python re; trailing backslash; unbalanced grouping.","solutions":["Test each regex in your tokens dict with re.compile() in isolation to find the bad one.","Fix the regex syntax (balance parens, correct escapes, valid character classes).","Use raw strings (r'...') for token regexes to avoid escape interpretation issues."],"exampleFix":"# before\nclass MyLexer(RegexLexer):\n    tokens = {'root': [(r'(unclosed', Text)]}\n# after\nclass MyLexer(RegexLexer):\n    tokens = {'root': [(r'\\(unclosed', Text)]}","handlingStrategy":"validation","validationCode":"import re\nfor state, rules in MyLexer.tokens.items():\n    for rule in rules:\n        if isinstance(rule, tuple):\n            re.compile(rule[0])  # raises early if invalid","typeGuard":"def all_rules_compile(tokens_dict) -> bool:\n    import re\n    for rules in tokens_dict.values():\n        for r in rules:\n            if isinstance(r, tuple):\n                try:\n                    re.compile(r[0])\n                except re.error:\n                    return False\n    return True","tryCatchPattern":"try:\n    class MyLexer(RegexLexer):\n        tokens = {...}\nexcept ValueError as e:\n    if 'uncompilable regex' in str(e):\n        # fix the offending regex then retry\n        raise\n    raise","preventionTips":["Use raw strings for token regexes.","Compile each regex standalone during lexer development.","Run a unit test that imports the lexer to surface compile errors at test time."],"tags":["python","pygments","lexer","regex","vendored"],"backgroundTag":null,"analyzedSha":"f399c3718970b1b0e2478dac5296eb62679a9b86","analyzedAt":"2026-08-08T23:01:42.227Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}