{"record":{"id":"db1e87b46b052672","repo":"pandas-dev/pandas","slug":"could-not-convert-name-to-a-valid-python-ident","errorCode":null,"errorMessage":"Could not convert '{name}' to a valid Python identifier.","messagePattern":"Could not convert '(.+?)' to a valid Python identifier\\.","errorType":"exception","errorClass":"SyntaxError","httpStatus":null,"severity":"error","filePath":"pandas/core/computation/parsing.py","lineNumber":77,"sourceCode":"        {\n            \" \": \"_\",\n            \"?\": \"_QUESTIONMARK_\",\n            \"!\": \"_EXCLAMATIONMARK_\",\n            \"$\": \"_DOLLARSIGN_\",\n            \"€\": \"_EUROSIGN_\",\n            \"°\": \"_DEGREESIGN_\",\n            \"'\": \"_SINGLEQUOTE_\",\n            '\"': \"_DOUBLEQUOTE_\",\n            \"#\": \"_HASH_\",\n            \"`\": \"_BACKTICK_\",\n        }\n    )\n\n    name = \"\".join([special_characters_replacements.get(char, char) for char in name])\n    name = f\"BACKTICK_QUOTED_STRING_{name}\"\n\n    if not name.isidentifier():\n        raise SyntaxError(f\"Could not convert '{name}' to a valid Python identifier.\")\n\n    return name\n\n\ndef clean_backtick_quoted_toks(tok: tuple[int, str]) -> tuple[int, str]:\n    \"\"\"\n    Clean up a column name if surrounded by backticks.\n\n    Backtick quoted string are indicated by a certain tokval value. If a string\n    is a backtick quoted token it will processed by\n    :func:`_create_valid_python_identifier` so that the parser can find this\n    string when the query is executed.\n    In this case the tok will get the NAME tokval.\n\n    Parameters\n    ----------\n    tok : tuple of int, str\n        ints correspond to the all caps constants in the tokenize module","sourceCodeStart":59,"sourceCodeEnd":95,"githubUrl":"https://github.com/pandas-dev/pandas/blob/3b7651241d4da534b3559b60ef128e1c34f54116/pandas/core/computation/parsing.py#L59-L95","documentation":"Raised in create_valid_python_identifier after every replacement pass still yields a string that fails str.isidentifier(). The function escapes non-ASCII (via backslashreplace + _UNICODE_/), maps special characters to token-name placeholders (e.g. '(' -> _LRB_), prefixes the result with 'BACKTICK_QUOTED_STRING_', and as a last resort raises SyntaxError. This powers the backtick-quoted column-name feature of query(); a failure means the column name contains a character the escape table cannot normalize into a valid identifier.","triggerScenarios":"df.query('`<weird col>` > 1') where the column name, after all substitutions, still contains a character that breaks identifier rules — e.g. certain Unicode combining marks, control characters, or characters not covered by tokenize.EXACT_TOKEN_TYPES or the manual override dict.","commonSituations":"Column names generated from external data (databases, CSVs) containing exotic Unicode (combining diacritics, zero-width joiners, emoji with variation selectors). Column names with leading digits after the BACKTICK_QUOTED_STRING_ prefix where a substitution leaves an invalid start character. pandas version upgrades that changed the escape table (GH 49633).","solutions":["Sanitize column names before querying: df.columns = df.columns.str.replace(r'[^\\w]', '_', regex=True).","Avoid backtick-quoted columns in query; use boolean indexing instead: df[df['<weird col>'] > 1].","Identify and remove the offending character: print([c for c in col if not ('_'+c).isidentifier()]).","Upgrade pandas — recent versions expanded the escape table (GH 49633) to cover more characters."],"exampleFix":"// before\ndf.query(\"`col\\u200bname` > 1\")  # zero-width space in name\n\n// after\ndf.columns = df.columns.str.replace(r\"\\W\", \"_\", regex=True)\ndf.query(\"`col_name` > 1\")\n# or skip query\ndf[df[\"col\\u200bname\"] > 1]","handlingStrategy":"validation","validationCode":"from pandas.core.computation.parsing import create_valid_python_identifier\n\ndef column_is_query_safe(col: str) -> bool:\n    try:\n        create_valid_python_identifier(col)\n        return True\n    except SyntaxError:\n        return False\n\nbad = [c for c in df.columns if not column_is_query_safe(c)]\nassert not bad, f'columns not query-safe: {bad}'","typeGuard":"from pandas.core.computation.parsing import create_valid_python_identifier\n\ndef is_query_safe_column(name: str) -> bool:\n    try:\n        create_valid_python_identifier(name)\n        return True\n    except SyntaxError:\n        return False","tryCatchPattern":"try:\n    df.query(f'`{col}` > 1')\nexcept SyntaxError as e:\n    if 'Could not convert' in str(e):\n        df = df.rename(columns={col: col.replace(' ', '_')})\n        df.query(f'`{col.replace(\" \", \"_\")}` > 1')\n    else:\n        raise","preventionTips":["Sanitize column names on ingestion: df.columns = df.columns.str.replace(r'\\W', '_', regex=True).","Validate every column name with create_valid_python_identifier before writing query strings.","Prefer boolean indexing df[df[col] > 1] for column names with exotic Unicode.","Keep pandas current — escape-table coverage improves with releases (GH 49633)."],"tags":["pandas","query","column-name","identifier","unicode","backtick"],"backgroundTag":null,"analyzedSha":"3b7651241d4da534b3559b60ef128e1c34f54116","analyzedAt":"2026-08-11T22:10:44.015Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}