pandas-dev/pandas · error · SyntaxError
Could not convert ' ' to a valid Python identifier.
Error message
Could not convert '{name}' to a valid Python identifier. What it means
Raised in create_valid_python_identifier after every replacement pass still yields a string that fails str.isidentifier(). The function escapes non-ASCII (via backslashreplace + _UNICODE_/), maps special characters to token-name placeholders (e.g. '(' -> _LRB_), prefixes the result with 'BACKTICK_QUOTED_STRING_', and as a last resort raises SyntaxError. This powers the backtick-quoted column-name feature of query(); a failure means the column name contains a character the escape table cannot normalize into a valid identifier.
Solutions
- Sanitize column names before querying: df.columns = df.columns.str.replace(r'[^\w]', '_', regex=True).
- Avoid backtick-quoted columns in query; use boolean indexing instead: df[df['<weird col>'] > 1].
- Identify and remove the offending character: print([c for c in col if not ('_'+c).isidentifier()]).
- Upgrade pandas — recent versions expanded the escape table (GH 49633) to cover more characters.
Example fix
// before
df.query("`col\u200bname` > 1") # zero-width space in name
// after
df.columns = df.columns.str.replace(r"\W", "_", regex=True)
df.query("`col_name` > 1")
# or skip query
df[df["col\u200bname"] > 1] Defensive patterns
Strategy: validation
Validate before calling
from pandas.core.computation.parsing import create_valid_python_identifier
def column_is_query_safe(col: str) -> bool:
try:
create_valid_python_identifier(col)
return True
except SyntaxError:
return False
bad = [c for c in df.columns if not column_is_query_safe(c)]
assert not bad, f'columns not query-safe: {bad}' Type guard
from pandas.core.computation.parsing import create_valid_python_identifier
def is_query_safe_column(name: str) -> bool:
try:
create_valid_python_identifier(name)
return True
except SyntaxError:
return False Try / catch
try:
df.query(f'`{col}` > 1')
except SyntaxError as e:
if 'Could not convert' in str(e):
df = df.rename(columns={col: col.replace(' ', '_')})
df.query(f'`{col.replace(" ", "_")}` > 1')
else:
raise Prevention
- Sanitize column names on ingestion: df.columns = df.columns.str.replace(r'\W', '_', regex=True).
- Validate every column name with create_valid_python_identifier before writing query strings.
- Prefer boolean indexing df[df[col] > 1] for column names with exotic Unicode.
- Keep pandas current — escape-table coverage improves with releases (GH 49633).
When it happens
Trigger: df.query('`<weird col>` > 1') where the column name, after all substitutions, still contains a character that breaks identifier rules — e.g. certain Unicode combining marks, control characters, or characters not covered by tokenize.EXACT_TOKEN_TYPES or the manual override dict.
Common situations: Column names generated from external data (databases, CSVs) containing exotic Unicode (combining diacritics, zero-width joiners, emoji with variation selectors). Column names with leading digits after the BACKTICK_QUOTED_STRING_ prefix where a substitution leaves an invalid start character. pandas version upgrades that changed the escape table (GH 49633).
Related errors
- cannot evaluate scalar only bool ops
- Function " " does not support keyword arguments
- Invalid function call
- keyword error in function call
- N-dimensional objects, where N > 2, are not supported with…
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/db1e87b46b052672.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/computation/parsing.py:77
{
" ": "_",
"?": "_QUESTIONMARK_",
"!": "_EXCLAMATIONMARK_",
"$": "_DOLLARSIGN_",
"€": "_EUROSIGN_",
"°": "_DEGREESIGN_",
"'": "_SINGLEQUOTE_",
'"': "_DOUBLEQUOTE_",
"#": "_HASH_",
"`": "_BACKTICK_",
}
)
name = "".join([special_characters_replacements.get(char, char) for char in name])
name = f"BACKTICK_QUOTED_STRING_{name}"
if not name.isidentifier():
raise SyntaxError(f"Could not convert '{name}' to a valid Python identifier.")
return name
def clean_backtick_quoted_toks(tok: tuple[int, str]) -> tuple[int, str]:
"""
Clean up a column name if surrounded by backticks.
Backtick quoted string are indicated by a certain tokval value. If a string
is a backtick quoted token it will processed by
:func:`_create_valid_python_identifier` so that the parser can find this
string when the query is executed.
In this case the tok will get the NAME tokval.
Parameters
----------
tok : tuple of int, str
ints correspond to the all caps constants in the tokenize moduleView on GitHub (pinned to 3b7651241d)