pandas-dev/pandas · error · SyntaxError

Could not convert ' ' to a valid Python identifier.

Error message

Could not convert '{name}' to a valid Python identifier.

What it means

Raised in create_valid_python_identifier after every replacement pass still yields a string that fails str.isidentifier(). The function escapes non-ASCII (via backslashreplace + _UNICODE_/), maps special characters to token-name placeholders (e.g. '(' -> _LRB_), prefixes the result with 'BACKTICK_QUOTED_STRING_', and as a last resort raises SyntaxError. This powers the backtick-quoted column-name feature of query(); a failure means the column name contains a character the escape table cannot normalize into a valid identifier.

Solutions

  1. Sanitize column names before querying: df.columns = df.columns.str.replace(r'[^\w]', '_', regex=True).
  2. Avoid backtick-quoted columns in query; use boolean indexing instead: df[df['<weird col>'] > 1].
  3. Identify and remove the offending character: print([c for c in col if not ('_'+c).isidentifier()]).
  4. Upgrade pandas — recent versions expanded the escape table (GH 49633) to cover more characters.

Example fix

// before
df.query("`col\u200bname` > 1")  # zero-width space in name

// after
df.columns = df.columns.str.replace(r"\W", "_", regex=True)
df.query("`col_name` > 1")
# or skip query
df[df["col\u200bname"] > 1]
Defensive patterns

Strategy: validation

Validate before calling

from pandas.core.computation.parsing import create_valid_python_identifier

def column_is_query_safe(col: str) -> bool:
    try:
        create_valid_python_identifier(col)
        return True
    except SyntaxError:
        return False

bad = [c for c in df.columns if not column_is_query_safe(c)]
assert not bad, f'columns not query-safe: {bad}'

Type guard

from pandas.core.computation.parsing import create_valid_python_identifier

def is_query_safe_column(name: str) -> bool:
    try:
        create_valid_python_identifier(name)
        return True
    except SyntaxError:
        return False

Try / catch

try:
    df.query(f'`{col}` > 1')
except SyntaxError as e:
    if 'Could not convert' in str(e):
        df = df.rename(columns={col: col.replace(' ', '_')})
        df.query(f'`{col.replace(" ", "_")}` > 1')
    else:
        raise

Prevention

When it happens

Trigger: df.query('`<weird col>` > 1') where the column name, after all substitutions, still contains a character that breaks identifier rules — e.g. certain Unicode combining marks, control characters, or characters not covered by tokenize.EXACT_TOKEN_TYPES or the manual override dict.

Common situations: Column names generated from external data (databases, CSVs) containing exotic Unicode (combining diacritics, zero-width joiners, emoji with variation selectors). Column names with leading digits after the BACKTICK_QUOTED_STRING_ prefix where a substitution leaves an invalid start character. pandas version upgrades that changed the escape table (GH 49633).

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/db1e87b46b052672. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/computation/parsing.py:77

        {
            " ": "_",
            "?": "_QUESTIONMARK_",
            "!": "_EXCLAMATIONMARK_",
            "$": "_DOLLARSIGN_",
            "€": "_EUROSIGN_",
            "°": "_DEGREESIGN_",
            "'": "_SINGLEQUOTE_",
            '"': "_DOUBLEQUOTE_",
            "#": "_HASH_",
            "`": "_BACKTICK_",
        }
    )

    name = "".join([special_characters_replacements.get(char, char) for char in name])
    name = f"BACKTICK_QUOTED_STRING_{name}"

    if not name.isidentifier():
        raise SyntaxError(f"Could not convert '{name}' to a valid Python identifier.")

    return name


def clean_backtick_quoted_toks(tok: tuple[int, str]) -> tuple[int, str]:
    """
    Clean up a column name if surrounded by backticks.

    Backtick quoted string are indicated by a certain tokval value. If a string
    is a backtick quoted token it will processed by
    :func:`_create_valid_python_identifier` so that the parser can find this
    string when the query is executed.
    In this case the tok will get the NAME tokval.

    Parameters
    ----------
    tok : tuple of int, str
        ints correspond to the all caps constants in the tokenize module

View on GitHub (pinned to 3b7651241d)