pandas-dev/pandas · error · ValueError

The numba engine only supports using string or numeric…

Error message

The numba engine only supports using string or numeric column names

What it means

Raised by `set_numba_data` (in pandas.core._numba.extensions) when an Index used with the numba engine holds object/string-dtype data that, when cast to a numpy string array, contains non-string elements. The numba engine requires column names be either numeric or pure strings; mixed object columns cannot be lowered into numba-compatible typed memory.

Solutions

  1. Normalize the Index to a single dtype (all strings or all ints) before invoking the numba engine: `df.columns = df.columns.astype(str)`.
  2. Drop or fill missing/None column labels before the call.
  3. Fall back to the default python engine if mixed-type columns are unavoidable.

Example fix

// before
df.groupby(col).mean(engine='numba')  # columns are mixed type

// after
df.columns = df.columns.astype(str)
df.groupby(col).mean(engine='numba')
Defensive patterns

Strategy: validation

Validate before calling

from pandas import lib
cols = df.columns
if cols.dtype in (object, 'string'):
    arr = cols.to_numpy()
    assert lib.is_string_array(arr.astype(object)), 'columns contain non-string entries'

Try / catch

try:
    df.groupby(col).mean(engine='numba')
except ValueError:
    df.groupby(col).mean(engine='cython')

Prevention

When it happens

Trigger: Running a groupby/transform/apply with `engine='numba'` where the column Index contains mixed types (e.g. ints and strings), NaNs in a string index, or unhashable objects stored as object dtype.

Common situations: Frames built from heterogeneous sources where the Index ended up as object dtype; nullable string columns with NaN cast to object; switching a workload to the numba engine on a frame not designed for it.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/b0a9ea1865a7a216. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/_numba/extensions.py:56

from pandas.core.indexes.base import Index
from pandas.core.indexing import _iLocIndexer
from pandas.core.internals import SingleBlockManager
from pandas.core.series import Series


# Helper function to hack around fact that Index casts numpy string dtype to object
#
# Idea is to set an attribute on an Index called _numba_data
# that is the original data, or the object data casted to numpy string dtype,
# with a context manager that is unset afterwards
@contextmanager
def set_numba_data(index: Index):
    numba_data = index._data
    if numba_data.dtype in (object, "string"):
        numba_data = np.asarray(numba_data)
        if not lib.is_string_array(numba_data):
            raise ValueError(
                "The numba engine only supports using string or numeric column names"
            )
        numba_data = numba_data.astype("U")
    try:
        index._numba_data = numba_data
        yield index
    finally:
        del index._numba_data


# TODO: Range index support
# (this currently lowers OK, but does not round-trip)
class IndexType(types.Type):
    """
    The type class for Index objects.
    """

    def __init__(self, dtype, layout, pyclass: any) -> None:

View on GitHub (pinned to 3b7651241d)