{"record":{"id":"b0a9ea1865a7a216","repo":"pandas-dev/pandas","slug":"the-numba-engine-only-supports-using-string-or-num","errorCode":null,"errorMessage":"The numba engine only supports using string or numeric column names","messagePattern":"The numba engine only supports using string or numeric column names","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"pandas/core/_numba/extensions.py","lineNumber":56,"sourceCode":"\nfrom pandas.core.indexes.base import Index\nfrom pandas.core.indexing import _iLocIndexer\nfrom pandas.core.internals import SingleBlockManager\nfrom pandas.core.series import Series\n\n\n# Helper function to hack around fact that Index casts numpy string dtype to object\n#\n# Idea is to set an attribute on an Index called _numba_data\n# that is the original data, or the object data casted to numpy string dtype,\n# with a context manager that is unset afterwards\n@contextmanager\ndef set_numba_data(index: Index):\n    numba_data = index._data\n    if numba_data.dtype in (object, \"string\"):\n        numba_data = np.asarray(numba_data)\n        if not lib.is_string_array(numba_data):\n            raise ValueError(\n                \"The numba engine only supports using string or numeric column names\"\n            )\n        numba_data = numba_data.astype(\"U\")\n    try:\n        index._numba_data = numba_data\n        yield index\n    finally:\n        del index._numba_data\n\n\n# TODO: Range index support\n# (this currently lowers OK, but does not round-trip)\nclass IndexType(types.Type):\n    \"\"\"\n    The type class for Index objects.\n    \"\"\"\n\n    def __init__(self, dtype, layout, pyclass: any) -> None:","sourceCodeStart":38,"sourceCodeEnd":74,"githubUrl":"https://github.com/pandas-dev/pandas/blob/3b7651241d4da534b3559b60ef128e1c34f54116/pandas/core/_numba/extensions.py#L38-L74","documentation":"Raised by `set_numba_data` (in pandas.core._numba.extensions) when an Index used with the numba engine holds object/string-dtype data that, when cast to a numpy string array, contains non-string elements. The numba engine requires column names be either numeric or pure strings; mixed object columns cannot be lowered into numba-compatible typed memory.","triggerScenarios":"Running a groupby/transform/apply with `engine='numba'` where the column Index contains mixed types (e.g. ints and strings), NaNs in a string index, or unhashable objects stored as object dtype.","commonSituations":"Frames built from heterogeneous sources where the Index ended up as object dtype; nullable string columns with NaN cast to object; switching a workload to the numba engine on a frame not designed for it.","solutions":["Normalize the Index to a single dtype (all strings or all ints) before invoking the numba engine: `df.columns = df.columns.astype(str)`.","Drop or fill missing/None column labels before the call.","Fall back to the default python engine if mixed-type columns are unavoidable."],"exampleFix":"// before\ndf.groupby(col).mean(engine='numba')  # columns are mixed type\n\n// after\ndf.columns = df.columns.astype(str)\ndf.groupby(col).mean(engine='numba')","handlingStrategy":"validation","validationCode":"from pandas import lib\ncols = df.columns\nif cols.dtype in (object, 'string'):\n    arr = cols.to_numpy()\n    assert lib.is_string_array(arr.astype(object)), 'columns contain non-string entries'","typeGuard":null,"tryCatchPattern":"try:\n    df.groupby(col).mean(engine='numba')\nexcept ValueError:\n    df.groupby(col).mean(engine='cython')","preventionTips":["Cast df.columns to a uniform dtype (str or int) before using engine='numba'.","Fall back to the default engine for frames with heterogeneous column types."],"tags":["numba","dtype","groupby","string-columns"],"backgroundTag":null,"analyzedSha":"3b7651241d4da534b3559b60ef128e1c34f54116","analyzedAt":"2026-08-11T22:10:44.015Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}