pandas-dev/pandas · error · ValueError
Column is backed by an extension array, which is not…
Error message
Column {colname} is backed by an extension array, which is not supported by the numba engine. What it means
Raised by `validate_values_for_numba` for a column whose dtype is an extension array (e.g. `Int64`, `Float64`, `boolean`, nullable types). Even if the dtype is numeric, numba operates on plain numpy buffers and cannot consume pandas ExtensionArrays, so pandas rejects them up front and names the offending column.
Solutions
- Cast extension columns to numpy equivalents: `df[col] = df[col].astype('int64')` (handle NA first).
- Use `df.select_dtypes(exclude='extension')` or filter out EA columns before the numba call.
- Fall back to the python engine for frames that must keep nullable dtypes.
Example fix
// before
df.apply(func, engine='numba') # df has Int64 (nullable) column
// after
df.astype({'col': 'int64'}).apply(func, engine='numba') Defensive patterns
Strategy: validation
Validate before calling
def no_ea_numba_apply(df, func, **kw):
ea_cols = [c for c, d in df.dtypes.items() if pd.api.types.is_extension_array_dtype(d)]
if ea_cols:
raise ValueError(f'Extension-array columns block numba: {ea_cols}')
return df.apply(func, engine='numba', **kw) Type guard
def frame_has_no_extension_arrays(df) -> bool:
return not any(pd.api.types.is_extension_array_dtype(d) for d in df.dtypes) Try / catch
try:
out = df.apply(func, engine='numba')
except ValueError as e:
if 'extension array' in str(e):
cast = df.copy()
for c in cast.columns:
if pd.api.types.is_extension_array_dtype(cast[c]):
cast[c] = cast[c].astype(cast[c].dtype._subtype if hasattr(cast[c].dtype, '_subtype') else 'float64')
out = cast.apply(func, engine='numba')
else:
raise Prevention
- Avoid convert_dtypes() before numba paths.
- Cast nullable Int64/Float64 to numpy int64/float64 after handling NaNs.
- Document which pipelines require plain numpy dtypes.
When it happens
Trigger: `df.apply(func, engine='numba')` where any column is a nullable/extension dtype (`'Int64'`, `'Float64'`, `'boolean'`, etc.). The check `is_extension_array_dtype(dtype)` fires after the numeric check, so it only triggers for numeric extension dtypes.
Common situations: Frames produced by `convert_dtypes()` (which yields nullable Int64/Float64/boolean), or by reading data with `dtype='Int64'`; migrating to nullable dtypes without realizing numba does not support them.
Related errors
- Column must have a numeric dtype. Found ' ' instead
- by_row= not allowed
- cannot broadcast result
- cannot diff on axis=
- Default 'empty' implementation is invalid for dtype=
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/612fe6aa6af0d5c1.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/apply.py:984
def generate_numba_apply_func(
func, nogil: bool = True, parallel: bool = False
) -> Callable[[npt.NDArray, Index, Index], dict[int, Any]]:
pass
@abc.abstractmethod
def apply_with_numba(self):
pass
def validate_values_for_numba(self) -> None:
# Validate column dtypes all OK
for colname, dtype in self.obj.dtypes.items():
if not is_numeric_dtype(dtype):
raise ValueError(
f"Column {colname} must have a numeric dtype. "
f"Found '{dtype}' instead"
)
if is_extension_array_dtype(dtype):
raise ValueError(
f"Column {colname} is backed by an extension array, "
f"which is not supported by the numba engine."
)
@abc.abstractmethod
def wrap_results_for_axis(
self, results: ResType, res_index: Index
) -> DataFrame | Series:
pass
# ---------------------------------------------------------------
@property
def res_columns(self) -> Index:
return self.result_columns
@property
def columns(self) -> Index:View on GitHub (pinned to 3b7651241d)