pandas-dev/pandas · error · ValueError

Column must have a numeric dtype. Found ' ' instead

Error message

Column {colname} must have a numeric dtype. Found '{dtype}' instead

What it means

Raised by `validate_values_for_numba` when a column of the DataFrame being passed to the numba engine has a non-numeric dtype (e.g. object, string, datetime, category). Numba JIT-compiles per-column numeric kernels and cannot handle arbitrary Python objects, so pandas validates every column dtype before invoking numba and reports the offending column and dtype.

Solutions

  1. Select only numeric columns before applying: `df.select_dtypes('number').apply(func, engine='numba')`.
  2. Coerce dtypes upstream: `df[col] = pd.to_numeric(df[col], errors='coerce')`.
  3. Drop or separate datetime/string columns and process them with the python engine.

Example fix

// before
df.apply(func, engine='numba')  # df has an object column
// after
df.select_dtypes('number').apply(func, engine='numba')
Defensive patterns

Strategy: validation

Validate before calling

def numeric_only_numba_apply(df, func, **kw):
    non_numeric = [c for c, d in df.dtypes.items() if not pd.api.types.is_numeric_dtype(d)]
    if non_numeric:
        raise ValueError(f'Non-numeric columns block numba: {non_numeric}')
    return df.apply(func, engine='numba', **kw)

Type guard

def frame_is_numeric(df) -> bool:
    return all(pd.api.types.is_numeric_dtype(d) for d in df.dtypes)

Try / catch

try:
    out = df.apply(func, engine='numba')
except ValueError as e:
    if 'numeric dtype' in str(e):
        out = df.select_dtypes('number').apply(func, engine='numba')
    else:
        raise

Prevention

When it happens

Trigger: `df.apply(func, engine='numba')` where `df` contains any non-numeric column (object, str, datetime64, category, bool-on some versions). The loop iterates `self.obj.dtypes.items()`.

Common situations: Mixed-type frames where an index or stray string column prevents numba compilation; CSVs that import numeric-looking columns as object due to NaNs/strings; datetime indexes that get included as columns after a reset_index.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/efcc50ca0bd23808. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/apply.py:979

        pass

    @staticmethod
    @functools.cache
    @abc.abstractmethod
    def generate_numba_apply_func(
        func, nogil: bool = True, parallel: bool = False
    ) -> Callable[[npt.NDArray, Index, Index], dict[int, Any]]:
        pass

    @abc.abstractmethod
    def apply_with_numba(self):
        pass

    def validate_values_for_numba(self) -> None:
        # Validate column dtypes all OK
        for colname, dtype in self.obj.dtypes.items():
            if not is_numeric_dtype(dtype):
                raise ValueError(
                    f"Column {colname} must have a numeric dtype. "
                    f"Found '{dtype}' instead"
                )
            if is_extension_array_dtype(dtype):
                raise ValueError(
                    f"Column {colname} is backed by an extension array, "
                    f"which is not supported by the numba engine."
                )

    @abc.abstractmethod
    def wrap_results_for_axis(
        self, results: ResType, res_index: Index
    ) -> DataFrame | Series:
        pass

    # ---------------------------------------------------------------

    @property

View on GitHub (pinned to 3b7651241d)