pandas-dev/pandas · error · TypeError

Cannot perform reduction

Error message

Cannot perform reduction '{name}' with string dtype

What it means

ArrowStringArray._reduce raises TypeError for any reduction not in (count, min, max, sum, argmin, argmax, any, all). PyArrow-backed string arrays support only those operations; numeric reductions like mean/median/std are undefined for strings.

Solutions

  1. Filter numeric columns before reducing: df.select_dtypes('number').mean().
  2. Cast the column to numeric with pd.to_numeric(..., errors='coerce') before reducing.
  3. Restrict reductions to count/min/max/sum/any/all on string columns.

Example fix

// before
s = pd.array(['1','2','3'], dtype='string[pyarrow]')
pd.Series(s).mean()  # TypeError
// after
pd.to_numeric(pd.Series(s), errors='coerce').mean()
Defensive patterns

Strategy: type-guard

Validate before calling

def safe_reduce_arrow(s, name):
    if pd.api.types.is_string_dtype(s) and name not in {'count','min','max','sum','argmin','argmax','any','all'}:
        raise TypeError(f"Reduction '{name}' not supported for string[pyarrow]")
    return getattr(s, name)()

Type guard

def arrow_string_reduce_safe(series) -> bool:
    return pd.api.types.is_numeric_dtype(series)

Try / catch

try:
    df.mean()
except TypeError as e:
    if 'string dtype' in str(e):
        df.select_dtypes('number').mean()
    else:
        raise

Prevention

When it happens

Trigger: df['str_col'].mean() where dtype is string[pyarrow]; df.median() including a pyarrow string column; calling _reduce('prod') directly.

Common situations: A pipeline computes a fixed list of aggregations on every column; converting a column to string[pyarrow] without updating downstream numeric reductions.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/4233d6b1e6b45dd7. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_arrow.py:640

        nv.validate_minmax_axis(axis, self.ndim)
        if self.dtype.na_value is np.nan and name in ["any", "all"]:
            if not skipna:
                nas = pc.is_null(self._pa_array)
                arr = pc.or_kleene(nas, pc.not_equal(self._pa_array, ""))
            else:
                arr = pc.not_equal(self._pa_array, "")
            result = ArrowExtensionArray(arr)._reduce(
                name, skipna=skipna, keepdims=keepdims, **kwargs
            )
            if keepdims:
                # ArrowExtensionArray will return a length-1 bool[pyarrow] array
                return result.astype(np.bool_)
            return result

        if name in ("count", "min", "max", "sum", "argmin", "argmax"):
            result = self._reduce_calc(name, skipna=skipna, keepdims=keepdims, **kwargs)
        else:
            raise TypeError(f"Cannot perform reduction '{name}' with string dtype")

        if name in ("argmin", "argmax") and isinstance(result, pa.Array):
            return self._convert_int_result(result)
        elif isinstance(result, pa.Array):
            return type(self)(result, dtype=self.dtype)
        else:
            return result

    def value_counts(self, dropna: bool = True) -> Series:
        result = super().value_counts(dropna=dropna)
        if self.dtype.na_value is np.nan:
            res_values = result._values.to_numpy()
            return result._constructor(
                res_values, index=result.index, name=result.name, copy=False
            )
        return result

    def _cmp_method(self, other, op):

View on GitHub (pinned to 3b7651241d)