pandas-dev/pandas · error · TypeError

Cannot perform reduction

Error message

Cannot perform reduction '{name}' with string dtype

What it means

StringArray._reduce raises TypeError for any reduction name not in the supported set (any, all, count, min, max, argmin, argmax, sum). String data has no meaningful numeric mean, median, std, prod, or sem, so pandas refuses rather than coerce strings to numbers and silently return NaN.

Solutions

  1. Filter numeric columns before reducing: df.select_dtypes('number').mean().
  2. Convert the column to numeric first: pd.to_numeric(df['col'], errors='coerce').mean().
  3. If strings encode numbers, strip/cast explicitly then reduce.
  4. Limit reductions to the supported set (count, min, max, sum, any, all) on string columns.

Example fix

// before
s = pd.Series(['1','2','3'], dtype='string')
s.mean()  # TypeError
// after
pd.to_numeric(s, errors='coerce').mean()  # 2.0
Defensive patterns

Strategy: type-guard

Validate before calling

def safe_reduce(s, name):
    if pd.api.types.is_string_dtype(s):
        if name not in {'count','min','max','sum','argmin','argmax','any','all'}:
            raise TypeError(f"Reduction '{name}' not defined for string dtype")
    return getattr(s, name)()

Type guard

def is_numeric_reduce_safe(series) -> bool:
    return pd.api.types.is_numeric_dtype(series)

Try / catch

try:
    df.mean()
except TypeError as e:
    if 'string dtype' in str(e):
        df.select_dtypes('number').mean()
    else:
        raise

Prevention

When it happens

Trigger: Calling df.mean(), df.median(), df.std(), df.var(), df.prod(), df.sem(), df.skew(), or df.kurt() on a column whose dtype is 'string' or 'string[pyarrow]'; also df._reduce('median') directly on a StringArray.

Common situations: A pipeline computes describe() or a fixed list of aggregations over every column without filtering dtypes; a CSV inferred as strings where numbers were expected; mixing categorical labels and numeric columns and calling mean() across the frame.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/81e47015614b963f. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_.py:978

        skipna: bool = True,
        keepdims: bool = False,
        axis: AxisInt | None = 0,
        **kwargs,
    ):
        if self.dtype.na_value is np.nan and name in ["any", "all"]:
            if name == "any":
                return nanops.nanany(self._ndarray, skipna=skipna)
            else:
                return nanops.nanall(self._ndarray, skipna=skipna)
        elif name == "count":
            return super().count()
        elif name in ["min", "max", "argmin", "argmax", "sum"]:
            result = getattr(self, name)(skipna=skipna, axis=axis, **kwargs)
            if keepdims:
                return self._from_sequence([result], dtype=self.dtype)
            return result

        raise TypeError(f"Cannot perform reduction '{name}' with string dtype")

    def _accumulate(self, name: str, *, skipna: bool = True, **kwargs) -> StringArray:
        """
        Return an ExtensionArray performing an accumulation operation.

        The underlying data type might change.

        Parameters
        ----------
        name : str
            Name of the function, supported values are:
            - cummin
            - cummax
            - cumsum
            - cumprod
        skipna : bool, default True
            If True, skip NA values.
        **kwargs

View on GitHub (pinned to 3b7651241d)