pandas-dev/pandas · error · TypeError
Cannot perform reduction
Error message
Cannot perform reduction '{name}' with string dtype What it means
StringArray._reduce raises TypeError for any reduction name not in the supported set (any, all, count, min, max, argmin, argmax, sum). String data has no meaningful numeric mean, median, std, prod, or sem, so pandas refuses rather than coerce strings to numbers and silently return NaN.
Solutions
- Filter numeric columns before reducing: df.select_dtypes('number').mean().
- Convert the column to numeric first: pd.to_numeric(df['col'], errors='coerce').mean().
- If strings encode numbers, strip/cast explicitly then reduce.
- Limit reductions to the supported set (count, min, max, sum, any, all) on string columns.
Example fix
// before s = pd.Series(['1','2','3'], dtype='string') s.mean() # TypeError // after pd.to_numeric(s, errors='coerce').mean() # 2.0
Defensive patterns
Strategy: type-guard
Validate before calling
def safe_reduce(s, name):
if pd.api.types.is_string_dtype(s):
if name not in {'count','min','max','sum','argmin','argmax','any','all'}:
raise TypeError(f"Reduction '{name}' not defined for string dtype")
return getattr(s, name)() Type guard
def is_numeric_reduce_safe(series) -> bool:
return pd.api.types.is_numeric_dtype(series) Try / catch
try:
df.mean()
except TypeError as e:
if 'string dtype' in str(e):
df.select_dtypes('number').mean()
else:
raise Prevention
- Filter numeric columns before applying numeric reductions.
- Validate dtypes of columns targeted by mean/median/std.
- Convert string-encoded numbers with pd.to_numeric before reducing.
When it happens
Trigger: Calling df.mean(), df.median(), df.std(), df.var(), df.prod(), df.sem(), df.skew(), or df.kurt() on a column whose dtype is 'string' or 'string[pyarrow]'; also df._reduce('median') directly on a StringArray.
Common situations: A pipeline computes describe() or a fixed list of aggregations over every column without filtering dtypes; a CSV inferred as strings where numbers were expected; mixing categorical labels and numeric columns and calling mean() across the frame.
Related errors
- Cannot perform reduction
- operation ' ' not supported for dtype
- Period type does not support
- setting an array element with a sequence.
- timedelta64 type does not support
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/81e47015614b963f.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/string_.py:978
skipna: bool = True,
keepdims: bool = False,
axis: AxisInt | None = 0,
**kwargs,
):
if self.dtype.na_value is np.nan and name in ["any", "all"]:
if name == "any":
return nanops.nanany(self._ndarray, skipna=skipna)
else:
return nanops.nanall(self._ndarray, skipna=skipna)
elif name == "count":
return super().count()
elif name in ["min", "max", "argmin", "argmax", "sum"]:
result = getattr(self, name)(skipna=skipna, axis=axis, **kwargs)
if keepdims:
return self._from_sequence([result], dtype=self.dtype)
return result
raise TypeError(f"Cannot perform reduction '{name}' with string dtype")
def _accumulate(self, name: str, *, skipna: bool = True, **kwargs) -> StringArray:
"""
Return an ExtensionArray performing an accumulation operation.
The underlying data type might change.
Parameters
----------
name : str
Name of the function, supported values are:
- cummin
- cummax
- cumsum
- cumprod
skipna : bool, default True
If True, skip NA values.
**kwargsView on GitHub (pinned to 3b7651241d)