pandas-dev/pandas · error · TypeError
Cannot perform reduction
Error message
Cannot perform reduction '{name}' with string dtype What it means
ArrowStringArray._reduce raises TypeError for any reduction not in (count, min, max, sum, argmin, argmax, any, all). PyArrow-backed string arrays support only those operations; numeric reductions like mean/median/std are undefined for strings.
Solutions
- Filter numeric columns before reducing: df.select_dtypes('number').mean().
- Cast the column to numeric with pd.to_numeric(..., errors='coerce') before reducing.
- Restrict reductions to count/min/max/sum/any/all on string columns.
Example fix
// before s = pd.array(['1','2','3'], dtype='string[pyarrow]') pd.Series(s).mean() # TypeError // after pd.to_numeric(pd.Series(s), errors='coerce').mean()
Defensive patterns
Strategy: type-guard
Validate before calling
def safe_reduce_arrow(s, name):
if pd.api.types.is_string_dtype(s) and name not in {'count','min','max','sum','argmin','argmax','any','all'}:
raise TypeError(f"Reduction '{name}' not supported for string[pyarrow]")
return getattr(s, name)() Type guard
def arrow_string_reduce_safe(series) -> bool:
return pd.api.types.is_numeric_dtype(series) Try / catch
try:
df.mean()
except TypeError as e:
if 'string dtype' in str(e):
df.select_dtypes('number').mean()
else:
raise Prevention
- Filter numeric columns before numeric reductions.
- Cast string-encoded numbers with pd.to_numeric first.
- Restrict reductions to the supported set on string[pyarrow] columns.
When it happens
Trigger: df['str_col'].mean() where dtype is string[pyarrow]; df.median() including a pyarrow string column; calling _reduce('prod') directly.
Common situations: A pipeline computes a fixed list of aggregations on every column; converting a column to string[pyarrow] without updating downstream numeric reductions.
Related errors
- Cannot perform reduction
- ArrowStringArray requires a PyArrow (chunked) array of…
- bad operand type for unary +
- Invalid value for dtype 'str'. Value should be a string or…
- Invalid value ' ' for dtype 'str'. Value should be a string…
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/4233d6b1e6b45dd7.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/string_arrow.py:640
nv.validate_minmax_axis(axis, self.ndim)
if self.dtype.na_value is np.nan and name in ["any", "all"]:
if not skipna:
nas = pc.is_null(self._pa_array)
arr = pc.or_kleene(nas, pc.not_equal(self._pa_array, ""))
else:
arr = pc.not_equal(self._pa_array, "")
result = ArrowExtensionArray(arr)._reduce(
name, skipna=skipna, keepdims=keepdims, **kwargs
)
if keepdims:
# ArrowExtensionArray will return a length-1 bool[pyarrow] array
return result.astype(np.bool_)
return result
if name in ("count", "min", "max", "sum", "argmin", "argmax"):
result = self._reduce_calc(name, skipna=skipna, keepdims=keepdims, **kwargs)
else:
raise TypeError(f"Cannot perform reduction '{name}' with string dtype")
if name in ("argmin", "argmax") and isinstance(result, pa.Array):
return self._convert_int_result(result)
elif isinstance(result, pa.Array):
return type(self)(result, dtype=self.dtype)
else:
return result
def value_counts(self, dropna: bool = True) -> Series:
result = super().value_counts(dropna=dropna)
if self.dtype.na_value is np.nan:
res_values = result._values.to_numpy()
return result._constructor(
res_values, index=result.index, name=result.name, copy=False
)
return result
def _cmp_method(self, other, op):View on GitHub (pinned to 3b7651241d)