pandas-dev/pandas · error · TypeError

dtype ' ' does not support operation

Error message

dtype '{self.dtype}' does not support operation '{how}'

What it means

_groupby_op explicitly rejects a list of mathematically-undefined reductions (prod, mean, median, cumsum, cumprod, std, sem, var, skew) for StringDtype-backed-by-arrow arrays, raising TypeError. Even though pyarrow has some kernels, these operations are semantically meaningless on strings, so pandas short-circuits before dispatch.

Solutions

  1. Filter columns by numeric dtype before grouping: df.select_dtypes('number').groupby(key).mean().
  2. Use a string-appropriate aggregation: .count, .first, .last, .sum (concatenation is allowed), .size.
  3. Build a per-column aggregation dictionary excluding string columns.

Example fix

// before
df = pd.DataFrame({"k": ["a", "a"], "v": ["x", "y"]}, dtype={"v": "string[pyarrow]"})
df.groupby("k")["v"].mean()
// after
df = pd.DataFrame({"k": ["a", "a"], "v": [1, 2]}, dtype={"v": "int64[pyarrow]"})
df.groupby("k")["v"].mean()
Defensive patterns

Strategy: validation

Validate before calling

STRING_UNSUPPORTED = {"prod","mean","median","cumsum","cumprod","std","sem","var","skew"}

def groupby_op_supported(dtype, how) -> bool:
    from pandas.core.arrays.string_ import StringDtype
    if isinstance(dtype, StringDtype) and how in STRING_UNSUPPORTED:
        return False
    return True

Type guard

def is_string_arrow_dtype(dtype) -> bool:
    from pandas.core.arrays.string_ import StringDtype
    return isinstance(dtype, StringDtype)

Try / catch

try:
    df.groupby("k")[col].mean()
except TypeError:
    # skip non-numeric columns
    df.select_dtypes("number").groupby("k").mean()

Prevention

When it happens

Trigger: Grouping a string[pyarrow] column and calling .prod()/.mean()/.median()/.cumsum()/.cumprod()/.std()/.sem()/.var()/.skew() on the groupby result.

Common situations: df.groupby('key').agg(['mean','std']) on a frame with string columns; applying numeric aggregations to every column without dtype filtering.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/fdd768c5c103304c. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/arrow/array.py:3539

        has_dropped_na: bool,
        min_count: int,
        ngroups: int,
        ids: npt.NDArray[np.intp],
        **kwargs,
    ):
        if isinstance(self.dtype, StringDtype):
            if how in [
                "prod",
                "mean",
                "median",
                "cumsum",
                "cumprod",
                "std",
                "sem",
                "var",
                "skew",
            ]:
                raise TypeError(
                    f"dtype '{self.dtype}' does not support operation '{how}'"
                )
            # Fall through to Arrow-native path below

        pa_type = self._pa_array.type

        # Try PyArrow-native path for decimal and string types where it's faster.
        # For integer/float/boolean, the fallback path via _to_masked() is faster.
        if (
            pa.types.is_decimal(pa_type)
            or pa.types.is_string(pa_type)
            or pa.types.is_large_string(pa_type)
        ):
            native_result = self._groupby_op_pyarrow(
                how=how,
                min_count=min_count,
                ngroups=ngroups,
                ids=ids,

View on GitHub (pinned to 3b7651241d)