pandas-dev/pandas · error · NotImplementedError

groupby first/last only supports 1D ExtensionArrays

Error message

groupby first/last only supports 1D ExtensionArrays

What it means

Raised by ExtensionArray._groupby_first_last when self.ndim != 1. The optimized index-based first/last implementation uses libgroupby.group_first_last_indexer which only handles 1D arrays, so 2D (or higher) ExtensionArrays are rejected rather than silently producing wrong-shaped output.

Solutions

  1. Flatten the ExtensionArray to 1D before groupby first/last, or operate per-column.
  2. If you control the EA subclass, ensure ndim always returns 1 for the array passed to groupby.
  3. Fall back to a non-EA representation (e.g. numpy object array) for the groupby operation.
  4. Override _groupby_first_last in your concrete EA subclass to handle the 2D case yourself.

Example fix

// before
grouped = df.groupby('key')[tensor_ea_2d_column].first()
// after
for col in tensor_ea_2d_column.flat_columns:
    df.groupby('key')[col].first()
Defensive patterns

Strategy: validation

Validate before calling

assert arr.ndim == 1, f'expected 1D EA, got ndim={arr.ndim}'

Type guard

def is_1d_ea(arr) -> bool:
    return isinstance(arr, ExtensionArray) and arr.ndim == 1

Try / catch

try:
    grouped.first()
except NotImplementedError as e:
    if '1D' in str(e):
        ...

Prevention

When it happens

Trigger: Calling groupby(...).first() or groupby(...).last() on a 2D ExtensionArray subclass (e.g. a custom EA with ndarray-shaped elements where ndim > 1). Reached when an ExtensionArray overrides __new__ to store 2D data, or when a third-party EA reports ndim=2.

Common situations: Custom third-party ExtensionArray types that wrap 2D data; experimental EAs for image/tensor data; subclassing ExtensionArray without keeping data 1D; version upgrades where this guard was added to fix silent shape bugs.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/e71dff4725e9d291. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/base.py:3208

    def _groupby_first_last(
        self,
        *,
        how: str,
        min_count: int,
        ngroups: int,
        ids: npt.NDArray[np.intp],
        skipna: bool = True,
    ) -> Self:
        """
        Optimized implementation of groupby first/last for ExtensionArrays.

        Uses an index-based approach: computes the index of the first/last
        non-NA element per group, then gathers results via take(). This avoids
        any dtype conversion and works for all EA types.
        """
        if self.ndim != 1:
            raise NotImplementedError(
                "groupby first/last only supports 1D ExtensionArrays"
            )
        isna_mask = np.asarray(self.isna(), dtype=np.uint8)

        result_indices, result_mask = libgroupby.group_first_last_indexer(
            labels=ids,
            mask=isna_mask,
            ngroups=ngroups,
            skipna=skipna,
            is_last=(how == "last"),
        )

        # Apply min_count: require at least min_count non-NA observations.
        # For first/last, the natural minimum is 1 (need at least one value).
        if min_count > 1:
            nobs = np.zeros(ngroups, dtype=np.int64)
            non_na_indices = np.where((~isna_mask.view(bool)) & (ids >= 0))[0]
            np.add.at(nobs, ids[non_na_indices], 1)

View on GitHub (pinned to 3b7651241d)