pandas-dev/pandas · error · NotImplementedError

The index/columns must be unique when raw=False and…

Error message

The index/columns must be unique when raw=False and engine='numba'

What it means

Raised by apply_series_numba when engine='numba' with raw=False and the underlying DataFrame/Series has non-unique index labels or duplicate column names. The numba path builds its result by positional assignment and then realigns by labels, which is impossible when labels repeat, so pandas refuses instead of producing silently misaligned output.

Solutions

  1. Deduplicate the index before applying: df = df.reset_index(drop=True) (or df[~df.index.duplicated(keep='first')]).
  2. Remove duplicate columns: df = df.loc[:, ~df.columns.duplicated()].
  3. Fall back to the default engine: df.apply(func) (no engine='numba').

Example fix

# before
df.apply(func, engine='numba')  # df.index has duplicates
# after
df.reset_index(drop=True).apply(func, engine='numba')
Defensive patterns

Strategy: validation

Validate before calling

def ensure_unique_for_numba(df):
    if df.index.has_duplicates:
        df = df.reset_index(drop=True)
    if getattr(df.columns, 'has_duplicates', False):
        df = df.loc[:, ~df.columns.duplicated()]
    return df

df = ensure_unique_for_numba(df)
df.apply(func, engine='numba')

Try / catch

try:
    df.apply(func, engine='numba')
except NotImplementedError as e:
    if 'must be unique' in str(e):
        df.reset_index(drop=True).apply(func, engine='numba')
    else:
        raise

Prevention

When it happens

Trigger: df.apply(func, engine='numba') where df.index.has_duplicates or df.columns.has_duplicates; same on a Series whose index has duplicates.

Common situations: Applying numba engine on pivoted/aggregated frames that retained duplicate index entries, or on a Series produced by .value_counts() / groupby sums that you forgot to reset.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/c5d974aad93ec3c0. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/apply.py:1311

        results = {}

        for i, v in enumerate(series_gen):
            results[i] = self.func(v, *self.args, **self.kwargs)
            if isinstance(results[i], ABCSeries):
                # If we have a view on v, we need to make a copy because
                #  series_generator will swap out the underlying data
                results[i] = results[i].copy(deep=False)

        return results, res_index

    def apply_series_numba(self):
        if self.engine_kwargs.get("parallel", False):
            raise NotImplementedError(
                "Parallel apply is not supported when raw=False and engine='numba'"
            )
        if not self.obj.index.is_unique or not self.columns.is_unique:
            raise NotImplementedError(
                "The index/columns must be unique when raw=False and engine='numba'"
            )
        self.validate_values_for_numba()
        results = self.apply_with_numba()
        return results, self.result_index

    def wrap_results(self, results: ResType, res_index: Index) -> DataFrame | Series:
        from pandas import Series

        # see if we can infer the results
        if len(results) > 0 and 0 in results and is_sequence(results[0]):
            return self.wrap_results_for_axis(results, res_index)

        # dict of scalars

        # the default dtype of an empty Series is `object`, but this
        # code can be hit by df.mean() where the result should have dtype
        # float64 even if it's an empty Series.

View on GitHub (pinned to 3b7651241d)