pandas-dev/pandas · error · ValueError

removals must all be in old categories

Error message

removals must all be in old categories: {not_included}

What it means

Raised by `remove_categories` when `removals` contains values that are not present in the current categories. Removing a non-existent label is treated as a programming error (likely a stale list) and is refused, with the offending entries reported via `not_included`.

Solutions

  1. Intersect the removal list with current categories: `to_remove = [r for r in removals if r in set(cat.categories)]; cat.remove_categories(to_remove)`.
  2. Use `set_categories` with the desired final set to sidestep the existence check.
  3. Verify each removal is in `cat.categories` before the call.
  4. Make cleanup scripts idempotent by filtering removals through current categories.

Example fix

# before
cat = pd.Categorical(['a', 'b'], categories=['a', 'b', 'c'])
cat = cat.remove_categories(['c', 'x'])  # ValueError: not_included={'x'}

# after
removals = ['c', 'x']
to_remove = [r for r in removals if r in cat.categories]
cat = cat.remove_categories(to_remove)
Defensive patterns

Strategy: validation

Validate before calling

def safe_remove_categories(cat, removals):
    if not is_list_like(removals):
        removals = [removals]
    existing = set(cat.categories)
    to_remove = [r for r in removals if r in existing]
    return cat.remove_categories(to_remove) if to_remove else cat

Type guard

def are_all_present(cat, removals) -> bool:
    existing = set(cat.categories)
    return all(r in existing for r in removals)

Try / catch

try:
    cat = cat.remove_categories(removals)
except ValueError as e:
    if 'must all be in old categories' in str(e):
        cat = cat.remove_categories([r for r in removals if r in set(cat.categories)])
    else:
        raise

Prevention

When it happens

Trigger: `cat.remove_categories(['x'])` where `'x'` is not in `cat.categories`; passing a list built from a different column or schema.

Common situations: Schema drift between runs; idempotent re-runs of a cleanup script; merging removal lists across differently-categorized DataFrames; refactoring that renamed labels upstream.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/8376b5e2168cb5f4. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/categorical.py:1492

        [NaN, 'c', 'b', 'c', NaN]
        Categories (2, str): ['b', 'c']
        """
        from pandas import Index

        if not is_list_like(removals):
            removals = [removals]

        removals = Index(removals).unique().dropna()
        new_categories = (
            self.dtype.categories.difference(removals, sort=False)
            if self.dtype.ordered is True
            else self.dtype.categories.difference(removals)
        )
        not_included = removals.difference(self.dtype.categories)

        if len(not_included) != 0:
            not_included = set(not_included)
            raise ValueError(f"removals must all be in old categories: {not_included}")

        return self.set_categories(new_categories, ordered=self.ordered, rename=False)

    def remove_unused_categories(self) -> Self:
        """
        Remove categories which are not used.

        This method is useful when working with datasets
        that undergo dynamic changes where categories may no longer be
        relevant, allowing to maintain a clean, efficient data structure.

        Returns
        -------
        Categorical
            Categorical with unused categories dropped.

        See Also
        --------

View on GitHub (pinned to 3b7651241d)