pandas-dev/pandas · error · ValueError
removals must all be in old categories
Error message
removals must all be in old categories: {not_included} What it means
Raised by `remove_categories` when `removals` contains values that are not present in the current categories. Removing a non-existent label is treated as a programming error (likely a stale list) and is refused, with the offending entries reported via `not_included`.
Solutions
- Intersect the removal list with current categories: `to_remove = [r for r in removals if r in set(cat.categories)]; cat.remove_categories(to_remove)`.
- Use `set_categories` with the desired final set to sidestep the existence check.
- Verify each removal is in `cat.categories` before the call.
- Make cleanup scripts idempotent by filtering removals through current categories.
Example fix
# before
cat = pd.Categorical(['a', 'b'], categories=['a', 'b', 'c'])
cat = cat.remove_categories(['c', 'x']) # ValueError: not_included={'x'}
# after
removals = ['c', 'x']
to_remove = [r for r in removals if r in cat.categories]
cat = cat.remove_categories(to_remove) Defensive patterns
Strategy: validation
Validate before calling
def safe_remove_categories(cat, removals):
if not is_list_like(removals):
removals = [removals]
existing = set(cat.categories)
to_remove = [r for r in removals if r in existing]
return cat.remove_categories(to_remove) if to_remove else cat Type guard
def are_all_present(cat, removals) -> bool:
existing = set(cat.categories)
return all(r in existing for r in removals) Try / catch
try:
cat = cat.remove_categories(removals)
except ValueError as e:
if 'must all be in old categories' in str(e):
cat = cat.remove_categories([r for r in removals if r in set(cat.categories)])
else:
raise Prevention
- Intersect `removals` with current categories before `remove_categories` for idempotent cleanups.
- Use `set_categories` with the desired final set to bypass existence checks.
- Validate each removal against `cat.categories` before the call.
When it happens
Trigger: `cat.remove_categories(['x'])` where `'x'` is not in `cat.categories`; passing a list built from a different column or schema.
Common situations: Schema drift between runs; idempotent re-runs of a cleanup script; merging removal lists across differently-categorized DataFrames; refactoring that renamed labels upstream.
Related errors
- Cannot cast dtype to
- Cannot convert float NaN to integer
- codes cannot contain NA values
- codes need to be array-like integers
- codes need to be between -1 and len(categories)-1
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/8376b5e2168cb5f4.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/categorical.py:1492
[NaN, 'c', 'b', 'c', NaN]
Categories (2, str): ['b', 'c']
"""
from pandas import Index
if not is_list_like(removals):
removals = [removals]
removals = Index(removals).unique().dropna()
new_categories = (
self.dtype.categories.difference(removals, sort=False)
if self.dtype.ordered is True
else self.dtype.categories.difference(removals)
)
not_included = removals.difference(self.dtype.categories)
if len(not_included) != 0:
not_included = set(not_included)
raise ValueError(f"removals must all be in old categories: {not_included}")
return self.set_categories(new_categories, ordered=self.ordered, rename=False)
def remove_unused_categories(self) -> Self:
"""
Remove categories which are not used.
This method is useful when working with datasets
that undergo dynamic changes where categories may no longer be
relevant, allowing to maintain a clean, efficient data structure.
Returns
-------
Categorical
Categorical with unused categories dropped.
See Also
--------View on GitHub (pinned to 3b7651241d)