pandas-dev/pandas · error · ValueError
new categories must not include old categories
Error message
new categories must not include old categories: {already_included} What it means
Raised by `add_categories` when any value in `new_categories` is already present in the current categories. Adding a duplicate would create ambiguous codes and silently inflate cardinality, so pandas rejects the call and reports the offending intersection.
Solutions
- Dedupe against existing categories first: `to_add = [c for c in new if c not in set(cat.categories)]; cat.add_categories(to_add)`.
- Use `set_categories` with the union if you want to force a final category set.
- Wrap repeated additions in an `if c not in cat.categories` guard.
- Make the operation idempotent: filter out `already_included` before calling.
Example fix
# before
cat = pd.Categorical(['b', 'c'], categories=['b', 'c'])
cat = cat.add_categories(['d', 'b']) # ValueError: already_included={'b'}
# after
new = ['d', 'b']
to_add = [c for c in new if c not in cat.categories]
cat = cat.add_categories(to_add) Defensive patterns
Strategy: validation
Validate before calling
def safe_add_categories(cat, new_categories):
if not is_list_like(new_categories):
new_categories = [new_categories]
existing = set(cat.categories)
to_add = [c for c in new_categories if c not in existing]
return cat.add_categories(to_add) if to_add else cat Type guard
def are_all_new(cat, new_categories) -> bool:
existing = set(cat.categories)
return all(c not in existing for c in new_categories) Try / catch
try:
cat = cat.add_categories(new)
except ValueError as e:
if 'must not include old categories' in str(e):
cat = cat.add_categories([c for c in new if c not in set(cat.categories)])
else:
raise Prevention
- Filter out already-present categories before `add_categories` to make additions idempotent.
- When merging category lists, dedupe against the existing set.
- Use `set_categories` with the union if you want a single declarative update.
When it happens
Trigger: `cat.add_categories(['a'])` where `'a'` is already in `cat.categories`, or `add_categories(['d', 'a'])` where `'a'` already exists.
Common situations: Idempotent re-application of category additions; merging category lists from multiple sources without deduping; loops that add categories one-by-one and re-process.
Related errors
- Cannot cast dtype to
- Cannot convert float NaN to integer
- codes cannot contain NA values
- codes need to be array-like integers
- codes need to be between -1 and len(categories)-1
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/781d8cbaeb1e16cd.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/categorical.py:1414
set_categories : Set the categories to the specified ones.
Examples
--------
>>> c = pd.Categorical(["c", "b", "c"])
>>> c
['c', 'b', 'c']
Categories (2, str): ['b', 'c']
>>> c.add_categories(["d", "a"])
['c', 'b', 'c']
Categories (4, str): ['b', 'c', 'd', 'a']
"""
if not is_list_like(new_categories):
new_categories = [new_categories]
already_included = set(new_categories) & set(self.dtype.categories)
if len(already_included) != 0:
raise ValueError(
f"new categories must not include old categories: {already_included}"
)
if hasattr(new_categories, "dtype"):
from pandas import Series
dtype = find_common_type(
[self.dtype.categories.dtype, new_categories.dtype]
)
new_categories = Series(
list(self.dtype.categories) + list(new_categories), dtype=dtype
)
else:
new_categories = list(self.dtype.categories) + list(new_categories)
new_dtype = CategoricalDtype(new_categories, self.ordered)
cat = self.copy()
codes = coerce_indexer_dtype(cat._ndarray, new_dtype.categories)View on GitHub (pinned to 3b7651241d)