pandas-dev/pandas · error · ValueError

new categories must not include old categories

Error message

new categories must not include old categories: {already_included}

What it means

Raised by `add_categories` when any value in `new_categories` is already present in the current categories. Adding a duplicate would create ambiguous codes and silently inflate cardinality, so pandas rejects the call and reports the offending intersection.

Solutions

  1. Dedupe against existing categories first: `to_add = [c for c in new if c not in set(cat.categories)]; cat.add_categories(to_add)`.
  2. Use `set_categories` with the union if you want to force a final category set.
  3. Wrap repeated additions in an `if c not in cat.categories` guard.
  4. Make the operation idempotent: filter out `already_included` before calling.

Example fix

# before
cat = pd.Categorical(['b', 'c'], categories=['b', 'c'])
cat = cat.add_categories(['d', 'b'])  # ValueError: already_included={'b'}

# after
new = ['d', 'b']
to_add = [c for c in new if c not in cat.categories]
cat = cat.add_categories(to_add)
Defensive patterns

Strategy: validation

Validate before calling

def safe_add_categories(cat, new_categories):
    if not is_list_like(new_categories):
        new_categories = [new_categories]
    existing = set(cat.categories)
    to_add = [c for c in new_categories if c not in existing]
    return cat.add_categories(to_add) if to_add else cat

Type guard

def are_all_new(cat, new_categories) -> bool:
    existing = set(cat.categories)
    return all(c not in existing for c in new_categories)

Try / catch

try:
    cat = cat.add_categories(new)
except ValueError as e:
    if 'must not include old categories' in str(e):
        cat = cat.add_categories([c for c in new if c not in set(cat.categories)])
    else:
        raise

Prevention

When it happens

Trigger: `cat.add_categories(['a'])` where `'a'` is already in `cat.categories`, or `add_categories(['d', 'a'])` where `'a'` already exists.

Common situations: Idempotent re-application of category additions; merging category lists from multiple sources without deduping; loops that add categories one-by-one and re-process.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/781d8cbaeb1e16cd. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/categorical.py:1414

        set_categories : Set the categories to the specified ones.

        Examples
        --------
        >>> c = pd.Categorical(["c", "b", "c"])
        >>> c
        ['c', 'b', 'c']
        Categories (2, str): ['b', 'c']

        >>> c.add_categories(["d", "a"])
        ['c', 'b', 'c']
        Categories (4, str): ['b', 'c', 'd', 'a']
        """

        if not is_list_like(new_categories):
            new_categories = [new_categories]
        already_included = set(new_categories) & set(self.dtype.categories)
        if len(already_included) != 0:
            raise ValueError(
                f"new categories must not include old categories: {already_included}"
            )

        if hasattr(new_categories, "dtype"):
            from pandas import Series

            dtype = find_common_type(
                [self.dtype.categories.dtype, new_categories.dtype]
            )
            new_categories = Series(
                list(self.dtype.categories) + list(new_categories), dtype=dtype
            )
        else:
            new_categories = list(self.dtype.categories) + list(new_categories)

        new_dtype = CategoricalDtype(new_categories, self.ordered)
        cat = self.copy()
        codes = coerce_indexer_dtype(cat._ndarray, new_dtype.categories)

View on GitHub (pinned to 3b7651241d)