pandas-dev/pandas · error · TypeError

Cannot setitem on a Categorical with a new category, set…

Error message

Cannot setitem on a Categorical with a new category, set the categories first

What it means

Raised in Categorical._validate_listlike when a value being assigned is not in the existing categories (and is not NaN). Categoricals are closed under their category set; introducing a new value would require extending the set, which pandas will not do implicitly. NaN is always allowed so setting missing values works.

Solutions

  1. Extend categories first: cat = cat.add_categories(['new_value']) then assign.
  2. Use set_categories with the full expected set: df['col'] = df['col'].cat.set_categories([... , 'new_value']).
  3. Assign np.nan for values you want to drop rather than recategorize.
  4. Rebuild the categorical from object data once new values are present: df['col'].astype(object).astype('category').

Example fix

# before
cat = pd.Categorical(['a','b'], categories=['a','b'])
cat[0] = 'c'  # TypeError

# after
cat = cat.add_categories(['c'])
cat[0] = 'c'
Defensive patterns

Strategy: validation

Validate before calling

new_vals = pd.Index(value).difference(df['col'].cat.categories)
if not new_vals.isna().all():
    df['col'] = df['col'].cat.add_categories(new_vals.dropna())
df['col'].iloc[i] = value

Type guard

def all_in_categories(values, cats: pd.Index) -> bool:
    return pd.Index(values).difference(cats).empty or pd.Index(values).difference(cats).isna().all()

Try / catch

try:
    df['col'].iloc[i] = value
except TypeError as e:
    if 'new category' in str(e):
        df['col'] = df['col'].cat.add_categories([value])
        df['col'].iloc[i] = value
    else:
        raise

Prevention

When it happens

Trigger: cat[0] = 'd' where 'd' is not a category. df['catcol'] = 'new_value' with 'new_value' absent from categories. fillna with a value not in the categories. .loc assignment introducing a previously unseen label.

Common situations: Appending new rows to a DataFrame whose column was categorized on a snapshot. Encoding unseen labels after a train/test split. Replacing a deprecated category value with a new one without updating categories.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/07711c7c5a3d799e. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/categorical.py:2483

                raise TypeError(
                    "Cannot set a Categorical with another, "
                    "without identical categories"
                )
            # dtype equality implies categories_match_up_to_permutation
            value = self._encode_with_my_categories(value)
            return value._codes

        from pandas import Index

        # tupleize_cols=False for e.g. test_fillna_iterable_category GH#41914
        to_add = Index._with_infer(value, tupleize_cols=False, copy=False).difference(
            self.categories
        )

        # no assignments of values not in categories, but it's always ok to set
        # something to np.nan
        if len(to_add) and not isna(to_add).all():
            raise TypeError(
                "Cannot setitem on a Categorical with a new "
                "category, set the categories first"
            )

        codes = self.categories.get_indexer(value)
        return codes.astype(self._ndarray.dtype, copy=False)

    def _reverse_indexer(self) -> dict[Hashable, npt.NDArray[np.intp]]:
        """
        Compute the inverse of a categorical, returning
        a dict of categories -> indexers.

        *This is an internal function*

        Returns
        -------
        Dict[Hashable, np.ndarray[np.intp]]
            dict of categories -> indexers

View on GitHub (pinned to 3b7651241d)