pandas-dev/pandas · error · TypeError
Cannot setitem on a Categorical with a new category, set…
Error message
Cannot setitem on a Categorical with a new category, set the categories first
What it means
Raised in Categorical._validate_listlike when a value being assigned is not in the existing categories (and is not NaN). Categoricals are closed under their category set; introducing a new value would require extending the set, which pandas will not do implicitly. NaN is always allowed so setting missing values works.
Solutions
- Extend categories first: cat = cat.add_categories(['new_value']) then assign.
- Use set_categories with the full expected set: df['col'] = df['col'].cat.set_categories([... , 'new_value']).
- Assign np.nan for values you want to drop rather than recategorize.
- Rebuild the categorical from object data once new values are present: df['col'].astype(object).astype('category').
Example fix
# before cat = pd.Categorical(['a','b'], categories=['a','b']) cat[0] = 'c' # TypeError # after cat = cat.add_categories(['c']) cat[0] = 'c'
Defensive patterns
Strategy: validation
Validate before calling
new_vals = pd.Index(value).difference(df['col'].cat.categories)
if not new_vals.isna().all():
df['col'] = df['col'].cat.add_categories(new_vals.dropna())
df['col'].iloc[i] = value Type guard
def all_in_categories(values, cats: pd.Index) -> bool:
return pd.Index(values).difference(cats).empty or pd.Index(values).difference(cats).isna().all() Try / catch
try:
df['col'].iloc[i] = value
except TypeError as e:
if 'new category' in str(e):
df['col'] = df['col'].cat.add_categories([value])
df['col'].iloc[i] = value
else:
raise Prevention
- Collect all expected labels up front and pass them to pd.Categorical(categories=...).
- Wrap incoming new data with add_categories before assignment in ETL jobs.
- For unseen labels in production, decide policy (add vs NaN) at the schema layer.
When it happens
Trigger: cat[0] = 'd' where 'd' is not a category. df['catcol'] = 'new_value' with 'new_value' absent from categories. fillna with a value not in the categories. .loc assignment introducing a previously unseen label.
Common situations: Appending new rows to a DataFrame whose column was categorized on a snapshot. Encoding unseen labels after a train/test split. Replacing a deprecated category value with a new one without updating categories.
Related errors
- Cannot set a Categorical with another, without identical…
- Cannot setitem on a Categorical with a new category
- Categoricals can only be compared if 'categories' are the…
- items in new_categories are not the same as in old…
- The categories must be provided in 'categories' or 'dtype'…
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/07711c7c5a3d799e.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/categorical.py:2483
raise TypeError(
"Cannot set a Categorical with another, "
"without identical categories"
)
# dtype equality implies categories_match_up_to_permutation
value = self._encode_with_my_categories(value)
return value._codes
from pandas import Index
# tupleize_cols=False for e.g. test_fillna_iterable_category GH#41914
to_add = Index._with_infer(value, tupleize_cols=False, copy=False).difference(
self.categories
)
# no assignments of values not in categories, but it's always ok to set
# something to np.nan
if len(to_add) and not isna(to_add).all():
raise TypeError(
"Cannot setitem on a Categorical with a new "
"category, set the categories first"
)
codes = self.categories.get_indexer(value)
return codes.astype(self._ndarray.dtype, copy=False)
def _reverse_indexer(self) -> dict[Hashable, npt.NDArray[np.intp]]:
"""
Compute the inverse of a categorical, returning
a dict of categories -> indexers.
*This is an internal function*
Returns
-------
Dict[Hashable, np.ndarray[np.intp]]
dict of categories -> indexersView on GitHub (pinned to 3b7651241d)