pandas-dev/pandas · error · TypeError
Cannot setitem on a Categorical with a new category
Error message
Cannot setitem on a Categorical with a new category ({fill_value}), set the categories first What it means
Raised by `_validate_setitem_value` (used by `__setitem__`, `fillna`, `where`, etc.) when assigning a fill value that is not a valid NA for the categories' dtype and is not present in the existing categories. Categoricals are closed under their category set — new labels cannot be introduced implicitly via item assignment; they must be added via `add_categories`/`set_categories` first.
Solutions
- Add the category before assigning: `cat = cat.add_categories(['z']); cat[0] = 'z'`.
- Use `set_categories` to expand the label set in one call.
- For `fillna`, ensure the fill value is in `cat.categories` or use a NaN-compatible value for the dtype.
- If the new label is not meaningful, map it to NaN instead: `cat[0] = np.nan`.
Example fix
# before
import numpy as np
cat = pd.Categorical(['a', 'b', None], categories=['a', 'b'])
cat = cat.fillna('unknown') # TypeError
# after
cat = cat.add_categories(['unknown']).fillna('unknown') Defensive patterns
Strategy: validation
Validate before calling
def safe_cat_setitem(cat, idx, value):
import numpy as np
from pandas.api.types import is_valid_na_for_dtype
if not is_valid_na_for_dtype(value, cat.categories.dtype) and value not in cat.categories:
cat = cat.add_categories([value])
cat[idx] = value
return cat Type guard
def is_assignable_to(cat, value) -> bool:
from pandas.api.types import is_valid_na_for_dtype
return is_valid_na_for_dtype(value, cat.categories.dtype) or value in cat.categories Try / catch
try:
cat[i] = value
except TypeError as e:
if 'new category' in str(e):
cat = cat.add_categories([value])
cat[i] = value
else:
raise Prevention
- Ensure any new label appears in `cat.categories` before assigning via `__setitem__`/`fillna`/`where`.
- Pre-declare all expected labels when constructing the Categorical.
- Map unexpected labels to `np.nan` if adding a category is not desired.
When it happens
Trigger: `cat[0] = 'z'` where `'z'` is not a current category; `cat.fillna('missing')` where `'missing'` is not in the categories; `cat.where(cond, 'sentinel')` with an unknown sentinel.
Common situations: Replacing missing values with a label that wasn't predeclared; assignment loops introducing new labels; downstream pipelines that expect free-form string assignment.
Related errors
- Cannot setitem on a Categorical with a new category, set…
- Cannot compare a Categorical for op
- Cannot set a Categorical with another, without identical…
- Categorical input must be list-like
- Categoricals can only be compared if 'categories' are the…
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/c13739540d9b512b.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/categorical.py:1722
Parameters
----------
fill_value : object
Returns
-------
fill_value : int
Raises
------
TypeError
"""
if is_valid_na_for_dtype(fill_value, self.categories.dtype):
fill_value = -1
elif fill_value in self.categories:
fill_value = self._unbox_scalar(fill_value)
else:
raise TypeError(
"Cannot setitem on a Categorical with a new "
f"category ({fill_value}), set the categories first"
) from None
return fill_value
@classmethod
def _validate_codes_for_dtype(cls, codes, *, dtype: CategoricalDtype) -> np.ndarray:
if isinstance(codes, ExtensionArray) and is_integer_dtype(codes.dtype):
# Avoid the implicit conversion of Int to object
if isna(codes).any():
raise ValueError("codes cannot contain NA values")
codes = codes.to_numpy(dtype=np.int64)
else:
codes = np.asarray(codes)
if len(codes) and codes.dtype.kind not in "iu":
raise ValueError("codes need to be array-like integers")
if len(codes) and (codes.max() >= len(dtype.categories) or codes.min() < -1):View on GitHub (pinned to 3b7651241d)