{"record":{"id":"07711c7c5a3d799e","repo":"pandas-dev/pandas","slug":"cannot-setitem-on-a-categorical-with-a-new-categor-07711c","errorCode":null,"errorMessage":"Cannot setitem on a Categorical with a new category, set the categories first","messagePattern":"Cannot setitem on a Categorical with a new category, set the categories first","errorType":"exception","errorClass":"TypeError","httpStatus":null,"severity":"error","filePath":"pandas/core/arrays/categorical.py","lineNumber":2483,"sourceCode":"                raise TypeError(\n                    \"Cannot set a Categorical with another, \"\n                    \"without identical categories\"\n                )\n            # dtype equality implies categories_match_up_to_permutation\n            value = self._encode_with_my_categories(value)\n            return value._codes\n\n        from pandas import Index\n\n        # tupleize_cols=False for e.g. test_fillna_iterable_category GH#41914\n        to_add = Index._with_infer(value, tupleize_cols=False, copy=False).difference(\n            self.categories\n        )\n\n        # no assignments of values not in categories, but it's always ok to set\n        # something to np.nan\n        if len(to_add) and not isna(to_add).all():\n            raise TypeError(\n                \"Cannot setitem on a Categorical with a new \"\n                \"category, set the categories first\"\n            )\n\n        codes = self.categories.get_indexer(value)\n        return codes.astype(self._ndarray.dtype, copy=False)\n\n    def _reverse_indexer(self) -> dict[Hashable, npt.NDArray[np.intp]]:\n        \"\"\"\n        Compute the inverse of a categorical, returning\n        a dict of categories -> indexers.\n\n        *This is an internal function*\n\n        Returns\n        -------\n        Dict[Hashable, np.ndarray[np.intp]]\n            dict of categories -> indexers","sourceCodeStart":2465,"sourceCodeEnd":2501,"githubUrl":"https://github.com/pandas-dev/pandas/blob/3b7651241d4da534b3559b60ef128e1c34f54116/pandas/core/arrays/categorical.py#L2465-L2501","documentation":"Raised in Categorical._validate_listlike when a value being assigned is not in the existing categories (and is not NaN). Categoricals are closed under their category set; introducing a new value would require extending the set, which pandas will not do implicitly. NaN is always allowed so setting missing values works.","triggerScenarios":"cat[0] = 'd' where 'd' is not a category. df['catcol'] = 'new_value' with 'new_value' absent from categories. fillna with a value not in the categories. .loc assignment introducing a previously unseen label.","commonSituations":"Appending new rows to a DataFrame whose column was categorized on a snapshot. Encoding unseen labels after a train/test split. Replacing a deprecated category value with a new one without updating categories.","solutions":["Extend categories first: cat = cat.add_categories(['new_value']) then assign.","Use set_categories with the full expected set: df['col'] = df['col'].cat.set_categories([... , 'new_value']).","Assign np.nan for values you want to drop rather than recategorize.","Rebuild the categorical from object data once new values are present: df['col'].astype(object).astype('category')."],"exampleFix":"# before\ncat = pd.Categorical(['a','b'], categories=['a','b'])\ncat[0] = 'c'  # TypeError\n\n# after\ncat = cat.add_categories(['c'])\ncat[0] = 'c'","handlingStrategy":"validation","validationCode":"new_vals = pd.Index(value).difference(df['col'].cat.categories)\nif not new_vals.isna().all():\n    df['col'] = df['col'].cat.add_categories(new_vals.dropna())\ndf['col'].iloc[i] = value","typeGuard":"def all_in_categories(values, cats: pd.Index) -> bool:\n    return pd.Index(values).difference(cats).empty or pd.Index(values).difference(cats).isna().all()","tryCatchPattern":"try:\n    df['col'].iloc[i] = value\nexcept TypeError as e:\n    if 'new category' in str(e):\n        df['col'] = df['col'].cat.add_categories([value])\n        df['col'].iloc[i] = value\n    else:\n        raise","preventionTips":["Collect all expected labels up front and pass them to pd.Categorical(categories=...).","Wrap incoming new data with add_categories before assignment in ETL jobs.","For unseen labels in production, decide policy (add vs NaN) at the schema layer."],"tags":["categorical","setitem","new-category","categories"],"backgroundTag":null,"analyzedSha":"3b7651241d4da534b3559b60ef128e1c34f54116","analyzedAt":"2026-08-11T22:10:44.015Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}