pandas-dev/pandas · error · TypeError
Categoricals can only be compared if 'categories' are the…
Error message
Categoricals can only be compared if 'categories' are the same.
What it means
Raised when comparing two Categoricals whose category sets do not agree up to permutation (checked via `_categories_match_up_to_permutation`). Even when both are ordered, the categories must be the same set; otherwise the code-by-code comparison would be meaningless because identical codes can map to different values.
Solutions
- Unify categories on both sides: `cat2 = cat2.set_categories(cat1.categories)` (and match `ordered`).
- Use `cat1.astype(cat2.dtype)` to align dtype before comparing.
- Compare underlying values directly: `np.asarray(cat1) == np.asarray(cat2)` if semantic equality on values is intended.
- Re-create both Categoricals from a shared, canonical category list.
Example fix
# before c1 = pd.Categorical(['a', 'b'], categories=['a', 'b', 'c'], ordered=True) c2 = pd.Categorical(['a', 'b'], categories=['a', 'b', 'd'], ordered=True) c1 < c2 # TypeError # after c2 = c2.set_categories(c1.categories) c1 < c2
Defensive patterns
Strategy: validation
Validate before calling
def align_categories(c1, c2):
if not c1._categories_match_up_to_permutation(c2):
c2 = c2.set_categories(c1.categories)
if c1.ordered:
c2 = c2.as_ordered()
return c1, c2 Type guard
def categories_compatible(c1, c2) -> bool:
return c1._categories_match_up_to_permutation(c2) Try / catch
try:
result = c1 < c2
except TypeError as e:
if "categories" in str(e):
c2 = c2.set_categories(c1.categories)
result = c1 < c2
else:
raise Prevention
- Standardize category dtype in one place and reuse the dtype object across columns/DataFrames.
- When concatenating DataFrames, unify category dtypes via `.astype(shared_dtype)` first.
- Document the canonical category list per categorical column in your schema.
When it happens
Trigger: `cat1 < cat2` where `cat1.categories` and `cat2.categories` differ in membership (not just ordering) — e.g., comparing a Categorical over `['a','b','c']` against one over `['a','b','d']`.
Common situations: Joining/concatenating DataFrames whose category columns were defined separately, comparing columns across versions of a dataset where categories drifted, or comparing after `add_categories`/`remove_categories` on only one side.
Related errors
- Cannot compare a Categorical for op
- Unordered Categoricals can only compare equality or not
- Cannot set a Categorical with another, without identical…
- Cannot setitem on a Categorical with a new category
- Cannot setitem on a Categorical with a new category, set…
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/3598517defe9c86c.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/categorical.py:150
@unpack_zerodim_and_defer(opname)
def func(self, other):
hashable = is_hashable(other)
if is_list_like(other) and len(other) != len(self) and not hashable:
# in hashable case we may have a tuple that is itself a category
raise ValueError("Lengths must match.")
if not self.ordered:
if opname in ["__lt__", "__gt__", "__le__", "__ge__"]:
raise TypeError(
"Unordered Categoricals can only compare equality or not"
)
if isinstance(other, Categorical):
# Two Categoricals can only be compared if the categories are
# the same (maybe up to ordering, depending on ordered)
msg = "Categoricals can only be compared if 'categories' are the same."
if not self._categories_match_up_to_permutation(other):
raise TypeError(msg)
if not self.ordered and not self.categories.equals(other.categories):
# both unordered and different order
other_codes = recode_for_categories(
other.codes, other.categories, self.categories, copy=False
)
else:
other_codes = other._codes
ret = op(self._codes, other_codes)
mask = (self._codes == -1) | (other_codes == -1)
if mask.any():
ret[mask] = fill_value
return ret
if hashable:
if other in self.categories:
i = self._unbox_scalar(other)View on GitHub (pinned to 3b7651241d)