pandas-dev/pandas · error · TypeError

Categoricals can only be compared if 'categories' are the…

Error message

Categoricals can only be compared if 'categories' are the same.

What it means

Raised when comparing two Categoricals whose category sets do not agree up to permutation (checked via `_categories_match_up_to_permutation`). Even when both are ordered, the categories must be the same set; otherwise the code-by-code comparison would be meaningless because identical codes can map to different values.

Solutions

  1. Unify categories on both sides: `cat2 = cat2.set_categories(cat1.categories)` (and match `ordered`).
  2. Use `cat1.astype(cat2.dtype)` to align dtype before comparing.
  3. Compare underlying values directly: `np.asarray(cat1) == np.asarray(cat2)` if semantic equality on values is intended.
  4. Re-create both Categoricals from a shared, canonical category list.

Example fix

# before
c1 = pd.Categorical(['a', 'b'], categories=['a', 'b', 'c'], ordered=True)
c2 = pd.Categorical(['a', 'b'], categories=['a', 'b', 'd'], ordered=True)
c1 < c2  # TypeError

# after
c2 = c2.set_categories(c1.categories)
c1 < c2
Defensive patterns

Strategy: validation

Validate before calling

def align_categories(c1, c2):
    if not c1._categories_match_up_to_permutation(c2):
        c2 = c2.set_categories(c1.categories)
        if c1.ordered:
            c2 = c2.as_ordered()
    return c1, c2

Type guard

def categories_compatible(c1, c2) -> bool:
    return c1._categories_match_up_to_permutation(c2)

Try / catch

try:
    result = c1 < c2
except TypeError as e:
    if "categories" in str(e):
        c2 = c2.set_categories(c1.categories)
        result = c1 < c2
    else:
        raise

Prevention

When it happens

Trigger: `cat1 < cat2` where `cat1.categories` and `cat2.categories` differ in membership (not just ordering) — e.g., comparing a Categorical over `['a','b','c']` against one over `['a','b','d']`.

Common situations: Joining/concatenating DataFrames whose category columns were defined separately, comparing columns across versions of a dataset where categories drifted, or comparing after `add_categories`/`remove_categories` on only one side.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/3598517defe9c86c. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/categorical.py:150

    @unpack_zerodim_and_defer(opname)
    def func(self, other):
        hashable = is_hashable(other)
        if is_list_like(other) and len(other) != len(self) and not hashable:
            # in hashable case we may have a tuple that is itself a category
            raise ValueError("Lengths must match.")

        if not self.ordered:
            if opname in ["__lt__", "__gt__", "__le__", "__ge__"]:
                raise TypeError(
                    "Unordered Categoricals can only compare equality or not"
                )
        if isinstance(other, Categorical):
            # Two Categoricals can only be compared if the categories are
            # the same (maybe up to ordering, depending on ordered)

            msg = "Categoricals can only be compared if 'categories' are the same."
            if not self._categories_match_up_to_permutation(other):
                raise TypeError(msg)

            if not self.ordered and not self.categories.equals(other.categories):
                # both unordered and different order
                other_codes = recode_for_categories(
                    other.codes, other.categories, self.categories, copy=False
                )
            else:
                other_codes = other._codes

            ret = op(self._codes, other_codes)
            mask = (self._codes == -1) | (other_codes == -1)
            if mask.any():
                ret[mask] = fill_value
            return ret

        if hashable:
            if other in self.categories:
                i = self._unbox_scalar(other)

View on GitHub (pinned to 3b7651241d)