pandas-dev/pandas · error · ValueError
The categories must be provided in 'categories' or 'dtype'…
Error message
The categories must be provided in 'categories' or 'dtype'. Both were None.
What it means
Raised by `Categorical.from_codes` when neither the `categories` argument nor a `CategoricalDtype` with non-null `.categories` was supplied. `from_codes` takes raw integer codes, so it has no way to infer category labels from the values themselves — the caller must tell it what each code means.
Solutions
- Pass an explicit `categories` list whose length exceeds the max code: `pd.Categorical.from_codes([0, 1, 0], categories=['a', 'b'])`.
- Pass a fully-formed `CategoricalDtype`: `from_codes(codes, dtype=dtype)`.
- Validate that the loaded dtype has non-null `.categories` before calling `from_codes`.
- If you have raw values (not codes), use `pd.Categorical(values)` instead.
Example fix
# before import pandas as pd cat = pd.Categorical.from_codes([0, 1, 0, 1]) # ValueError # after cat = pd.Categorical.from_codes([0, 1, 0, 1], categories=['a', 'b'])
Defensive patterns
Strategy: validation
Validate before calling
def from_codes_safe(codes, categories=None, dtype=None):
if categories is None and (dtype is None or dtype.categories is None):
raise ValueError('supply categories= or a CategoricalDtype with categories')
return pd.Categorical.from_codes(codes, categories=categories, dtype=dtype) Type guard
def has_categories(categories, dtype) -> bool:
return categories is not None or (dtype is not None and dtype.categories is not None) Prevention
- Always pass `categories=` to `from_codes` — there is no inference path.
- Assert `dtype.categories is not None` when loading a dtype from disk before calling `from_codes`.
- If you have value labels rather than codes, use `pd.Categorical(values)`.
When it happens
Trigger: Calling `pd.Categorical.from_codes([0, 1, 0])` with no `categories=` and no `dtype=`, or passing a `CategoricalDtype(categories=None)`.
Common situations: Refactoring code that previously built a Categorical via the normal constructor (which can infer categories); deserialization pipelines that read codes from a binary blob but forgot to also read the label table; copy-paste errors omitting the keyword.
Related errors
- codes cannot contain NA values
- codes need to be array-like integers
- codes need to be between -1 and len(categories)-1
- items in new_categories are not the same as in old…
- Cannot cast dtype to
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/6b5ea518977723c1.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/categorical.py:796
codes : The category codes of the categorical.
CategoricalIndex : An Index with an underlying ``Categorical``.
Examples
--------
>>> dtype = pd.CategoricalDtype(["a", "b"], ordered=True)
>>> pd.Categorical.from_codes(codes=[0, 1, 0, 1], dtype=dtype)
['a', 'b', 'a', 'b']
Categories (2, str): ['a' < 'b']
"""
dtype = CategoricalDtype._from_values_or_dtype(
categories=categories, ordered=ordered, dtype=dtype
)
if dtype.categories is None:
msg = (
"The categories must be provided in 'categories' or "
"'dtype'. Both were None."
)
raise ValueError(msg)
if validate:
# beware: non-valid codes may segfault
codes = cls._validate_codes_for_dtype(codes, dtype=dtype)
return cls._simple_new(codes, dtype=dtype)
# ------------------------------------------------------------------
# Categories/Codes/Ordered
@property
def categories(self) -> Index:
"""
The categories of this categorical.
Setting assigns new values to each category (effectively a rename of
each individual category).
View on GitHub (pinned to 3b7651241d)