pandas-dev/pandas · error · TypeError
Input must be list-like
Error message
Input must be list-like
What it means
Raised by factorize_from_iterables (categorical.py:3217) when the input values is not list-like. factorize requires an iterable of values to assign integer codes; a scalar has nothing to enumerate. The is_list_like guard rejects scalars (int, str-as-scalar, None) before any encoding happens.
Solutions
- Wrap the scalar in a list: pd.factorize([value]).
- Ensure the variable feeding factorize is always iterable; coerce with np.asarray or pd.Index.
- For strings intended as a sequence of characters, pass list(s) explicitly.
- Add an isinstance(x, (list, tuple, np.ndarray, pd.Series)) guard upstream.
Example fix
# before
pd.factorize('a') # TypeError: Input must be list-like
# after
pd.factorize(['a']) # (array([0]), Index(['a'], dtype='object')) Defensive patterns
Strategy: validation
Validate before calling
import numpy as np
if not isinstance(values, (list, tuple, np.ndarray, pd.Series, pd.Index)):
values = [values]
codes, uniques = pd.factorize(values) Type guard
from collections.abc import Iterable
import pandas as pd
def ensure_listlike(x):
if isinstance(x, (str, bytes)):
return [x]
if isinstance(x, Iterable):
return list(x)
return [x] Prevention
- Wrap scalar inputs in a list before factorize.
- Document factorize entry points as list-like only.
When it happens
Trigger: pd.factorize(5), pd.factorize('a') (single string, not a list of strings). pd.Categorical.from_codes with a scalar codes argument. Internal calls from pd.cut/qcut when the input collapses to a scalar. Passing a single datetime instead of a DatetimeIndex.
Common situations: User passes a bare string expecting it to be treated as a one-element sequence. Variable that was expected to be a list is actually a single value from upstream logic. Calling factorize inside a loop where some iterations yield a scalar.
Related errors
- Cannot compare a Categorical for op
- invalid na_position
- 'values' is not ordered, please explicitly specify the…
- > 1 ndim Categorical are not supported at this time
- Accumulation not supported for
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/06995c22537ec8e8.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/categorical.py:3217
"""
Factorize an input `values` into `categories` and `codes`. Preserves
categorical dtype in `categories`.
Parameters
----------
values : list-like
Returns
-------
codes : ndarray
categories : Index
If `values` has a categorical dtype, then `categories` is
a CategoricalIndex keeping the categories and order of `values`.
"""
from pandas import CategoricalIndex
if not is_list_like(values):
raise TypeError("Input must be list-like")
categories: Index
vdtype = getattr(values, "dtype", None)
if isinstance(vdtype, CategoricalDtype):
values = extract_array(values)
# The Categorical we want to build has the same categories
# as values but its codes are by def [0, ..., len(n_categories) - 1]
cat_codes = np.arange(len(values.categories), dtype=values.codes.dtype)
cat = Categorical.from_codes(cat_codes, dtype=values.dtype, validate=False)
categories = CategoricalIndex(cat)
codes = values.codes
else:
# The value of ordered is irrelevant since we don't use cat as such,
# but only the resulting categories, the order of which is independent
# from ordered. Set ordered to False as default. See GH #15457
cat = Categorical(values, ordered=False)View on GitHub (pinned to 3b7651241d)