pandas-dev/pandas · error · TypeError

Input must be list-like

Error message

Input must be list-like

What it means

Raised by factorize_from_iterables (categorical.py:3217) when the input values is not list-like. factorize requires an iterable of values to assign integer codes; a scalar has nothing to enumerate. The is_list_like guard rejects scalars (int, str-as-scalar, None) before any encoding happens.

Solutions

  1. Wrap the scalar in a list: pd.factorize([value]).
  2. Ensure the variable feeding factorize is always iterable; coerce with np.asarray or pd.Index.
  3. For strings intended as a sequence of characters, pass list(s) explicitly.
  4. Add an isinstance(x, (list, tuple, np.ndarray, pd.Series)) guard upstream.

Example fix

# before
pd.factorize('a')  # TypeError: Input must be list-like

# after
pd.factorize(['a'])  # (array([0]), Index(['a'], dtype='object'))
Defensive patterns

Strategy: validation

Validate before calling

import numpy as np
if not isinstance(values, (list, tuple, np.ndarray, pd.Series, pd.Index)):
    values = [values]
codes, uniques = pd.factorize(values)

Type guard

from collections.abc import Iterable
import pandas as pd

def ensure_listlike(x):
    if isinstance(x, (str, bytes)):
        return [x]
    if isinstance(x, Iterable):
        return list(x)
    return [x]

Prevention

When it happens

Trigger: pd.factorize(5), pd.factorize('a') (single string, not a list of strings). pd.Categorical.from_codes with a scalar codes argument. Internal calls from pd.cut/qcut when the input collapses to a scalar. Passing a single datetime instead of a DatetimeIndex.

Common situations: User passes a bare string expecting it to be treated as a one-element sequence. Variable that was expected to be a list is actually a single value from upstream logic. Calling factorize inside a loop where some iterations yield a scalar.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/06995c22537ec8e8. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/categorical.py:3217

    """
    Factorize an input `values` into `categories` and `codes`. Preserves
    categorical dtype in `categories`.

    Parameters
    ----------
    values : list-like

    Returns
    -------
    codes : ndarray
    categories : Index
        If `values` has a categorical dtype, then `categories` is
        a CategoricalIndex keeping the categories and order of `values`.
    """
    from pandas import CategoricalIndex

    if not is_list_like(values):
        raise TypeError("Input must be list-like")

    categories: Index

    vdtype = getattr(values, "dtype", None)
    if isinstance(vdtype, CategoricalDtype):
        values = extract_array(values)
        # The Categorical we want to build has the same categories
        # as values but its codes are by def [0, ..., len(n_categories) - 1]
        cat_codes = np.arange(len(values.categories), dtype=values.codes.dtype)
        cat = Categorical.from_codes(cat_codes, dtype=values.dtype, validate=False)

        categories = CategoricalIndex(cat)
        codes = values.codes
    else:
        # The value of ordered is irrelevant since we don't use cat as such,
        # but only the resulting categories, the order of which is independent
        # from ordered. Set ordered to False as default. See GH #15457
        cat = Categorical(values, ordered=False)

View on GitHub (pinned to 3b7651241d)