pandas-dev/pandas · error · ValueError

values should be unique if codes is not None

Error message

values should be unique if codes is not None

What it means

safe_sort() requires that the values array contains no duplicates when a codes argument is provided — otherwise the code remapping is ambiguous (which duplicate should a code point to?). This specific check is in the counting-sort fast path, used for integer values spanning a modest range. The counting pass detects duplicates by comparing the cumulative count of unique values against the total length. If they differ, duplicates exist and a ValueError is raised.

Solutions

  1. Deduplicate values before calling safe_sort: unique_values = np.unique(values); safe_sort(unique_values, codes=codes).
  2. If you are certain values are unique and want to skip the check, pass assume_unique=True.
  3. Verify that the values array you are passing is actually the set of unique categories, not the raw data.

Example fix

# before
safe_sort(np.array([1, 2, 2, 3]), codes=np.array([0, 1, 2]))

# after
unique_vals = np.unique(np.array([1, 2, 2, 3]))  # [1, 2, 3]
safe_sort(unique_vals, codes=np.array([0, 1, 2]))
Defensive patterns

Strategy: validation

Validate before calling

def safe_safe_sort(values, codes=None, assume_unique=False, **kwargs):
    if codes is not None and not assume_unique:
        arr = np.asarray(values)
        if arr.dtype.kind in 'iu':
            if len(np.unique(arr)) != len(arr):
                raise ValueError("values must be unique when codes is not None")
    return pd.core.algorithms.safe_sort(values, codes=codes, assume_unique=assume_unique, **kwargs)

Try / catch

try:
    result = safe_sort(values, codes)
except ValueError as e:
    if "values should be unique" in str(e):
        values = np.unique(values)
        result = safe_sort(values, codes, assume_unique=True)
    else:
        raise

Prevention

When it happens

Trigger: Calling safe_sort with integer values that contain duplicates and a non-None codes array, without setting assume_unique=True. Passing a factorized array where the unique values were not actually deduplicated before calling safe_sort.

Common situations: Using safe_sort in categorical reordering logic where the categories array accidentally contains a repeated value. Chaining operations where a previous step was expected to unique-ify the values but did not. Passing the wrong array (e.g., the full data instead of unique categories) as values.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/eaede111f49e5ba4. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/algorithms.py:1752

    codes = ensure_platform_int(np.asarray(codes))

    # ranks[i] gives the position of values[i] in `ordered`
    if use_counting:
        arr = cast("np.ndarray", values)
        if arr.dtype.kind == "i":
            # go through int64 so differences don't overflow narrower signed
            #  dtypes; int64 wraparound is exact since the true differences
            #  are within rng_size
            shifted = arr.astype(np.int64, copy=False) - vmin
        else:
            # unsigned: differences always fit the unsigned dtype
            shifted = arr - vmin
        present = np.zeros(rng_size, dtype=bool)
        present[shifted] = True
        counts = present.cumsum(dtype=np.intp)
        # the counting pass gives the uniqueness check for free
        if not assume_unique and counts[-1] != len(values):
            raise ValueError("values should be unique if codes is not None")
        ranks = counts[shifted]
        ranks -= 1
    else:
        if not assume_unique and not len(unique(values)) == len(values):
            raise ValueError("values should be unique if codes is not None")

        if sorter is not None:
            # sorter is a permutation, so scatter is a faster equivalent to
            #  `ranks = sorter.argsort()`
            ranks = np.empty(len(sorter), dtype=np.intp)
            ranks[sorter] = np.arange(len(sorter), dtype=np.intp)
        else:
            # mixed types
            # error: Argument 1 to "_get_hashtable_algo" has incompatible type
            # "Union[Index, ExtensionArray, ndarray[Any, Any]]"; expected
            # "ndarray[Any, Any]"
            hash_klass, values = _get_hashtable_algo(values)  # type: ignore[arg-type]
            t = hash_klass(len(values))

View on GitHub (pinned to 3b7651241d)