pandas-dev/pandas · error · ValueError
values should be unique if codes is not None
Error message
values should be unique if codes is not None
What it means
safe_sort() requires that the values array contains no duplicates when a codes argument is provided — otherwise the code remapping is ambiguous (which duplicate should a code point to?). This specific check is in the counting-sort fast path, used for integer values spanning a modest range. The counting pass detects duplicates by comparing the cumulative count of unique values against the total length. If they differ, duplicates exist and a ValueError is raised.
Solutions
- Deduplicate values before calling safe_sort: unique_values = np.unique(values); safe_sort(unique_values, codes=codes).
- If you are certain values are unique and want to skip the check, pass assume_unique=True.
- Verify that the values array you are passing is actually the set of unique categories, not the raw data.
Example fix
# before safe_sort(np.array([1, 2, 2, 3]), codes=np.array([0, 1, 2])) # after unique_vals = np.unique(np.array([1, 2, 2, 3])) # [1, 2, 3] safe_sort(unique_vals, codes=np.array([0, 1, 2]))
Defensive patterns
Strategy: validation
Validate before calling
def safe_safe_sort(values, codes=None, assume_unique=False, **kwargs):
if codes is not None and not assume_unique:
arr = np.asarray(values)
if arr.dtype.kind in 'iu':
if len(np.unique(arr)) != len(arr):
raise ValueError("values must be unique when codes is not None")
return pd.core.algorithms.safe_sort(values, codes=codes, assume_unique=assume_unique, **kwargs) Try / catch
try:
result = safe_sort(values, codes)
except ValueError as e:
if "values should be unique" in str(e):
values = np.unique(values)
result = safe_sort(values, codes, assume_unique=True)
else:
raise Prevention
- Deduplicate values with np.unique() before calling safe_sort with codes.
- Pass assume_unique=True only when you have verified uniqueness yourself.
- Verify that the values array is the set of categories, not the raw data.
When it happens
Trigger: Calling safe_sort with integer values that contain duplicates and a non-None codes array, without setting assume_unique=True. Passing a factorized array where the unique values were not actually deduplicated before calling safe_sort.
Common situations: Using safe_sort in categorical reordering logic where the categories array accidentally contains a repeated value. Chaining operations where a previous step was expected to unique-ify the values but did not. Passing the wrong array (e.g., the full data instead of unique categories) as values.
Related errors
- Only list-like objects or None are allowed to be passed to…
- Only np.ndarray, ExtensionArray, and Index objects are…
- by_row= not allowed
- cannot broadcast result
- cannot combine transform and aggregation operations
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/eaede111f49e5ba4.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/algorithms.py:1752
codes = ensure_platform_int(np.asarray(codes))
# ranks[i] gives the position of values[i] in `ordered`
if use_counting:
arr = cast("np.ndarray", values)
if arr.dtype.kind == "i":
# go through int64 so differences don't overflow narrower signed
# dtypes; int64 wraparound is exact since the true differences
# are within rng_size
shifted = arr.astype(np.int64, copy=False) - vmin
else:
# unsigned: differences always fit the unsigned dtype
shifted = arr - vmin
present = np.zeros(rng_size, dtype=bool)
present[shifted] = True
counts = present.cumsum(dtype=np.intp)
# the counting pass gives the uniqueness check for free
if not assume_unique and counts[-1] != len(values):
raise ValueError("values should be unique if codes is not None")
ranks = counts[shifted]
ranks -= 1
else:
if not assume_unique and not len(unique(values)) == len(values):
raise ValueError("values should be unique if codes is not None")
if sorter is not None:
# sorter is a permutation, so scatter is a faster equivalent to
# `ranks = sorter.argsort()`
ranks = np.empty(len(sorter), dtype=np.intp)
ranks[sorter] = np.arange(len(sorter), dtype=np.intp)
else:
# mixed types
# error: Argument 1 to "_get_hashtable_algo" has incompatible type
# "Union[Index, ExtensionArray, ndarray[Any, Any]]"; expected
# "ndarray[Any, Any]"
hash_klass, values = _get_hashtable_algo(values) # type: ignore[arg-type]
t = hash_klass(len(values))View on GitHub (pinned to 3b7651241d)