{"record":{"id":"eaede111f49e5ba4","repo":"pandas-dev/pandas","slug":"values-should-be-unique-if-codes-is-not-none","errorCode":null,"errorMessage":"values should be unique if codes is not None","messagePattern":"values should be unique if codes is not None","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"pandas/core/algorithms.py","lineNumber":1752,"sourceCode":"    codes = ensure_platform_int(np.asarray(codes))\n\n    # ranks[i] gives the position of values[i] in `ordered`\n    if use_counting:\n        arr = cast(\"np.ndarray\", values)\n        if arr.dtype.kind == \"i\":\n            # go through int64 so differences don't overflow narrower signed\n            #  dtypes; int64 wraparound is exact since the true differences\n            #  are within rng_size\n            shifted = arr.astype(np.int64, copy=False) - vmin\n        else:\n            # unsigned: differences always fit the unsigned dtype\n            shifted = arr - vmin\n        present = np.zeros(rng_size, dtype=bool)\n        present[shifted] = True\n        counts = present.cumsum(dtype=np.intp)\n        # the counting pass gives the uniqueness check for free\n        if not assume_unique and counts[-1] != len(values):\n            raise ValueError(\"values should be unique if codes is not None\")\n        ranks = counts[shifted]\n        ranks -= 1\n    else:\n        if not assume_unique and not len(unique(values)) == len(values):\n            raise ValueError(\"values should be unique if codes is not None\")\n\n        if sorter is not None:\n            # sorter is a permutation, so scatter is a faster equivalent to\n            #  `ranks = sorter.argsort()`\n            ranks = np.empty(len(sorter), dtype=np.intp)\n            ranks[sorter] = np.arange(len(sorter), dtype=np.intp)\n        else:\n            # mixed types\n            # error: Argument 1 to \"_get_hashtable_algo\" has incompatible type\n            # \"Union[Index, ExtensionArray, ndarray[Any, Any]]\"; expected\n            # \"ndarray[Any, Any]\"\n            hash_klass, values = _get_hashtable_algo(values)  # type: ignore[arg-type]\n            t = hash_klass(len(values))","sourceCodeStart":1734,"sourceCodeEnd":1770,"githubUrl":"https://github.com/pandas-dev/pandas/blob/3b7651241d4da534b3559b60ef128e1c34f54116/pandas/core/algorithms.py#L1734-L1770","documentation":"safe_sort() requires that the values array contains no duplicates when a codes argument is provided — otherwise the code remapping is ambiguous (which duplicate should a code point to?). This specific check is in the counting-sort fast path, used for integer values spanning a modest range. The counting pass detects duplicates by comparing the cumulative count of unique values against the total length. If they differ, duplicates exist and a ValueError is raised.","triggerScenarios":"Calling safe_sort with integer values that contain duplicates and a non-None codes array, without setting assume_unique=True. Passing a factorized array where the unique values were not actually deduplicated before calling safe_sort.","commonSituations":"Using safe_sort in categorical reordering logic where the categories array accidentally contains a repeated value. Chaining operations where a previous step was expected to unique-ify the values but did not. Passing the wrong array (e.g., the full data instead of unique categories) as values.","solutions":["Deduplicate values before calling safe_sort: unique_values = np.unique(values); safe_sort(unique_values, codes=codes).","If you are certain values are unique and want to skip the check, pass assume_unique=True.","Verify that the values array you are passing is actually the set of unique categories, not the raw data."],"exampleFix":"# before\nsafe_sort(np.array([1, 2, 2, 3]), codes=np.array([0, 1, 2]))\n\n# after\nunique_vals = np.unique(np.array([1, 2, 2, 3]))  # [1, 2, 3]\nsafe_sort(unique_vals, codes=np.array([0, 1, 2]))","handlingStrategy":"validation","validationCode":"def safe_safe_sort(values, codes=None, assume_unique=False, **kwargs):\n    if codes is not None and not assume_unique:\n        arr = np.asarray(values)\n        if arr.dtype.kind in 'iu':\n            if len(np.unique(arr)) != len(arr):\n                raise ValueError(\"values must be unique when codes is not None\")\n    return pd.core.algorithms.safe_sort(values, codes=codes, assume_unique=assume_unique, **kwargs)","typeGuard":null,"tryCatchPattern":"try:\n    result = safe_sort(values, codes)\nexcept ValueError as e:\n    if \"values should be unique\" in str(e):\n        values = np.unique(values)\n        result = safe_sort(values, codes, assume_unique=True)\n    else:\n        raise","preventionTips":["Deduplicate values with np.unique() before calling safe_sort with codes.","Pass assume_unique=True only when you have verified uniqueness yourself.","Verify that the values array is the set of categories, not the raw data."],"tags":["pandas","safe-sort","uniqueness","internal-api","valueerror"],"backgroundTag":null,"analyzedSha":"3b7651241d4da534b3559b60ef128e1c34f54116","analyzedAt":"2026-08-11T22:10:44.015Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}