{"record":{"id":"f9d465f22ec115b2","repo":"pandas-dev/pandas","slug":"searchsorted-requires-array-to-be-sorted-which-is-f9d465","errorCode":null,"errorMessage":"searchsorted requires array to be sorted, which is impossible with NAs present.","messagePattern":"searchsorted requires array to be sorted, which is impossible with NAs present\\.","errorType":"exception","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"pandas/core/arrays/string_.py","lineNumber":1198,"sourceCode":"        \"\"\"\n\n        # GH#65837: avoid O(n) scan; NA confined to array ends in sorted data.\n        # When sorter is given, the sorted order is ndarray[sorter], so check\n        # the first/last positions via sorter instead of raw ndarray positions.\n        ndarray = self._ndarray\n        if len(ndarray):\n            if sorter is None:\n                has_na = libmissing.checknull(ndarray[0]) or libmissing.checknull(\n                    ndarray[-1]\n                )\n            else:\n                has_na = libmissing.checknull(\n                    ndarray[sorter[0]]\n                ) or libmissing.checknull(ndarray[sorter[-1]])\n        else:\n            has_na = False\n        if has_na:\n            raise ValueError(\n                \"searchsorted requires array to be sorted, which is impossible \"\n                \"with NAs present.\"\n            )\n        return super().searchsorted(value=value, side=side, sorter=sorter)\n\n    def _cmp_method(self, other, op):\n        from pandas.arrays import (\n            ArrowExtensionArray,\n            BooleanArray,\n        )\n\n        if (\n            isinstance(other, BaseStringArray)\n            and self.dtype.na_value is not libmissing.NA\n            and other.dtype.na_value is libmissing.NA\n        ):\n            # NA has priority of NaN semantics\n            return op(self.astype(other.dtype, copy=False), other)","sourceCodeStart":1180,"sourceCodeEnd":1216,"githubUrl":"https://github.com/pandas-dev/pandas/blob/3b7651241d4da534b3559b60ef128e1c34f54116/pandas/core/arrays/string_.py#L1180-L1216","documentation":"StringArray.searchsorted refuses to operate when NA values are present at the sorted boundaries, because searchsorted requires a strictly ordered array and NAs cannot be ordered relative to strings. pandas checks the first/last elements (or the sorter-indexed extremes) for NA and aborts before producing a meaningless insertion index.","triggerScenarios":"Calling s.searchsorted(value) on a string Series that contains pd.NA or np.nan at the start or end; passing a sorter that places an NA at a boundary; calling searchsorted on unsorted data containing NAs.","commonSituations":"Running bisect-style lookups on a column that has missing values; sorting dropped NAs to the end and then calling searchsorted; using searchsorted on a partially-cleaned column without dropna first.","solutions":["Drop or fill NA values before calling searchsorted: s.dropna().searchsorted(value).","Ensure the array is fully sorted and NA-free; verify with s.isna().any().","If NAs must stay, exclude them and pass a sorter for the non-NA subset only."],"exampleFix":"// before\ns = pd.Series(['a','b',pd.NA], dtype='string')\ns.searchsorted('b')  # ValueError\n// after\nclean = s.dropna().sort_values()\nclean.searchsorted('b')","handlingStrategy":"validation","validationCode":"def safe_searchsorted(s, value):\n    if s.isna().any():\n        s = s.dropna()\n    if not s.is_monotonic_increasing and not s.is_monotonic_decreasing:\n        s = s.sort_values().reset_index(drop=True)\n    return s.searchsorted(value)","typeGuard":"def is_searchsortable(s) -> bool:\n    return not s.isna().any() and (s.is_monotonic_increasing or s.is_monotonic_decreasing)","tryCatchPattern":"try:\n    s.searchsorted(value)\nexcept ValueError as e:\n    if 'requires array to be sorted' in str(e):\n        s.dropna().sort_values().searchsorted(value)\n    else:\n        raise","preventionTips":["Drop NAs before searchsorted.","Sort the Series first and confirm monotonicity.","Pass a sorter for pre-sorted non-NA subsets."],"tags":["pandas","string-array","searchsorted","missing-values","sorting"],"backgroundTag":null,"analyzedSha":"3b7651241d4da534b3559b60ef128e1c34f54116","analyzedAt":"2026-08-11T22:10:44.015Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}