pandas-dev/pandas · error · ValueError

searchsorted requires array to be sorted, which is…

Error message

searchsorted requires array to be sorted, which is impossible with NAs present.

What it means

StringArray.searchsorted refuses to operate when NA values are present at the sorted boundaries, because searchsorted requires a strictly ordered array and NAs cannot be ordered relative to strings. pandas checks the first/last elements (or the sorter-indexed extremes) for NA and aborts before producing a meaningless insertion index.

Solutions

  1. Drop or fill NA values before calling searchsorted: s.dropna().searchsorted(value).
  2. Ensure the array is fully sorted and NA-free; verify with s.isna().any().
  3. If NAs must stay, exclude them and pass a sorter for the non-NA subset only.

Example fix

// before
s = pd.Series(['a','b',pd.NA], dtype='string')
s.searchsorted('b')  # ValueError
// after
clean = s.dropna().sort_values()
clean.searchsorted('b')
Defensive patterns

Strategy: validation

Validate before calling

def safe_searchsorted(s, value):
    if s.isna().any():
        s = s.dropna()
    if not s.is_monotonic_increasing and not s.is_monotonic_decreasing:
        s = s.sort_values().reset_index(drop=True)
    return s.searchsorted(value)

Type guard

def is_searchsortable(s) -> bool:
    return not s.isna().any() and (s.is_monotonic_increasing or s.is_monotonic_decreasing)

Try / catch

try:
    s.searchsorted(value)
except ValueError as e:
    if 'requires array to be sorted' in str(e):
        s.dropna().sort_values().searchsorted(value)
    else:
        raise

Prevention

When it happens

Trigger: Calling s.searchsorted(value) on a string Series that contains pd.NA or np.nan at the start or end; passing a sorter that places an NA at a boundary; calling searchsorted on unsorted data containing NAs.

Common situations: Running bisect-style lookups on a column that has missing values; sorting dropped NAs to the end and then calling searchsorted; using searchsorted on a partially-cleaned column without dropna first.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/f9d465f22ec115b2. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_.py:1198

        """

        # GH#65837: avoid O(n) scan; NA confined to array ends in sorted data.
        # When sorter is given, the sorted order is ndarray[sorter], so check
        # the first/last positions via sorter instead of raw ndarray positions.
        ndarray = self._ndarray
        if len(ndarray):
            if sorter is None:
                has_na = libmissing.checknull(ndarray[0]) or libmissing.checknull(
                    ndarray[-1]
                )
            else:
                has_na = libmissing.checknull(
                    ndarray[sorter[0]]
                ) or libmissing.checknull(ndarray[sorter[-1]])
        else:
            has_na = False
        if has_na:
            raise ValueError(
                "searchsorted requires array to be sorted, which is impossible "
                "with NAs present."
            )
        return super().searchsorted(value=value, side=side, sorter=sorter)

    def _cmp_method(self, other, op):
        from pandas.arrays import (
            ArrowExtensionArray,
            BooleanArray,
        )

        if (
            isinstance(other, BaseStringArray)
            and self.dtype.na_value is not libmissing.NA
            and other.dtype.na_value is libmissing.NA
        ):
            # NA has priority of NaN semantics
            return op(self.astype(other.dtype, copy=False), other)

View on GitHub (pinned to 3b7651241d)