pandas-dev/pandas · error · ValueError
searchsorted requires array to be sorted, which is…
Error message
searchsorted requires array to be sorted, which is impossible with NAs present.
What it means
StringArray.searchsorted refuses to operate when NA values are present at the sorted boundaries, because searchsorted requires a strictly ordered array and NAs cannot be ordered relative to strings. pandas checks the first/last elements (or the sorter-indexed extremes) for NA and aborts before producing a meaningless insertion index.
Solutions
- Drop or fill NA values before calling searchsorted: s.dropna().searchsorted(value).
- Ensure the array is fully sorted and NA-free; verify with s.isna().any().
- If NAs must stay, exclude them and pass a sorter for the non-NA subset only.
Example fix
// before
s = pd.Series(['a','b',pd.NA], dtype='string')
s.searchsorted('b') # ValueError
// after
clean = s.dropna().sort_values()
clean.searchsorted('b') Defensive patterns
Strategy: validation
Validate before calling
def safe_searchsorted(s, value):
if s.isna().any():
s = s.dropna()
if not s.is_monotonic_increasing and not s.is_monotonic_decreasing:
s = s.sort_values().reset_index(drop=True)
return s.searchsorted(value) Type guard
def is_searchsortable(s) -> bool:
return not s.isna().any() and (s.is_monotonic_increasing or s.is_monotonic_decreasing) Try / catch
try:
s.searchsorted(value)
except ValueError as e:
if 'requires array to be sorted' in str(e):
s.dropna().sort_values().searchsorted(value)
else:
raise Prevention
- Drop NAs before searchsorted.
- Sort the Series first and confirm monotonicity.
- Pass a sorter for pre-sorted non-NA subsets.
When it happens
Trigger: Calling s.searchsorted(value) on a string Series that contains pd.NA or np.nan at the start or end; passing a sorter that places an NA at a boundary; calling searchsorted on unsorted data containing NAs.
Common situations: Running bisect-style lookups on a column that has missing values; sorting dropped NAs to the end and then calling searchsorted; using searchsorted on a partially-cleaned column without dropna first.
Related errors
- searchsorted requires array to be sorted, which is…
- Cannot multiply StringArray by bools. Explicitly cast to…
- Cannot perform reduction
- Lengths of operands do not match
- operation ' ' not supported for dtype
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/f9d465f22ec115b2.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/string_.py:1198
"""
# GH#65837: avoid O(n) scan; NA confined to array ends in sorted data.
# When sorter is given, the sorted order is ndarray[sorter], so check
# the first/last positions via sorter instead of raw ndarray positions.
ndarray = self._ndarray
if len(ndarray):
if sorter is None:
has_na = libmissing.checknull(ndarray[0]) or libmissing.checknull(
ndarray[-1]
)
else:
has_na = libmissing.checknull(
ndarray[sorter[0]]
) or libmissing.checknull(ndarray[sorter[-1]])
else:
has_na = False
if has_na:
raise ValueError(
"searchsorted requires array to be sorted, which is impossible "
"with NAs present."
)
return super().searchsorted(value=value, side=side, sorter=sorter)
def _cmp_method(self, other, op):
from pandas.arrays import (
ArrowExtensionArray,
BooleanArray,
)
if (
isinstance(other, BaseStringArray)
and self.dtype.na_value is not libmissing.NA
and other.dtype.na_value is libmissing.NA
):
# NA has priority of NaN semantics
return op(self.astype(other.dtype, copy=False), other)View on GitHub (pinned to 3b7651241d)