pandas-dev/pandas · error · TypeError

expected a string object, not

Error message

expected a string object, not {type(pat).__name__}

What it means

Raised by ArrowExtensionArray._str_rsplit when `pat` is supplied but is not a `str`. The arrow split path is dispatched to pyarrow.compute.split_pattern, which requires a literal string pattern; passing bytes, a list, None-as-other-type, or any non-string object is rejected up front with a TypeError naming the offending type.

Solutions

  1. Pass a literal str separator: decode bytes with `.decode()` first.
  2. If you intended a regex/compiled-pattern split, note rsplit on the arrow path only supports literal patterns — supply the pattern as a str (still treated literally here).

Example fix

# before
ser.str.rsplit(pat=b",")
# after
ser.str.rsplit(pat=",")
Defensive patterns

Strategy: type-guard

Validate before calling

if pat is not None and not isinstance(pat, str):
    raise TypeError(f"pat must be str, got {type(pat).__name__}")
ser.str.rsplit(pat=pat)

Type guard

def is_str_or_none(pat) -> bool:
    return pat is None or isinstance(pat, str)

Try / catch

try:
    out = ser.str.rsplit(pat=pat)
except TypeError as e:
    if "expected a string object" in str(e):
        out = ser.str.rsplit(pat=str(pat))
    else:
        raise

Prevention

When it happens

Trigger: Calling `ser.str.rsplit(pat=b",")` (bytes), `pat=44` (int), or any non-string separator on a string[pyarrow] Series. Note `pat=None` (whitespace split) is allowed; the check only fires for non-None non-str values.

Common situations: Reusing a separator constant declared as bytes; data flowing in from a binary source; passing a compiled pattern object where a literal string is expected.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/d71f4e8cd2dabb0d. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/arrow/array.py:3777

        if n in {-1, 0}:
            n = None
        if pat is None:
            split_func = pc.utf8_split_whitespace
        elif regex is True:
            split_func = functools.partial(pc.split_pattern_regex, pattern=pat)
        elif regex is False:
            split_func = functools.partial(pc.split_pattern, pattern=pat)
        # GH#58321: regex is None — infer: single-char literal, multi-char regex
        elif len(pat) == 1:
            split_func = functools.partial(pc.split_pattern, pattern=pat)
        else:
            split_func = functools.partial(pc.split_pattern_regex, pattern=pat)
        return self._from_pyarrow_array(split_func(self._pa_array, max_splits=n))

    def _str_rsplit(self, pat: str | None = None, n: int | None = -1) -> Self:
        if pat is not None and not isinstance(pat, str):
            msg = f"expected a string object, not {type(pat).__name__}"
            raise TypeError(msg)
        if n in {-1, 0}:
            n = None
        if pat is None:
            return self._from_pyarrow_array(
                pc.utf8_split_whitespace(self._pa_array, max_splits=n, reverse=True)
            )
        return self._from_pyarrow_array(
            pc.split_pattern(self._pa_array, pat, max_splits=n, reverse=True)
        )

    def _str_translate(self, table: dict[int, str]) -> Self:
        predicate = lambda val: val.translate(table)
        result = self._apply_elementwise(predicate)
        return self._from_pyarrow_array(pa.chunked_array(result))

    def _str_wrap(self, width: int, **kwargs) -> Self:
        kwargs["width"] = width
        tw = textwrap.TextWrapper(**kwargs)

View on GitHub (pinned to 3b7651241d)