pandas-dev/pandas · error · TypeError

Invalid value for dtype 'str'. Value should be a string or…

Error message

Invalid value for dtype 'str'. Value should be a string or missing value (or array of those).

What it means

ArrowStringArray._validate_setitem_value rejects array-likes whose contents are not all strings (or NAs). After converting the value to a numpy object array, it checks lib.is_string_array(value, skipna=True); any non-string non-NA element triggers TypeError.

Solutions

  1. Coerce each element to str: s.iloc[:] = [str(v) for v in values].
  2. Use pd.NA in place of None/NaN for missing entries.
  3. Validate the source array with pd.api.types.is_string_dtype before assignment.

Example fix

// before
s.iloc[:] = [1, 2, 3]  # TypeError
// after
s.iloc[:] = [str(v) for v in [1, 2, 3]]
Defensive patterns

Strategy: validation

Validate before calling

def safe_setitem_array(arr, key, values):
    import numpy as np
    values = np.asarray(values, dtype=object)
    if len(values) and not all(isinstance(v, str) or v is pd.NA for v in values):
        values = np.array([str(v) if not (isinstance(v, str) or v is pd.NA) else v for v in values], dtype=object)
    arr[key] = values

Type guard

def is_string_or_na_array(values) -> bool:
    import numpy as np
    arr = np.asarray(values, dtype=object)
    return all(isinstance(v, str) or v is pd.NA for v in arr)

Try / catch

try:
    s.iloc[:] = values
except TypeError as e:
    if "Invalid value for dtype 'str'" in str(e):
        s.iloc[:] = [str(v) for v in values]
    else:
        raise

Prevention

When it happens

Trigger: s.iloc[:] = [1, 'a', 'b'] on a string[pyarrow] Series; s.iloc[:] = np.array([1.0, 2.0]) (non-string ndarray); assigning a list with mixed types.

Common situations: Assigning a list comprehension that yields mixed types; vectorized assignment from another column with the wrong dtype; bulk replacement with computed non-string values.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/4554fdfe59597734. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_arrow.py:354

        if is_scalar(value):
            if isna(value):
                value = None
            elif not isinstance(value, str):
                raise TypeError(
                    f"Invalid value '{value}' for dtype 'str'. Value should be a "
                    f"string or missing value, got '{type(value).__name__}' instead."
                )
        elif isinstance(value, type(self)):
            pass
        else:
            if not is_array_like_deprecate_non_pandas(value):
                value = np.asarray(value, dtype=object)
            else:
                value = np.asarray(value)
            if len(value) and not (
                value.ndim == 1 and lib.is_string_array(value, skipna=True)
            ):
                raise TypeError(
                    "Invalid value for dtype 'str'. Value should be a "
                    "string or missing value (or array of those)."
                )
        return super()._validate_setitem_value(value)

    def isin(self, values: ArrayLike) -> npt.NDArray[np.bool_]:
        value_set = [
            pa_scalar.as_py()
            for pa_scalar in [pa.scalar(value, from_pandas=True) for value in values]
            if pa_scalar.type in (pa.string(), pa.null(), pa.large_string())
        ]

        # short-circuit to return all False array.
        if not value_set:
            return np.zeros(len(self), dtype=bool)

        result = pc.is_in(
            self._pa_array, value_set=pa.array(value_set, type=self._pa_array.type)

View on GitHub (pinned to 3b7651241d)