pandas-dev/pandas · error · TypeError

Invalid value ' ' for dtype 'str'. Value should be a string…

Error message

Invalid value '{item}' for dtype 'str'. Value should be a string or missing value, got '{type(item).__name__}' instead.

What it means

ArrowStringArray.insert rejects any item that is neither a Python str nor libmissing.NA. Because the dtype is strictly str, inserting integers, floats, lists, or None-as-nothing raises TypeError to prevent silent coercion that would corrupt the typed array.

Solutions

  1. Convert the item to str before inserting: arr.insert(loc, str(value)).
  2. Use libmissing.NA (or pd.NA) for missing values; if na_value is np.nan, np.nan is accepted and converted to NA.
  3. Filter out or route non-string items to a different column.

Example fix

// before
arr.insert(0, 42)  # TypeError
// after
arr.insert(0, str(42))
# missing
arr.insert(0, pd.NA)
Defensive patterns

Strategy: type-guard

Validate before calling

def safe_insert(arr, loc, item):
    if not isinstance(item, str) and item is not pd.NA:
        item = str(item)
    return arr.insert(loc, item)

Type guard

def is_insertable_string_item(item) -> bool:
    return isinstance(item, str) or item is pd.NA

Try / catch

try:
    arr.insert(loc, item)
except TypeError as e:
    if "Invalid value" in str(e):
        arr.insert(loc, str(item))
    else:
        raise

Prevention

When it happens

Trigger: arr.insert(loc, 5); arr.insert(loc, 3.14); arr.insert(loc, None) when na_value is not np.nan; arr.insert(loc, ['a']) on an ArrowStringArray.

Common situations: Inserting a value from an untyped source (JSON, CSV cell) into a string column without converting; mixing dtypes in a loop that inserts heterogeneous rows.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/099ff22eaa53b4fb. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_arrow.py:315

    @classmethod
    def _from_sequence_of_strings(
        cls, strings, *, dtype: ExtensionDtype, copy: bool = False
    ) -> Self:
        return cls._from_sequence(strings, dtype=dtype, copy=copy)

    @property
    def dtype(self) -> StringDtype:  # type: ignore[override]
        """
        An instance of 'string[pyarrow]'.
        """
        return self._dtype

    def insert(self, loc: int, item) -> ArrowStringArray:
        if self.dtype.na_value is np.nan and item is np.nan:
            item = libmissing.NA
        if not isinstance(item, str) and item is not libmissing.NA:
            raise TypeError(
                f"Invalid value '{item}' for dtype 'str'. Value should be a "
                f"string or missing value, got '{type(item).__name__}' instead."
            )
        return super().insert(loc, item)

    def _convert_bool_result(self, values, na=lib.no_default, method_name=None):
        validate_na_arg(na, name="na")
        if self.dtype.na_value is np.nan:
            if na is lib.no_default or isna(na):
                # NaN propagates as False
                values = values.fill_null(False)
            else:
                values = values.fill_null(na)
            return values.to_numpy()
        elif na is not lib.no_default and not isna(na):  # pyright: ignore [reportGeneralTypeIssues]
            values = values.fill_null(na)
        return BooleanDtype().__from_arrow__(values)

View on GitHub (pinned to 3b7651241d)