pandas-dev/pandas · error · ValueError

StringArray requires a sequence of strings or NaN. Got

Error message

StringArray requires a sequence of strings or NaN. Got '{self._ndarray.dtype}' dtype instead.

What it means

Thrown by StringArray._validate in pandas/core/arrays/string_.py:744 — the NaN-flavored counterpart of error 453. Fires when na_value is np.nan and the backing ndarray dtype is not 'object'. Note: the source string at line 746 is missing the 'f' prefix, so the placeholder '{self._ndarray.dtype}' is emitted literally rather than interpolated; this is a known minor bug in the message formatting and does not affect the exception type.

Solutions

  1. Use pd.array(numeric_arr, dtype=str) which converts via _from_sequence.
  2. Cast to object first: numeric_arr.astype(object) — but ensure contents are strings or NaN to avoid error 454.
  3. Cast numerics to strings: numeric_arr.astype(str).astype(object).

Example fix

// before
import pandas as pd, numpy as np
dtype = pd.StringDtype(storage='python', na_value=np.nan)
pd.arrays.StringArray(np.array([1,2,3]), dtype=dtype)  # raises ValueError

// after
pd.array([1,2,3], dtype=str)
Defensive patterns

Strategy: validation

Validate before calling

import numpy as np
def to_object_string_array(values):
    arr = np.asarray(values)
    if arr.dtype != object:
        arr = arr.astype(str).astype(object)
    return arr

Type guard

import numpy as np
def is_object_ndarray(arr) -> bool:
    return getattr(arr, 'dtype', None) == object

Try / catch

null

Prevention

When it happens

Trigger: Constructing pd.arrays.StringArray(np.array([1,2,3]), dtype=StringDtype(na_value=np.nan)) — numeric ndarray with the NaN-variant dtype. Triggered via the infer_string future when a non-object ndarray reaches the constructor.

Common situations: Same as error 453 but in code paths where the NaN-flavored StringDtype is selected (e.g., pd.options.future.infer_string = True and dtype=str).

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/660b5b4c4e953ab3. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_.py:744

            if self._ndarray.dtype != "object":
                raise ValueError(
                    "StringArray requires a sequence of strings or pandas.NA. Got "
                    f"'{self._ndarray.dtype}' dtype instead."
                )
            # Check to see if need to convert Na values to pd.NA
            if self._ndarray.ndim > 2:
                # Ravel if ndims > 2 b/c no cythonized version available
                lib.convert_nans_to_NA(self._ndarray.ravel("K"))
            else:
                lib.convert_nans_to_NA(self._ndarray)
        else:
            # Validate that we only store NaN or strings.
            if len(self._ndarray) and not lib.is_string_array(
                self._ndarray, skipna=True
            ):
                raise ValueError("StringArray requires a sequence of strings or NaN")
            if self._ndarray.dtype != "object":
                raise ValueError(
                    "StringArray requires a sequence of strings "
                    "or NaN. Got '{self._ndarray.dtype}' dtype instead."
                )
            # TODO validate or force NA/None to NaN

    def _validate_scalar(self, value):
        # used by NDArrayBackedExtensionIndex.insert
        if isna(value):
            return self.dtype.na_value
        elif not isinstance(value, str):
            raise TypeError(
                f"Invalid value '{value}' for dtype '{self.dtype}'. Value should be a "
                f"string or missing value, got '{type(value).__name__}' instead."
            )
        return value

    @classmethod
    def _from_sequence(

View on GitHub (pinned to 3b7651241d)