pandas-dev/pandas · error · ValueError

StringArray requires a sequence of strings or NaN

Error message

StringArray requires a sequence of strings or NaN

What it means

Thrown by StringArray._validate in pandas/core/arrays/string_.py:742 — the NaN-flavored counterpart of error 452. Fires when na_value is np.nan (rather than pandas.NA) and the object ndarray contains values that are neither strings nor missing sentinels. The two code paths exist because StringDtype supports both NA and NaN missing markers, each with its own validation branch.

Solutions

  1. Coerce values to strings upstream: np.array([str(x) for x in values], dtype=object).
  2. Use pd.array(values, dtype=str) or dtype='string[python]' which route through _from_sequence and stringify.
  3. Drop non-string entries or replace them with np.nan before construction.

Example fix

// before
import pandas as pd, numpy as np
dtype = pd.StringDtype(storage='python', na_value=np.nan)
pd.arrays.StringArray(np.array([1, 'a'], dtype=object), dtype=dtype)  # raises

// after
pd.array([1, 'a'], dtype=str)
Defensive patterns

Strategy: validation

Validate before calling

import pandas as pd
def build_nan_variant_string(values):
    # _from_sequence coerces numerics to strings; direct constructor does not
    dtype = pd.StringDtype(storage='python', na_value=np.nan)
    return pd.array(values, dtype=dtype)

Type guard

import pandas._libs.lib as lib
def all_strings_or_nan(arr) -> bool:
    return lib.is_string_array(arr, skipna=True)

Try / catch

null

Prevention

When it happens

Trigger: Constructing StringArray with dtype=StringDtype(na_value=np.nan) and an ndarray containing ints/bools/lists. Triggered via pd.Series(..., dtype=str) when the future infer_string option is enabled (which selects the NaN variant).

Common situations: Enabling pd.options.future.infer_string = True changes 'str' dtype to StringDtype(na_value=np.nan), then mixed-type inputs surface here. Direct construction with the NaN variant dtype.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/7418c00975b63c5f. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_.py:742

                    "StringArray requires a sequence of strings or pandas.NA"
                )
            if self._ndarray.dtype != "object":
                raise ValueError(
                    "StringArray requires a sequence of strings or pandas.NA. Got "
                    f"'{self._ndarray.dtype}' dtype instead."
                )
            # Check to see if need to convert Na values to pd.NA
            if self._ndarray.ndim > 2:
                # Ravel if ndims > 2 b/c no cythonized version available
                lib.convert_nans_to_NA(self._ndarray.ravel("K"))
            else:
                lib.convert_nans_to_NA(self._ndarray)
        else:
            # Validate that we only store NaN or strings.
            if len(self._ndarray) and not lib.is_string_array(
                self._ndarray, skipna=True
            ):
                raise ValueError("StringArray requires a sequence of strings or NaN")
            if self._ndarray.dtype != "object":
                raise ValueError(
                    "StringArray requires a sequence of strings "
                    "or NaN. Got '{self._ndarray.dtype}' dtype instead."
                )
            # TODO validate or force NA/None to NaN

    def _validate_scalar(self, value):
        # used by NDArrayBackedExtensionIndex.insert
        if isna(value):
            return self.dtype.na_value
        elif not isinstance(value, str):
            raise TypeError(
                f"Invalid value '{value}' for dtype '{self.dtype}'. Value should be a "
                f"string or missing value, got '{type(value).__name__}' instead."
            )
        return value

View on GitHub (pinned to 3b7651241d)