pandas-dev/pandas · error · ValueError

StringArray requires a sequence of strings or pandas.NA

Error message

StringArray requires a sequence of strings or pandas.NA

What it means

Thrown by StringArray._validate in pandas/core/arrays/string_.py:723 during construction when na_value is pandas.NA and the supplied object ndarray contains at least one value that is neither a string nor a missing sentinel (as detected by lib.is_string_array with skipna=True). StringArray (python storage) requires pure-string content; numeric or mixed content is rejected at the boundary.

Solutions

  1. Use pd.array(values, dtype='string') or pd.Series(values, dtype='string') — these coerce non-string scalars to str before validation.
  2. Pre-stringify the input: np.array([str(x) for x in values], dtype=object).
  3. If non-string data is legitimate, use dtype='object' or a categorical instead of 'string'.

Example fix

// before
import pandas as pd, numpy as np
pd.arrays.StringArray(np.array([1, 'a'], dtype=object))  # raises ValueError

// after
pd.array([1, 'a'], dtype='string')  # coerces 1 -> '1'
Defensive patterns

Strategy: validation

Validate before calling

import numpy as np, pandas as pd
def build_string_array(values):
    # _from_sequence coerces; direct constructor does not
    return pd.array(values, dtype='string')

Type guard

import pandas as pd, numpy as np
def all_strings_or_na(arr: np.ndarray) -> bool:
    import pandas._libs.lib as lib
    return lib.is_string_array(arr, skipna=True)

Try / catch

null

Prevention

When it happens

Trigger: Directly constructing pd.arrays.StringArray(np.array([1, 'a', None], dtype=object)) or passing ints/floats. Calling pd.Series([1,2], dtype='string') normally routes through _from_sequence which coerces, but bypassing it via the constructor hits _validate. Inserting via internals that rebuild from a raw ndarray.

Common situations: User reaches for the low-level StringArray constructor instead of pd.array(values, dtype='string') which auto-stringifies. Mixed-type object arrays from CSV reading fed directly. Migration code that built object arrays previously and now targets StringArray.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/7c0382811fd14c45. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_.py:723

        values = extract_array(values)

        super().__init__(values, copy=copy)
        if not isinstance(values, type(self)):
            self._validate(dtype)
        NDArrayBacked.__init__(
            self,
            self._ndarray,
            dtype,
        )

    def _validate(self, dtype: StringDtype) -> None:
        """Validate that we only store NA or strings."""

        if dtype._na_value is libmissing.NA:
            if len(self._ndarray) and not lib.is_string_array(
                self._ndarray, skipna=True
            ):
                raise ValueError(
                    "StringArray requires a sequence of strings or pandas.NA"
                )
            if self._ndarray.dtype != "object":
                raise ValueError(
                    "StringArray requires a sequence of strings or pandas.NA. Got "
                    f"'{self._ndarray.dtype}' dtype instead."
                )
            # Check to see if need to convert Na values to pd.NA
            if self._ndarray.ndim > 2:
                # Ravel if ndims > 2 b/c no cythonized version available
                lib.convert_nans_to_NA(self._ndarray.ravel("K"))
            else:
                lib.convert_nans_to_NA(self._ndarray)
        else:
            # Validate that we only store NaN or strings.
            if len(self._ndarray) and not lib.is_string_array(
                self._ndarray, skipna=True
            ):

View on GitHub (pinned to 3b7651241d)