pandas-dev/pandas · error · ValueError
StringArray requires a sequence of strings or pandas.NA. Got
Error message
StringArray requires a sequence of strings or pandas.NA. Got '{self._ndarray.dtype}' dtype instead. What it means
Thrown by StringArray._validate in pandas/core/arrays/string_.py:727 when na_value is pandas.NA and the backing ndarray's dtype is not 'object'. StringArray stores strings as Python objects in an object-dtype ndarray; a numeric/structured dtype means the data was not prepared correctly, so construction is rejected before any string operation runs.
Solutions
- Use pd.array(numeric_arr, dtype='string') which handles conversion through _from_sequence.
- Convert to object dtype first: pd.arrays.StringArray(numeric_arr.astype(object)) — but note non-string objects still trip error 452.
- Cast numerics to strings explicitly: numeric_arr.astype(str).astype(object).
Example fix
// before import pandas as pd, numpy as np pd.arrays.StringArray(np.array([1,2,3])) # raises ValueError // after pd.array([1,2,3], dtype='string')
Defensive patterns
Strategy: validation
Validate before calling
import numpy as np
def ensure_object_ndarray(values):
arr = np.asarray(values)
if arr.dtype != object:
arr = arr.astype(object)
return arr Type guard
import numpy as np
def is_object_ndarray(arr) -> bool:
return getattr(arr, 'dtype', None) == object Try / catch
null
Prevention
- Convert numeric ndarrays to strings via .astype(str) before handing to StringArray.
- Use pd.array(..., dtype='string') which handles dtype conversion internally.
- Do not bypass _from_sequence with the raw constructor unless data is already object-dtype strings.
When it happens
Trigger: Constructing pd.arrays.StringArray(np.array([1,2,3])) (int dtype ndarray) directly. Passing a float64 ndarray from a numeric column to the StringArray constructor without object conversion. Internal code that hands a non-object ndarray to StringArray.
Common situations: User assumes StringArray casts numeric arrays the way pd.Series(..., dtype='string') does — the low-level constructor does not. Reusing a numeric buffer for a string column without an intermediate astype(object).
Related errors
- StringArray requires a sequence of strings or NaN. Got
- Cannot pass both a timezone-aware dtype and tz=None
- Cannot perform reduction
- cannot supply both a tz and a dtype with a tz
- cannot supply both a tz and a timezone-naive dtype (i.e…
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/3ec8e507a5945e5a.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/string_.py:727
self._validate(dtype)
NDArrayBacked.__init__(
self,
self._ndarray,
dtype,
)
def _validate(self, dtype: StringDtype) -> None:
"""Validate that we only store NA or strings."""
if dtype._na_value is libmissing.NA:
if len(self._ndarray) and not lib.is_string_array(
self._ndarray, skipna=True
):
raise ValueError(
"StringArray requires a sequence of strings or pandas.NA"
)
if self._ndarray.dtype != "object":
raise ValueError(
"StringArray requires a sequence of strings or pandas.NA. Got "
f"'{self._ndarray.dtype}' dtype instead."
)
# Check to see if need to convert Na values to pd.NA
if self._ndarray.ndim > 2:
# Ravel if ndims > 2 b/c no cythonized version available
lib.convert_nans_to_NA(self._ndarray.ravel("K"))
else:
lib.convert_nans_to_NA(self._ndarray)
else:
# Validate that we only store NaN or strings.
if len(self._ndarray) and not lib.is_string_array(
self._ndarray, skipna=True
):
raise ValueError("StringArray requires a sequence of strings or NaN")
if self._ndarray.dtype != "object":
raise ValueError(
"StringArray requires a sequence of strings "View on GitHub (pinned to 3b7651241d)