{"record":{"id":"7418c00975b63c5f","repo":"pandas-dev/pandas","slug":"stringarray-requires-a-sequence-of-strings-or-nan","errorCode":null,"errorMessage":"StringArray requires a sequence of strings or NaN","messagePattern":"StringArray requires a sequence of strings or NaN","errorType":"exception","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"pandas/core/arrays/string_.py","lineNumber":742,"sourceCode":"                    \"StringArray requires a sequence of strings or pandas.NA\"\n                )\n            if self._ndarray.dtype != \"object\":\n                raise ValueError(\n                    \"StringArray requires a sequence of strings or pandas.NA. Got \"\n                    f\"'{self._ndarray.dtype}' dtype instead.\"\n                )\n            # Check to see if need to convert Na values to pd.NA\n            if self._ndarray.ndim > 2:\n                # Ravel if ndims > 2 b/c no cythonized version available\n                lib.convert_nans_to_NA(self._ndarray.ravel(\"K\"))\n            else:\n                lib.convert_nans_to_NA(self._ndarray)\n        else:\n            # Validate that we only store NaN or strings.\n            if len(self._ndarray) and not lib.is_string_array(\n                self._ndarray, skipna=True\n            ):\n                raise ValueError(\"StringArray requires a sequence of strings or NaN\")\n            if self._ndarray.dtype != \"object\":\n                raise ValueError(\n                    \"StringArray requires a sequence of strings \"\n                    \"or NaN. Got '{self._ndarray.dtype}' dtype instead.\"\n                )\n            # TODO validate or force NA/None to NaN\n\n    def _validate_scalar(self, value):\n        # used by NDArrayBackedExtensionIndex.insert\n        if isna(value):\n            return self.dtype.na_value\n        elif not isinstance(value, str):\n            raise TypeError(\n                f\"Invalid value '{value}' for dtype '{self.dtype}'. Value should be a \"\n                f\"string or missing value, got '{type(value).__name__}' instead.\"\n            )\n        return value\n","sourceCodeStart":724,"sourceCodeEnd":760,"githubUrl":"https://github.com/pandas-dev/pandas/blob/3b7651241d4da534b3559b60ef128e1c34f54116/pandas/core/arrays/string_.py#L724-L760","documentation":"Thrown by StringArray._validate in pandas/core/arrays/string_.py:742 — the NaN-flavored counterpart of error 452. Fires when na_value is np.nan (rather than pandas.NA) and the object ndarray contains values that are neither strings nor missing sentinels. The two code paths exist because StringDtype supports both NA and NaN missing markers, each with its own validation branch.","triggerScenarios":"Constructing StringArray with dtype=StringDtype(na_value=np.nan) and an ndarray containing ints/bools/lists. Triggered via pd.Series(..., dtype=str) when the future infer_string option is enabled (which selects the NaN variant).","commonSituations":"Enabling pd.options.future.infer_string = True changes 'str' dtype to StringDtype(na_value=np.nan), then mixed-type inputs surface here. Direct construction with the NaN variant dtype.","solutions":["Coerce values to strings upstream: np.array([str(x) for x in values], dtype=object).","Use pd.array(values, dtype=str) or dtype='string[python]' which route through _from_sequence and stringify.","Drop non-string entries or replace them with np.nan before construction."],"exampleFix":"// before\nimport pandas as pd, numpy as np\ndtype = pd.StringDtype(storage='python', na_value=np.nan)\npd.arrays.StringArray(np.array([1, 'a'], dtype=object), dtype=dtype)  # raises\n\n// after\npd.array([1, 'a'], dtype=str)","handlingStrategy":"validation","validationCode":"import pandas as pd\ndef build_nan_variant_string(values):\n    # _from_sequence coerces numerics to strings; direct constructor does not\n    dtype = pd.StringDtype(storage='python', na_value=np.nan)\n    return pd.array(values, dtype=dtype)","typeGuard":"import pandas._libs.lib as lib\ndef all_strings_or_nan(arr) -> bool:\n    return lib.is_string_array(arr, skipna=True)","tryCatchPattern":"null","preventionTips":["Use pd.array(..., dtype=str) which routes through _from_sequence and coerces.","When infer_string is enabled, pre-stringify numeric columns.","Validate contents with lib.is_string_array before the raw constructor."],"tags":["string-array","validation","nan","infer-string"],"backgroundTag":null,"analyzedSha":"3b7651241d4da534b3559b60ef128e1c34f54116","analyzedAt":"2026-08-11T22:10:44.015Z","contentChangedAt":null,"schemaVersion":2},"datasetVersion":"2026-09-23T08:17:48.524Z"}