pandas-dev/pandas · error · TypeError
Invalid value ' ' for dtype ' '. Value should be a string…
Error message
Invalid value '{value}' for dtype '{self.dtype}'. Value should be a string or missing value, got '{type(value).__name__}' instead. What it means
Thrown by StringArray._validate_scalar in pandas/core/arrays/string_.py:755 — used by NDArrayBackedExtensionIndex.insert. When a scalar being inserted is not NA/NaN and not a Python str, the method rejects it because StringArray can only hold strings plus its designated missing marker. The message names the offending value, the dtype, and the actual type received.
Solutions
- Stringify the scalar before inserting: idx.insert(loc, str(5)).
- Coerce missing sentinels explicitly: insert(loc, pd.NA) or insert(loc, np.nan) depending on the dtype's na_value.
- If mixed types are legitimate, build the index with dtype='object' instead of 'string'.
Example fix
// before idx = pd.Index(['a','b'], dtype='string') idx.insert(1, 5) # raises TypeError // after idx.insert(1, '5')
Defensive patterns
Strategy: validation
Validate before calling
import pandas as pd
def safe_insert(index_obj, loc, value):
if isinstance(getattr(index_obj, 'dtype', None), pd.StringDtype):
if not (pd.isna(value) or isinstance(value, str)):
value = str(value)
return index_obj.insert(loc, value) Type guard
def is_string_scalar(v) -> bool:
import pandas as pd
return isinstance(v, str) or pd.isna(v) Try / catch
null
Prevention
- Stringify scalars before inserting into a string-dtype Index.
- Use pd.NA (or np.nan for the NaN variant) when the intent is a missing entry.
- When appending indices of different dtypes, normalize types upstream.
When it happens
Trigger: Calling .insert(loc, 5) on an Index backed by StringArray. Internally triggered by Index.insert / Index.append when concatenating indices that introduce a non-string scalar. Building a string index element-by-element with mixed types.
Common situations: Appending a numeric index level to a string index. insert() called by alignment/reindex machinery that picks up a stray int. Concatenating two differently-typed indices where pandas picks the string dtype but the other side contributes non-strings.
Related errors
- Invalid value for dtype 'str'. Value should be a string or…
- can only convert an array of size 1 to a Python scalar
- can only insert Interval objects and NA into an…
- Cannot change data-type for string array.
- Cannot modify read-only array
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/39acc5d84f7c073d.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/string_.py:755
else:
# Validate that we only store NaN or strings.
if len(self._ndarray) and not lib.is_string_array(
self._ndarray, skipna=True
):
raise ValueError("StringArray requires a sequence of strings or NaN")
if self._ndarray.dtype != "object":
raise ValueError(
"StringArray requires a sequence of strings "
"or NaN. Got '{self._ndarray.dtype}' dtype instead."
)
# TODO validate or force NA/None to NaN
def _validate_scalar(self, value):
# used by NDArrayBackedExtensionIndex.insert
if isna(value):
return self.dtype.na_value
elif not isinstance(value, str):
raise TypeError(
f"Invalid value '{value}' for dtype '{self.dtype}'. Value should be a "
f"string or missing value, got '{type(value).__name__}' instead."
)
return value
@classmethod
def _from_sequence(
cls, scalars, *, dtype: Dtype | None = None, copy: bool = False
) -> Self:
if dtype and not (isinstance(dtype, str) and dtype == "string"):
dtype = pandas_dtype(dtype)
assert isinstance(dtype, StringDtype) and dtype.storage == "python"
elif using_string_dtype():
dtype = StringDtype(storage="python", na_value=np.nan)
else:
dtype = StringDtype(storage="python")
from pandas.core.arrays.masked import BaseMaskedArrayView on GitHub (pinned to 3b7651241d)