pandas-dev/pandas · error · ValueError

'na_value' must be np.nan or pd.NA, got

Error message

'na_value' must be np.nan or pd.NA, got {na_value}

What it means

Thrown by StringDtype.__init__ in pandas/core/arrays/string_.py:226 when na_value is neither a NaN float nor pandas.NA. StringDtype only supports these two sentinel missing values because its downstream comparisons and is_string_array checks key on identity (na_value is np.nan or na_value is libmissing.NA).

Solutions

  1. Use pd.NA (default) for the modern nullable string type: pd.StringDtype() or pd.StringDtype(na_value=pd.NA).
  2. Use np.nan for the NaN-flavored variant: pd.StringDtype(na_value=np.nan).
  3. If a custom sentinel is truly needed, use a Categorical or object dtype instead.

Example fix

// before
pd.StringDtype(na_value='')  # raises ValueError

// after
pd.StringDtype(na_value=pd.NA)
Defensive patterns

Strategy: validation

Validate before calling

def safe_na_value(v):
    import numpy as np, pandas as pd
    if v is pd.NA or (isinstance(v, float) and np.isnan(v)):
        return v
    raise ValueError('na_value must be pd.NA or np.nan')

Type guard

import numpy as np, pandas as pd
def valid_na_value(v) -> bool:
    return v is pd.NA or (isinstance(v, float) and np.isnan(v))

Try / catch

null

Prevention

When it happens

Trigger: Calling pd.StringDtype(na_value=0), na_value='', na_value=None, or na_value='NA' (string). Constructing a StringDtype with a custom missing-value sentinel.

Common situations: User assumes any falsy value can represent missingness. Passing None expecting it to be coerced to NA (None is not is_nan and is not libmissing.NA). Migration from object dtype where None was the missing marker.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/8c1e35cfa6d9debf. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/string_.py:226

                    storage = "python"

        # validate options
        if storage not in {"python", "pyarrow"}:
            raise ValueError(
                f"Storage must be 'python' or 'pyarrow'. Got {storage} instead."
            )
        if storage == "pyarrow" and not HAS_PYARROW:
            raise ImportError(
                f"pyarrow>={PYARROW_MIN_VERSION} is required for PyArrow "
                "backed StringArray."
            )

        if isinstance(na_value, float) and np.isnan(na_value):
            # when passed a NaN value, always set to np.nan to ensure we use
            # a consistent NaN value (and we can use `dtype.na_value is np.nan`)
            na_value = np.nan
        elif na_value is not libmissing.NA:
            raise ValueError(f"'na_value' must be np.nan or pd.NA, got {na_value}")

        self._storage = cast("str", storage)
        self._na_value = na_value

    def __repr__(self) -> str:
        storage = "" if self.storage == "pyarrow" else "storage='python', "
        return f"<StringDtype({storage}na_value={self._na_value})>"

    def __eq__(self, other: object) -> bool:
        # we need to override the base class __eq__ because na_value (NA or NaN)
        # cannot be checked with normal `==`
        if isinstance(other, str):
            # TODO should dtype == "string" work for the NaN variant?
            if other == "string" or other == self.name:  # noqa: PLR1714 (repeated-equality-comparison)
                return True
            try:
                other = self.construct_from_string(other)
            except (TypeError, ImportError):

View on GitHub (pinned to 3b7651241d)