pandas-dev/pandas · error · ValueError

cannot convert float NaN to bool

Error message

cannot convert float NaN to bool

What it means

Raised by BaseMaskedArray.astype when the target dtype is boolean kind ('b') and the masked array has missing values. numpy's astype_nansafe converts np.nan to True, which would silently corrupt results, so pandas raises up front to surface the ambiguity.

Solutions

  1. Cast to nullable boolean dtype: arr.astype('boolean') to preserve NA.
  2. Decide on a sentinel first: arr.fillna(False).astype('bool') or arr.dropna().astype('bool').
  3. Use pd.isna(arr) to build an explicit mask instead of forcing bool.

Example fix

// before
arr = pd.array([True, None, False], dtype='boolean')
arr.astype('bool')   # raises
// after
arr.fillna(False).astype('bool')
Defensive patterns

Strategy: validation

Validate before calling

if dtype.kind == 'b' and arr._hasna:
    raise ValueError('Refusing bool cast with NA; fill or drop first')
out = arr.astype(dtype)

Type guard

def can_astype_bool(arr) -> bool:
    return not arr._hasna

Try / catch

try:
    out = arr.astype('bool')
except ValueError as e:
    if 'NaN to bool' in str(e):
        out = arr.fillna(False).astype('bool')
    else:
        raise

Prevention

When it happens

Trigger: Calling arr.astype('bool') or arr.astype(np.bool_) on any masked ExtensionArray with self._hasna True (nullable boolean, Int, or Float with NA).

Common situations: Converting a nullable boolean column to numpy bool for ML features or conditional masks; coercing nullable Int/Float flags to bool after a groupby; pipelines that assume a clean boolean column.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/4caf0ffe893e5593. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/masked.py:808

        na_value: float | np.datetime64 | lib.NoDefault

        # coerce
        if dtype.kind == "f":
            # In astype, we consider dtype=float to also mean na_value=np.nan
            na_value = np.nan
        elif dtype.kind == "M":
            unit = np.datetime_data(dtype)[0]
            na_value = np.datetime64("NaT", unit)  # type: ignore[call-overload]
        else:
            na_value = lib.no_default

        # to_numpy will also raise, but we get somewhat nicer exception messages here
        if dtype.kind in "iu" and self._hasna:
            raise ValueError("cannot convert NA to integer")
        if dtype.kind == "b" and self._hasna:
            # careful: astype_nansafe converts np.nan to True
            raise ValueError("cannot convert float NaN to bool")

        data = self.to_numpy(dtype=dtype, na_value=na_value, copy=copy)
        return data

    __array_priority__ = 1000  # higher than ndarray so ops dispatch to us

    def __array__(
        self, dtype: NpDtype | None = None, copy: bool | None = None
    ) -> np.ndarray:
        """
        the array interface, return my values
        We return an object array here to preserve our scalar values
        """
        if copy is False:
            if not self._hasna:
                # special case, here we can simply return the underlying data
                result = np.array(self._data, dtype=dtype, copy=copy)
                # If the ExtensionArray is readonly, make the numpy array readonly too

View on GitHub (pinned to 3b7651241d)