pandas-dev/pandas · error · ValueError

'data' must have a single column, not

Error message

'data' must have a single column, not '{ncol}'

What it means

Raised by SparseArray.from_spmatrix when the source scipy sparse matrix has more than one column. The Series/SparseArray model is 1-D, so only single-column matrices (shape (N, 1)) can be flattened into a SparseArray.

Solutions

  1. Select a single column: SparseArray.from_spmatrix(mat[:, col]).
  2. For multi-column data, use pd.DataFrame.sparse.from_spmatrix(mat) instead.
  3. Reshape to (N, 1) explicitly if you truly have one vector.

Example fix

# before
pd.arrays.SparseArray.from_spmatrix(mat)  # mat.shape == (5, 3)
# after
pd.arrays.SparseArray.from_spmatrix(mat[:, 0])
# or, for all columns
pd.DataFrame.sparse.from_spmatrix(mat)
Defensive patterns

Strategy: validation

Validate before calling

def matrix_is_single_column(mat) -> bool:
    return mat.ndim == 2 and mat.shape[1] == 1

Type guard

def single_column_sparse_matrix(mat) -> bool:
    return getattr(mat, 'shape', (0, 0))[1] == 1

Try / catch

try:
    arr = pd.arrays.SparseArray.from_spmatrix(mat)
except ValueError as e:
    if 'single column' in str(e):
        arr = pd.arrays.SparseArray.from_spmatrix(mat[:, [0]])
    else:
        raise

Prevention

When it happens

Trigger: pd.arrays.SparseArray.from_spmatrix(scipy.sparse.csr_matrix((3,2))); passing a multi-column CSC/CSR/COO matrix.

Common situations: Treating a 2-D sparse matrix as a 1-D vector; wanting one column but constructing the matrix with the wrong shape; forgetting to slice the matrix first.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/3cfb7d535856150d. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/sparse/array.py:558

            sparse matrix with a single column.

        Returns
        -------
        SparseArray

        Examples
        --------
        >>> import scipy.sparse
        >>> mat = scipy.sparse.coo_matrix((4, 1))
        >>> pd.arrays.SparseArray.from_spmatrix(mat)
        <SparseArray>
        [0.0, 0.0, 0.0, 0.0]
        Length: 4, dtype: Sparse[float64, 0.0]
        """
        length, ncol = data.shape

        if ncol != 1:
            raise ValueError(f"'data' must have a single column, not '{ncol}'")

        # our sparse index classes require that the positions be strictly
        # increasing. So we need to sort loc, and arr accordingly.
        data_csc = data.tocsc()
        data_csc.sort_indices()
        arr = data_csc.data
        idx = data_csc.indices

        zero = np.array(0, dtype=arr.dtype).item()
        dtype = SparseDtype(arr.dtype, zero)
        index = IntIndex(length, idx)

        return cls._simple_new(arr, index, dtype)

    def __array__(
        self, dtype: NpDtype | None = None, copy: bool | None = None
    ) -> np.ndarray:
        if self.sp_index.ngaps == 0:

View on GitHub (pinned to 3b7651241d)