pandas-dev/pandas · error · ValueError

Column length mismatch

Error message

Column length mismatch: {len(columns)} vs. {K}

What it means

Raised by SparseFrameAccessor._prep_index (used by from_spmatrix) when len(columns) does not equal K, the second dimension of the source sparse matrix. The provided column labels must exactly cover the matrix's column count.

Solutions

  1. Pass columns=None to let pandas assign a RangeIndex.
  2. Build columns from shape: columns=[f'c{i}' for i in range(mat.shape[1])].
  3. Validate len(columns) == mat.shape[1] before calling.

Example fix

# before
pd.DataFrame.sparse.from_spmatrix(mat, columns=['a','b'])  # mat is (5,3)
# after
cols = [f'c{i}' for i in range(mat.shape[1])]
pd.DataFrame.sparse.from_spmatrix(mat, columns=cols)
Defensive patterns

Strategy: validation

Validate before calling

def columns_match_matrix(columns, mat) -> bool:
    return columns is None or len(columns) == mat.shape[1]

Type guard

def valid_columns_for(columns, mat) -> bool:
    return columns is None or len(list(columns)) == mat.shape[1]

Try / catch

try:
    df = pd.DataFrame.sparse.from_spmatrix(mat, columns=columns)
except ValueError as e:
    if 'Column length mismatch' in str(e):
        df = pd.DataFrame.sparse.from_spmatrix(mat)  # default RangeIndex
    else:
        raise

Prevention

When it happens

Trigger: pd.DataFrame.sparse.from_spmatrix(mat, columns=['a','b']) where mat.shape == (N, 3); passing a Index/Series of column names with the wrong length.

Common situations: Hard-coded column list that drifted from matrix shape; reusing columns from a different matrix; transposing the matrix but not the labels.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/190a11920fe0c126. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/sparse/accessor.py:502

    @staticmethod
    def _prep_index(data, index, columns):
        from pandas.core.indexes.api import (
            default_index,
            ensure_index,
        )

        N, K = data.shape
        if index is None:
            index = default_index(N)
        else:
            index = ensure_index(index)
        if columns is None:
            columns = default_index(K)
        else:
            columns = ensure_index(columns)

        if len(columns) != K:
            raise ValueError(f"Column length mismatch: {len(columns)} vs. {K}")
        if len(index) != N:
            raise ValueError(f"Index length mismatch: {len(index)} vs. {N}")
        return index, columns

View on GitHub (pinned to 3b7651241d)