pandas-dev/pandas · error · KeyError

Columns not found

Error message

Columns not found: {str(bad_keys)[1:-1]}

What it means

Raised by SelectionMixin.__getitem__ (the bracket selector on groupby/rolling/expanding/resample window objects) when a list/tuple/Series/Index/ndarray of column names is passed and at least one name does not exist in the underlying DataFrame's columns. The bad keys are computed as set(key).difference(self.obj.columns) and rendered without the surrounding list brackets. It exists so column selection on grouped/windowed data fails fast with the offending names instead of producing empty groups.

Solutions

  1. Inspect df.columns and correct the list to only contain existing labels.
  2. Intersect before selecting: cols = [c for c in wanted if c in df.columns]; grouped[cols].
  3. Normalize names on read: df = pd.read_csv(...).rename(columns=lambda c: c.strip()).
  4. Guard with set difference: missing = set(wanted) - set(df.columns); assert not missing, missing.

Example fix

// before
grouped = df.groupby('id')[['id', 'amout']]
// after
wanted = ['id', 'amount']
missing = set(wanted) - set(df.columns)
assert not missing, f'missing cols: {missing}'
grouped = df.groupby('id')[wanted]
Defensive patterns

Strategy: validation

Validate before calling

wanted = ['a', 'b', 'c']
missing = set(wanted) - set(df.columns)
if missing:
    raise KeyError(f'columns not in df: {missing}')
grouped = df.groupby('id')[wanted]

Type guard

def safe_columns(df, wanted):
    cols = list(wanted) if not isinstance(wanted, str) else [wanted]
    missing = [c for c in cols if c not in df.columns]
    if missing:
        raise KeyError(missing)
    return cols

Prevention

When it happens

Trigger: Calling df.groupby('a')[['a','missing']] or rolling/expanding/resample objects indexed with a list where any label is absent, e.g. grouped[['b','typo']]. Also triggered by passing an np.ndarray or Index of labels containing a misspelled/dropped column, or after a rename/drop that left stale names in user code.

Common situations: Renaming columns (df.rename) or dropping them after copy-pasting a selection list; typos in column names; case mismatches ('Name' vs 'name'); trailing/leading whitespace introduced by reading CSVs; code written against an older schema.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/1523009fbd0aa686. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/base.py:219

        if isinstance(self.obj, ABCSeries):
            return self.obj

        if self._selection is not None:
            return self.obj[self._selection_list]

        if len(self.exclusions) > 0:
            return self.obj._drop_axis(self.exclusions, axis=1)
        else:
            return self.obj

    def __getitem__(self, key):
        if self._selection is not None:
            raise IndexError(f"Column(s) {self._selection} already selected")

        if isinstance(key, (list, tuple, ABCSeries, ABCIndex, np.ndarray)):
            if len(self.obj.columns.intersection(key)) != len(set(key)):
                bad_keys = list(set(key).difference(self.obj.columns))
                raise KeyError(f"Columns not found: {str(bad_keys)[1:-1]}")
            return self._gotitem(list(key), ndim=2)

        else:
            if key not in self.obj:
                raise KeyError(f"Column not found: {key}")
            ndim = self.obj[key].ndim
            return self._gotitem(key, ndim=ndim)

    def _gotitem(self, key, ndim: int, subset=None):
        """
        sub-classes to define
        return a sliced object

        Parameters
        ----------
        key : str / list of selections
        ndim : {1, 2}
            requested ndim of result

View on GitHub (pinned to 3b7651241d)