pandas-dev/pandas · error · ValueError

must contain a symbolic group name.

Error message

{pat=} must contain a symbolic group name.

What it means

Raised by ArrowExtensionArray._str_extract when the supplied regex pattern contains no symbolic (named) group. The pyarrow path returns a struct whose fields are addressed positionally, but pandas' expand=True result is keyed by group names from `re.compile(pat).groupindex`, so at least one named group is mandatory. This is stricter than the object-dtype path, which accepts positional groups and synthesizes integer column names.

Solutions

  1. Add a symbolic group name to every capture group you want returned, e.g. `r"(?P<num>\d+)"`.
  2. If you cannot change the pattern, cast to object dtype (`ser.astype(object)`) before `.str.extract` to use the positional-group behavior.

Example fix

# before
ser.str.extract(r"(\d+)-(\d+)")
# after
ser.str.extract(r"(?P<a>\d+)-(?P<b>\d+)")
Defensive patterns

Strategy: validation

Validate before calling

import re
groups = re.compile(pat).groupindex
if not groups:
    raise ValueError(f"pattern {pat!r} needs at least one named group for arrow str.extract")
ser.str.extract(pat)

Type guard

def has_named_group(pat: str) -> bool:
    return len(re.compile(pat).groupindex) > 0

Try / catch

try:
    out = ser.str.extract(pat)
except ValueError as e:
    if "symbolic group name" in str(e):
        # add a default name and retry on object dtype
        out = ser.astype(object).str.extract(pat)
    else:
        raise

Prevention

When it happens

Trigger: Calling `ser.str.extract(r"(\d+)")` (positional group only) on a pyarrow-backed string Series. The check `len(re.compile(pat).groupindex.keys()) == 0` fires whenever no `(?P<name>...)` group is present.

Common situations: Porting `str.extract` code that relied on integer-numbered capture groups from an object Series to a string[pyarrow] Series; patterns written for the numpy path.

Related errors


AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11). Data as JSON: /api/errors/39a606025cc08c5c. Report an issue: GitHub.

Appendix: source

Thrown at pandas/core/arrays/arrow/array.py:3697

        result = self._apply_elementwise(predicate)
        return self._from_pyarrow_array(pa.chunked_array(result))

    def _str_casefold(self) -> Self:
        predicate = lambda val: val.casefold()
        result = self._apply_elementwise(predicate)
        return self._from_pyarrow_array(pa.chunked_array(result))

    def _str_encode(self, encoding: str, errors: str = "strict") -> Self:
        predicate = lambda val: val.encode(encoding, errors)
        result = self._apply_elementwise(predicate)
        return self._from_pyarrow_array(pa.chunked_array(result))

    def _str_extract(self, pat: str, flags: int = 0, expand: bool = True):
        if flags:
            raise NotImplementedError("Only flags=0 is implemented.")
        groups = re.compile(pat).groupindex.keys()
        if len(groups) == 0:
            raise ValueError(f"{pat=} must contain a symbolic group name.")
        result = pc.extract_regex(self._pa_array, pat)
        if expand:
            return {
                col: self._from_pyarrow_array(pc.struct_field(result, [i]))
                for col, i in zip(groups, range(result.type.num_fields), strict=True)
            }
        else:
            return type(self)(pc.struct_field(result, [0]))

    def _str_findall(self, pat: str, flags: int = 0) -> Self:
        regex = re.compile(pat, flags=flags)
        predicate = lambda val: regex.findall(val)
        result = self._apply_elementwise(predicate)
        return self._from_pyarrow_array(pa.chunked_array(result))

    def _str_get_dummies(self, sep: str = "|", dtype: NpDtype | None = None):
        if dtype is None:
            dtype = np.bool_

View on GitHub (pinned to 3b7651241d)