pandas-dev/pandas · error · ValueError
must contain a symbolic group name.
Error message
{pat=} must contain a symbolic group name. What it means
Raised by ArrowExtensionArray._str_extract when the supplied regex pattern contains no symbolic (named) group. The pyarrow path returns a struct whose fields are addressed positionally, but pandas' expand=True result is keyed by group names from `re.compile(pat).groupindex`, so at least one named group is mandatory. This is stricter than the object-dtype path, which accepts positional groups and synthesizes integer column names.
Solutions
- Add a symbolic group name to every capture group you want returned, e.g. `r"(?P<num>\d+)"`.
- If you cannot change the pattern, cast to object dtype (`ser.astype(object)`) before `.str.extract` to use the positional-group behavior.
Example fix
# before ser.str.extract(r"(\d+)-(\d+)") # after ser.str.extract(r"(?P<a>\d+)-(?P<b>\d+)")
Defensive patterns
Strategy: validation
Validate before calling
import re
groups = re.compile(pat).groupindex
if not groups:
raise ValueError(f"pattern {pat!r} needs at least one named group for arrow str.extract")
ser.str.extract(pat) Type guard
def has_named_group(pat: str) -> bool:
return len(re.compile(pat).groupindex) > 0 Try / catch
try:
out = ser.str.extract(pat)
except ValueError as e:
if "symbolic group name" in str(e):
# add a default name and retry on object dtype
out = ser.astype(object).str.extract(pat)
else:
raise Prevention
- Always use (?P<name>...) named groups in str.extract patterns
- Add a unit test asserting named groups when targeting pyarrow strings
When it happens
Trigger: Calling `ser.str.extract(r"(\d+)")` (positional group only) on a pyarrow-backed string Series. The check `len(re.compile(pat).groupindex.keys()) == 0` fires whenever no `(?P<name>...)` group is present.
Common situations: Porting `str.extract` code that relied on integer-numbered capture groups from an object Series to a string[pyarrow] Series; patterns written for the numpy path.
Related errors
- Only flags=0 is implemented.
- expected a string object, not
- ambiguous is not supported.
- is not supported
- as_unit not implemented for
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/39a606025cc08c5c.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/arrow/array.py:3697
result = self._apply_elementwise(predicate)
return self._from_pyarrow_array(pa.chunked_array(result))
def _str_casefold(self) -> Self:
predicate = lambda val: val.casefold()
result = self._apply_elementwise(predicate)
return self._from_pyarrow_array(pa.chunked_array(result))
def _str_encode(self, encoding: str, errors: str = "strict") -> Self:
predicate = lambda val: val.encode(encoding, errors)
result = self._apply_elementwise(predicate)
return self._from_pyarrow_array(pa.chunked_array(result))
def _str_extract(self, pat: str, flags: int = 0, expand: bool = True):
if flags:
raise NotImplementedError("Only flags=0 is implemented.")
groups = re.compile(pat).groupindex.keys()
if len(groups) == 0:
raise ValueError(f"{pat=} must contain a symbolic group name.")
result = pc.extract_regex(self._pa_array, pat)
if expand:
return {
col: self._from_pyarrow_array(pc.struct_field(result, [i]))
for col, i in zip(groups, range(result.type.num_fields), strict=True)
}
else:
return type(self)(pc.struct_field(result, [0]))
def _str_findall(self, pat: str, flags: int = 0) -> Self:
regex = re.compile(pat, flags=flags)
predicate = lambda val: regex.findall(val)
result = self._apply_elementwise(predicate)
return self._from_pyarrow_array(pa.chunked_array(result))
def _str_get_dummies(self, sep: str = "|", dtype: NpDtype | None = None):
if dtype is None:
dtype = np.bool_View on GitHub (pinned to 3b7651241d)