pandas-dev/pandas · error · ValueError
invalid normalization form
Error message
invalid normalization form
What it means
ValueError raised in _str_normalize of the ArrowStringArray mixin when the form argument is not one of the four Unicode normalization forms ('NFC', 'NFD', 'NFKC', 'NFKD'). The pyarrow backend validates the form before dispatching to pc.utf8_normalize or the unicodedata fallback.
Solutions
- Pass one of the exact uppercase forms: 'NFC', 'NFD', 'NFKC', 'NFKD'.
- Normalize user input upstream: form = form.strip().upper() and validate against the four allowed values.
- Provide a dropdown / enum in the UI so only valid forms are submitted.
Example fix
# before
s.str.normalize(user_form) # user_form may be 'nfc'
# after
allowed = {'NFC','NFD','NFKC','NFKD'}
s.str.normalize(user_form.strip().upper() if user_form.strip().upper() in allowed else 'NFC') Defensive patterns
Strategy: validation
Validate before calling
ALLOWED = {'NFC', 'NFD', 'NFKC', 'NFKD'}
def safe_normalize(s, form):
form = str(form).strip().upper()
if form not in ALLOWED:
raise ValueError(f'form must be one of {ALLOWED}; got {form!r}')
return s.str.normalize(form) Type guard
def is_valid_normalization(form) -> bool:
return isinstance(form, str) and form.strip().upper() in {'NFC','NFD','NFKC','NFKD'} Try / catch
try:
s.str.normalize(form)
except ValueError as e:
if 'normalization form' in str(e):
s.str.normalize('NFC') # safe default
else:
raise Prevention
- Constrain UI/config inputs for normalization form to the four canonical names.
- Normalize user input with .strip().upper() and validate against an allowlist.
When it happens
Trigger: s.str.normalize('NFC ') (trailing space); s.str.normalize('nfc') (lowercase); s.str.normalize('XYZ') on a pyarrow-backed string Series.
Common situations: User-supplied normalization form strings from config or CLI args that don't exactly match the canonical uppercase names.
Related errors
- Invalid side: . Side must be one of 'left', 'right', 'both
- contains not implemented with
- replace is not supported with a re.Pattern, callable repl…
- ambiguous is not supported.
- is not supported
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/7c47d6e2bcc405e8.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/_arrow_string_mixins.py:182
pa_pad = partial(pc.utf8_center, lean_left_on_odd_padding=lean_left)
else:
raise ValueError(
f"Invalid side: {side}. Side must be one of 'left', 'right', 'both'"
)
return self._from_pyarrow_array(
pa_pad(self._pa_array, width=width, padding=fillchar)
)
def _str_zfill(self, width: int) -> Self:
if pa_version_under21p0:
predicate = lambda val: val.zfill(width)
result = self._apply_elementwise(predicate)
return self._from_pyarrow_array(pa.chunked_array(result))
return self._from_pyarrow_array(pc.utf8_zfill(self._pa_array, width))
def _str_normalize(self, form: Literal["NFC", "NFD", "NFKC", "NFKD"]) -> Self:
if form not in ("NFC", "NFD", "NFKC", "NFKD"):
raise ValueError("invalid normalization form")
if form in ("NFC", "NFKC"):
# GH#64359 pc.utf8_normalize only decomposes; it skips the canonical
# composition step, so for the composing forms it returns decomposed
# output. Fall back to unicodedata for these.
predicate = lambda val: unicodedata.normalize(form, val)
result = self._apply_elementwise(predicate)
return self._from_pyarrow_array(pa.chunked_array(result))
return self._from_pyarrow_array(pc.utf8_normalize(self._pa_array, form=form))
def _str_get(self, i: int) -> Self:
lengths = pc.utf8_length(self._pa_array)
if i >= 0:
out_of_bounds = pc.greater_equal(i, lengths)
start = i
stop = i + 1
step = 1
else:
out_of_bounds = pc.greater(-i, lengths)View on GitHub (pinned to 3b7651241d)