pandas-dev/pandas · error · ImportError
pyarrow>= is required for PyArrow backed StringArray.
Error message
pyarrow>={PYARROW_MIN_VERSION} is required for PyArrow backed StringArray. What it means
Thrown by StringDtype.__init__ in pandas/core/arrays/string_.py:216 when storage='pyarrow' is requested but the optional pyarrow dependency is unavailable (or older than PYARROW_MIN_VERSION). pandas raises ImportError rather than ValueError because this is a missing-dependency condition, distinct from an invalid configuration.
Solutions
- Install/upgrade pyarrow: pip install -U 'pyarrow>={minimum}'. Run python -c 'import pyarrow; print(pyarrow.__version__)' to confirm.
- Fall back to python storage: pd.StringDtype(storage='python') or dtype='string[python]'.
- If using a requirements file, add pyarrow with a lower bound matching pandas' PYARROW_MIN_VERSION.
Example fix
// before # pyarrow not installed pd.Series(['a'], dtype='string[pyarrow]') # raises ImportError // after pip install 'pyarrow>=15.0' pd.Series(['a'], dtype='string[pyarrow]')
Defensive patterns
Strategy: fallback
Validate before calling
try:
import pyarrow # noqa: F401
HAS_PA = True
except ImportError:
HAS_PA = False
storage = 'pyarrow' if HAS_PA else 'python'
dtype = pd.StringDtype(storage=storage) Type guard
null
Try / catch
try:
dtype = pd.StringDtype(storage='pyarrow')
except ImportError:
dtype = pd.StringDtype(storage='python') Prevention
- Pin pyarrow in requirements with a lower bound matching pandas' minimum.
- Feature-detect pyarrow in code that may run in minimal environments.
- Provide a python-storage fallback in user-facing code when pyarrow is optional.
When it happens
Trigger: Using dtype='string[pyarrow]' or pd.StringDtype(storage='pyarrow') in an environment where pyarrow is not installed or is below the minimum supported version. Setting mode.string_storage='pyarrow' without pyarrow installed.
Common situations: Lightweight CI image without pyarrow. Downgrading pyarrow in a pinned environment. Fresh venv where pandas was installed without the pyarrow extra (pip install pandas without 'pyarrow' extra, then using pyarrow-backed strings).
Related errors
- pyarrow>= is required for PyArrow backed…
- pyarrow>= is required for PyArrow backed…
- dtype ' ' does not support operation
- dtype ' ' does not support operation 'quantile
- Unable to import required dependency
AI-assisted analysis of pandas-dev/pandas@3b7651241d (2026-08-11).
Data as JSON: /api/errors/bc7d9a164cbfd58d.
Report an issue: GitHub.
Appendix: source
Thrown at pandas/core/arrays/string_.py:216
storage: str | None = None,
na_value: libmissing.NAType | float = libmissing.NA,
) -> None:
# infer defaults
if storage is None:
storage = config["mode"]["string_storage"]
if storage == "auto":
if HAS_PYARROW:
storage = "pyarrow"
else:
storage = "python"
# validate options
if storage not in {"python", "pyarrow"}:
raise ValueError(
f"Storage must be 'python' or 'pyarrow'. Got {storage} instead."
)
if storage == "pyarrow" and not HAS_PYARROW:
raise ImportError(
f"pyarrow>={PYARROW_MIN_VERSION} is required for PyArrow "
"backed StringArray."
)
if isinstance(na_value, float) and np.isnan(na_value):
# when passed a NaN value, always set to np.nan to ensure we use
# a consistent NaN value (and we can use `dtype.na_value is np.nan`)
na_value = np.nan
elif na_value is not libmissing.NA:
raise ValueError(f"'na_value' must be np.nan or pd.NA, got {na_value}")
self._storage = cast("str", storage)
self._na_value = na_value
def __repr__(self) -> str:
storage = "" if self.storage == "pyarrow" else "storage='python', "
return f"<StringDtype({storage}na_value={self._na_value})>"
View on GitHub (pinned to 3b7651241d)