vllm-project/vllm · error · RuntimeError
Failed to initialize PyNcclEplbCommunicator ({exc}).
Error message
Failed to initialize PyNcclEplbCommunicator ({exc}). What it means
Error "Failed to initialize PyNcclEplbCommunicator ({exc})." thrown in vllm-project/vllm.
Source
Thrown at vllm/distributed/eplb/eplb_communicator.py:725
if unsupported_dtypes:
raise RuntimeError(
"EPLB communicator 'pynccl' requested but expert weights contain "
"unsupported dtypes: "
f"({', '.join(str(dtype) for dtype in unsupported_dtypes)})."
)
device_comm = group_coordinator.device_communicator
pynccl_comm = (
getattr(device_comm, "pynccl_comm", None)
if device_comm is not None
else None
)
if pynccl_comm is None or pynccl_comm.disabled or not pynccl_comm.available:
raise RuntimeError("EPLB communicator 'pynccl' requested but unavailable.")
try:
return PyNcclEplbCommunicator(pynccl_comm=pynccl_comm)
except Exception as exc:
raise RuntimeError(
f"Failed to initialize PyNcclEplbCommunicator ({exc})."
) from exc
is_stateless = isinstance(group_coordinator, StatelessGroupCoordinator)
if is_stateless:
if backend == "nixl":
pass # handled below with defer_remote_setup=True
elif backend not in ("torch_nccl", "pynccl"):
raise ValueError(
f"Elastic EP requires 'torch_nccl', 'pynccl', or 'nixl' "
f"EPLB communicator (got '{backend}')."
)
else:
if backend == "torch_nccl":
logger.warning(
"Stateless elastic EP requires PyNCCL backend. "
"Forcing EPLB communicator to 'pynccl'."
)View on GitHub (pinned to c794754062)
Solutions
- Inspect the wrapped exception in the error for the root cause of PyNcclEplbCommunicator initialization failure.
When it happens
Trigger: Raised at vllm/distributed/eplb/eplb_communicator.py:725 when validation fails: Failed to initialize PyNcclEplbCommunicator. Typically triggered by an incompatible or incomplete vLLM configuration, an unsupported platform/backend combination, or a runtime resource/dependency that is missing.
Common situations: Commonly encountered at vllm/distributed/eplb/eplb_communicator.py:725 during vLLM startup/config validation or runtime setup when: (1) conflicting CLI flags or config fields are combined, (2) the current platform (CUDA/ROCm/CPU/XPU) or installed optional packages do not support the requested feature, or (3) a required value is absent or out of range. Resolve by correcting the configuration as described in the message, or by selecting a supported alternative.
AI-assisted analysis of vllm-project/vllm@c794754062 (2026-08-14).
Data as JSON: /api/errors/455d4646aeb4097e.
Report an issue: GitHub.