sgl-project/sglang · error · ValueError
QKV and cos/sin tensors must be on the same CUDA device
Error message
QKV and cos/sin tensors must be on the same CUDA device
What it means
All QKV tensors and the cos/sin RoPE tables must live on the same CUDA device. The kernel launches on img_q's device and cross-device pointers would be invalid.
Source
Thrown at python/sglang/kernels/ops/diffusion/rope/hunyuan_qkv_pack_triton.py:187
txt_tokens = txt_q.shape[1]
expected_img = (batch, img_tokens, num_heads, head_dim)
expected_txt = (batch, txt_tokens, num_heads, head_dim)
if any(tuple(x.shape) != expected_img for x in (img_q, img_k, img_v)):
raise ValueError("image QKV shapes must match")
if any(tuple(x.shape) != expected_txt for x in (txt_q, txt_k, txt_v)):
raise ValueError("text QKV shapes must match")
if any(x.stride(-1) != 1 for x in tensors):
raise ValueError("QKV last dimensions must be contiguous")
if head_dim <= 0 or head_dim > 128 or head_dim % 2:
raise ValueError("head_dim must be positive, even, and <= 128")
if cos.ndim != 2 or sin.ndim != 2 or cos.shape != sin.shape:
raise ValueError("cos and sin must have matching [S, D/2] shapes")
if cos.shape[0] < img_tokens or cos.shape[1] != head_dim // 2:
raise ValueError("cos/sin shape does not cover image tokens and head_dim")
if not cos.is_cuda or not sin.is_cuda or cos.stride(-1) != 1 or sin.stride(-1) != 1:
raise ValueError("cos and sin must be CUDA and last-dim contiguous")
if cos.device != img_q.device or sin.device != img_q.device:
raise ValueError("QKV and cos/sin tensors must be on the same CUDA device")
total_tokens = img_tokens + txt_tokens
storage = torch.empty(
(3, batch, total_tokens, num_heads, head_dim),
device=img_q.device,
dtype=img_q.dtype,
)
args = []
for x in tensors:
args.extend((x.stride(0), x.stride(1), x.stride(2)))
with torch.cuda.device(img_q.device):
_hunyuan_qkv_rope_pack_kernel[
lambda meta: (
batch * total_tokens,
triton.cdiv(num_heads, meta["BLOCK_HEADS"]),
)
](
*tensors,View on GitHub (pinned to 0132848349)
Solutions
- Move cos/sin to img_q.device before the call.
- Create per-rank RoPE tables using the local device in TP code.
- Set CUDA_VISIBLE_DEVICES or torch.cuda.set_device consistently.
Example fix
// before
cos = cos.to('cuda:0') # while img_q is on cuda:1
// after
cos = cos.to(img_q.device)
sin = sin.to(img_q.device) Defensive patterns
Strategy: validation
Validate before calling
assert cos.device == img_q.device and sin.device == img_q.device
Prevention
- Use img_q.device (not hardcoded 'cuda') when moving tensors.
- Build per-rank tables using the local device under TP.
When it happens
Trigger: img_q on cuda:1 while cos was created on cuda:0 (or vice versa) in multi-GPU / tensor-parallel Hunyuan inference.
Common situations: Tensor-parallel or pipeline-parallel setups where RoPE tables are cached on device 0 but projections run on another rank's device.
Related errors
- All inputs must be on the same device.
- {name}_block_cnt and {name}_block_idx must be on the same de
- All inputs must be on the same device.
- indices must be on q's device {device}, got {indices.device}
- topk_length must be a CUDA tensor
AI-assisted analysis of sgl-project/sglang@0132848349 (2026-08-28).
Data as JSON: /api/errors/bda904f1c9d352e1.
Report an issue: GitHub.