{"record":{"id":"18ff79f14027495c","repo":"sgl-project/sglang","slug":"qkv-last-dimensions-must-be-contiguous","errorCode":null,"errorMessage":"QKV last dimensions must be contiguous","messagePattern":"QKV last dimensions must be contiguous","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"python/sglang/kernels/ops/diffusion/rope/hunyuan_qkv_pack_triton.py","lineNumber":177,"sourceCode":"    sin: torch.Tensor,\n) -> tuple[torch.Tensor, torch.Tensor, torch.Tensor]:\n    tensors = (img_q, img_k, img_v, txt_q, txt_k, txt_v)\n    if any(x.ndim != 4 for x in tensors):\n        raise ValueError(\"QKV tensors must have shape [B, S, H, D]\")\n    if any(not x.is_cuda or x.dtype != torch.bfloat16 for x in tensors):\n        raise ValueError(\"QKV tensors must be CUDA bfloat16 tensors\")\n    if any(x.device != img_q.device for x in tensors):\n        raise ValueError(\"QKV tensors must be on the same CUDA device\")\n    batch, img_tokens, num_heads, head_dim = img_q.shape\n    txt_tokens = txt_q.shape[1]\n    expected_img = (batch, img_tokens, num_heads, head_dim)\n    expected_txt = (batch, txt_tokens, num_heads, head_dim)\n    if any(tuple(x.shape) != expected_img for x in (img_q, img_k, img_v)):\n        raise ValueError(\"image QKV shapes must match\")\n    if any(tuple(x.shape) != expected_txt for x in (txt_q, txt_k, txt_v)):\n        raise ValueError(\"text QKV shapes must match\")\n    if any(x.stride(-1) != 1 for x in tensors):\n        raise ValueError(\"QKV last dimensions must be contiguous\")\n    if head_dim <= 0 or head_dim > 128 or head_dim % 2:\n        raise ValueError(\"head_dim must be positive, even, and <= 128\")\n    if cos.ndim != 2 or sin.ndim != 2 or cos.shape != sin.shape:\n        raise ValueError(\"cos and sin must have matching [S, D/2] shapes\")\n    if cos.shape[0] < img_tokens or cos.shape[1] != head_dim // 2:\n        raise ValueError(\"cos/sin shape does not cover image tokens and head_dim\")\n    if not cos.is_cuda or not sin.is_cuda or cos.stride(-1) != 1 or sin.stride(-1) != 1:\n        raise ValueError(\"cos and sin must be CUDA and last-dim contiguous\")\n    if cos.device != img_q.device or sin.device != img_q.device:\n        raise ValueError(\"QKV and cos/sin tensors must be on the same CUDA device\")\n\n    total_tokens = img_tokens + txt_tokens\n    storage = torch.empty(\n        (3, batch, total_tokens, num_heads, head_dim),\n        device=img_q.device,\n        dtype=img_q.dtype,\n    )\n    args = []","sourceCodeStart":159,"sourceCodeEnd":195,"githubUrl":"https://github.com/sgl-project/sglang/blob/0132848349585cfe6aae51c4941cbae872505f8a/python/sglang/kernels/ops/diffusion/rope/hunyuan_qkv_pack_triton.py#L159-L195","documentation":"The Hunyuan QKV RoPE pack Triton kernel requires every Q/K/V tensor (image and text) to have unit stride in the last dimension (contiguous head_dim). The Triton kernel indexes head elements assuming a dense innermost dimension, so non-contiguous inputs would read wrong memory.","triggerScenarios":"Calling hunyuan_qkv_rope_pack (directly or via _hunyuan_pack_qkv) with any of img_q/img_k/img_v/txt_q/txt_k/txt_v produced by a non-contiguous operation, e.g. a transpose, narrow, or slice such as x[:, :, ::2] giving stride(-1) != 1.","commonSituations":"Passing split half tensors from interleaved RoPE-style slicing (x[..., :d/2] is fine but x[..., ::2] is not), or transposed attention projections from a custom attention module.","solutions":["Make each tensor contiguous before the call: q = q.contiguous() etc.","Restructure upstream slicing to produce last-dim-contiguous views (slice leading dims, not the innermost one).","Check x.stride(-1) == 1 in a debug assert before calling to find the offending tensor."],"exampleFix":"// before\nq, k, v = x[..., ::2], x[..., 1::2]  # stride(-1) == 2\nout = hunyuan_qkv_rope_pack(q, k, v, qt, kt, vt, cos, sin)\n// after\nq, k, v = x[..., ::2].contiguous(), x[..., 1::2].contiguous()\nout = hunyuan_qkv_rope_pack(q, k, v, qt, kt, vt, cos, sin)","handlingStrategy":"validation","validationCode":"def qkv_ok(*ts):\n    return all(t.stride(-1) == 1 for t in ts)\nassert qkv_ok(img_q, img_k, img_v, txt_q, txt_k, txt_v)","typeGuard":null,"tryCatchPattern":"try:\n    out = hunyuan_qkv_rope_pack(...)\nexcept ValueError as e:\n    if 'contiguous' in str(e):\n        tensors = [t.contiguous() for t in tensors]\n        out = hunyuan_qkv_rope_pack(...)\n    else:\n        raise","preventionTips":["Call .contiguous() on tensors produced by slicing/transpose.","Assert stride(-1)==1 in unit tests for attention inputs."],"tags":["triton","rope","tensor-contiguity","hunyuan"],"backgroundTag":"non-contiguous-tensor-stride","analyzedSha":"0132848349585cfe6aae51c4941cbae872505f8a","analyzedAt":"2026-08-28T05:10:05.995Z","schemaVersion":2},"datasetVersion":"2026-08-28T06:17:29.519Z"}