{"record":{"id":"597daef9eeadd2f8","repo":"sgl-project/sglang","slug":"seqused-q-tensor-must-be-int32","errorCode":null,"errorMessage":"seqused_q tensor must be Int32","messagePattern":"seqused_q tensor must be Int32","errorType":"validation","errorClass":"TypeError","httpStatus":null,"severity":"error","filePath":"python/sglang/kernels/ops/attention/flash_attn/cute/flash_fwd.py","lineNumber":221,"sourceCode":"        # Get the data type and check if it is fp16 or bf16\n        if const_expr(self.is_split_kv):\n            # SplitKV writes float32 partial outputs; Q/K/V still fp16/bf16.\n            if const_expr(not (mQ_type == mK_type == mV_type)):\n                raise TypeError(\"Q/K/V must have the same data type\")\n            if const_expr(mO_type != Float32):\n                raise TypeError(\"SplitKV partial output (mO) must be Float32\")\n        elif const_expr(not (mQ_type == mK_type == mV_type == mO_type)):\n            raise TypeError(\"All tensors must have the same data type\")\n        if const_expr(mQ_type not in [cutlass.Float16, cutlass.BFloat16]):\n            raise TypeError(\"Only Float16 or BFloat16 is supported\")\n        if const_expr(mLSE_type not in [None, Float32]):\n            raise TypeError(\"LSE tensor must be Float32\")\n        if const_expr(mCuSeqlensQ_type not in [None, Int32]):\n            raise TypeError(\"cu_seqlens_q tensor must be Int32\")\n        if const_expr(mCuSeqlensK_type not in [None, Int32]):\n            raise TypeError(\"cu_seqlens_k tensor must be Int32\")\n        if const_expr(mSeqUsedQ_type not in [None, Int32]):\n            raise TypeError(\"seqused_q tensor must be Int32\")\n        if const_expr(mSeqUsedK_type not in [None, Int32]):\n            raise TypeError(\"seqused_k tensor must be Int32\")\n        assert mQ_type == self.dtype\n\n    def _setup_attributes(self):\n        # ///////////////////////////////////////////////////////////////////////////////\n        # Shared memory layout: Q/K/V\n        # ///////////////////////////////////////////////////////////////////////////////\n        (\n            sQ_layout_atom,\n            sK_layout_atom,\n            sV_layout_atom,\n            sO_layout_atom,\n            sP_layout_atom,\n        ) = self._get_smem_layout_atom()\n        self.sQ_layout = cute.tile_to_shape(\n            sQ_layout_atom,\n            (self.tile_m, self.tile_hdim),","sourceCodeStart":203,"sourceCodeEnd":239,"githubUrl":"https://github.com/sgl-project/sglang/blob/0132848349585cfe6aae51c4941cbae872505f8a/python/sglang/kernels/ops/attention/flash_attn/cute/flash_fwd.py#L203-L239","documentation":"The optional seqused_q tensor (per-batch actual query sequence lengths, used to mask padding) must be Int32. Passing None is allowed when seqused masking is not used.","triggerScenarios":"Providing seqused_q with dtype Int64 or anything other than Int32 to FlashAttentionForward.","commonSituations":"Passing seq_lens tensors straight from a dataloader (commonly int64); enabling page/padding masking with tensors prepared for a different backend that accepts int64.","solutions":["Cast: seqused_q = seqused_q.to(torch.int32)","Allocate length tables with dtype=torch.int32 from the start","Also cast seqused_k (checked next)"],"exampleFix":"// before\nseqused_q = seq_lens  # int64 from dataloader\n// after\nseqused_q = seq_lens.to(torch.int32)","handlingStrategy":"type-guard","validationCode":"if seqused_q is not None:\n    assert seqused_q.dtype == torch.int32","typeGuard":"def int32_or_none(t) -> bool:\n    return t is None or t.dtype == torch.int32","tryCatchPattern":null,"preventionTips":["Convert dataloader int64 lengths to int32 at the boundary","Keep seqused_q/seqused_k dtype checks in one validation function"],"tags":["cuda","dtype","flash-attention","masking","int32"],"backgroundTag":null,"analyzedSha":"0132848349585cfe6aae51c4941cbae872505f8a","analyzedAt":"2026-08-28T05:10:05.995Z","schemaVersion":2},"datasetVersion":"2026-08-28T06:17:29.519Z"}