sgl-project/sglang · error · TypeError
cu_seqlens_q tensor must be Int32
Error message
cu_seqlens_q tensor must be Int32
What it means
The optional cu_seqlens_q tensor (cumulative sequence lengths for varlen batched queries) must have Int32 element type, matching the kernel's index arithmetic. Passing None is allowed when not using varlen mode.
Source
Thrown at python/sglang/kernels/ops/attention/flash_attn/cute/flash_fwd.py:217
mCuSeqlensK_type: Type[cutlass.Numeric] | None,
mSeqUsedQ_type: Type[cutlass.Numeric] | None,
mSeqUsedK_type: Type[cutlass.Numeric] | None,
):
# Get the data type and check if it is fp16 or bf16
if const_expr(self.is_split_kv):
# SplitKV writes float32 partial outputs; Q/K/V still fp16/bf16.
if const_expr(not (mQ_type == mK_type == mV_type)):
raise TypeError("Q/K/V must have the same data type")
if const_expr(mO_type != Float32):
raise TypeError("SplitKV partial output (mO) must be Float32")
elif const_expr(not (mQ_type == mK_type == mV_type == mO_type)):
raise TypeError("All tensors must have the same data type")
if const_expr(mQ_type not in [cutlass.Float16, cutlass.BFloat16]):
raise TypeError("Only Float16 or BFloat16 is supported")
if const_expr(mLSE_type not in [None, Float32]):
raise TypeError("LSE tensor must be Float32")
if const_expr(mCuSeqlensQ_type not in [None, Int32]):
raise TypeError("cu_seqlens_q tensor must be Int32")
if const_expr(mCuSeqlensK_type not in [None, Int32]):
raise TypeError("cu_seqlens_k tensor must be Int32")
if const_expr(mSeqUsedQ_type not in [None, Int32]):
raise TypeError("seqused_q tensor must be Int32")
if const_expr(mSeqUsedK_type not in [None, Int32]):
raise TypeError("seqused_k tensor must be Int32")
assert mQ_type == self.dtype
def _setup_attributes(self):
# ///////////////////////////////////////////////////////////////////////////////
# Shared memory layout: Q/K/V
# ///////////////////////////////////////////////////////////////////////////////
(
sQ_layout_atom,
sK_layout_atom,
sV_layout_atom,
sO_layout_atom,
sP_layout_atom,View on GitHub (pinned to 0132848349)
Solutions
- Cast to int32: cu_seqlens_q = cu_seqlens_q.to(torch.int32)
- Or compute directly in int32: torch.zeros(..., dtype=torch.int32) and accumulate
- Verify cu_seqlens_k has the same treatment (it has an identical check)
Example fix
// before cu_seqlens_q = torch.cumsum(seq_lens, 0) # int64 // after cu_seqlens_q = torch.cumsum(seq_lens, 0).to(torch.int32)
Defensive patterns
Strategy: type-guard
Validate before calling
if cu_seqlens_q is not None:
assert cu_seqlens_q.dtype == torch.int32 Type guard
def int32_or_none(t) -> bool:
return t is None or t.dtype == torch.int32 Prevention
- Wrap cumsum: lambda x: torch.cumsum(x, 0).to(torch.int32)
- Store all attention metadata as int32 tensors in your batch preparation step
When it happens
Trigger: Supplying cu_seqlens_q as Int64 (the default for torch.cumsum output) or Int16/UInt32 to FlashAttentionForward.
Common situations: Computing cu_seqlens with torch.cumsum(seq_lens, dim=0) which yields int64 and passing it directly; converting data between frameworks that default to int64 indices.
Related errors
- cu_seqlens_k tensor must be Int32
- seqused_q tensor must be Int32
- seqused_k tensor must be Int32
- SplitKV partial output (mO) must be Float32
- All tensors must have the same data type
AI-assisted analysis of sgl-project/sglang@0132848349 (2026-08-28).
Data as JSON: /api/errors/28b0986ade595e89.
Report an issue: GitHub.