sgl-project/sglang · error · TypeError

cu_seqlens_q tensor must be Int32

Error message

cu_seqlens_q tensor must be Int32

What it means

The optional cu_seqlens_q tensor (cumulative sequence lengths for varlen batched queries) must have Int32 element type, matching the kernel's index arithmetic. Passing None is allowed when not using varlen mode.

Source

Thrown at python/sglang/kernels/ops/attention/flash_attn/cute/flash_fwd.py:217

        mCuSeqlensK_type: Type[cutlass.Numeric] | None,
        mSeqUsedQ_type: Type[cutlass.Numeric] | None,
        mSeqUsedK_type: Type[cutlass.Numeric] | None,
    ):
        # Get the data type and check if it is fp16 or bf16
        if const_expr(self.is_split_kv):
            # SplitKV writes float32 partial outputs; Q/K/V still fp16/bf16.
            if const_expr(not (mQ_type == mK_type == mV_type)):
                raise TypeError("Q/K/V must have the same data type")
            if const_expr(mO_type != Float32):
                raise TypeError("SplitKV partial output (mO) must be Float32")
        elif const_expr(not (mQ_type == mK_type == mV_type == mO_type)):
            raise TypeError("All tensors must have the same data type")
        if const_expr(mQ_type not in [cutlass.Float16, cutlass.BFloat16]):
            raise TypeError("Only Float16 or BFloat16 is supported")
        if const_expr(mLSE_type not in [None, Float32]):
            raise TypeError("LSE tensor must be Float32")
        if const_expr(mCuSeqlensQ_type not in [None, Int32]):
            raise TypeError("cu_seqlens_q tensor must be Int32")
        if const_expr(mCuSeqlensK_type not in [None, Int32]):
            raise TypeError("cu_seqlens_k tensor must be Int32")
        if const_expr(mSeqUsedQ_type not in [None, Int32]):
            raise TypeError("seqused_q tensor must be Int32")
        if const_expr(mSeqUsedK_type not in [None, Int32]):
            raise TypeError("seqused_k tensor must be Int32")
        assert mQ_type == self.dtype

    def _setup_attributes(self):
        # ///////////////////////////////////////////////////////////////////////////////
        # Shared memory layout: Q/K/V
        # ///////////////////////////////////////////////////////////////////////////////
        (
            sQ_layout_atom,
            sK_layout_atom,
            sV_layout_atom,
            sO_layout_atom,
            sP_layout_atom,

View on GitHub (pinned to 0132848349)

Solutions

  1. Cast to int32: cu_seqlens_q = cu_seqlens_q.to(torch.int32)
  2. Or compute directly in int32: torch.zeros(..., dtype=torch.int32) and accumulate
  3. Verify cu_seqlens_k has the same treatment (it has an identical check)

Example fix

// before
cu_seqlens_q = torch.cumsum(seq_lens, 0)  # int64
// after
cu_seqlens_q = torch.cumsum(seq_lens, 0).to(torch.int32)
Defensive patterns

Strategy: type-guard

Validate before calling

if cu_seqlens_q is not None:
    assert cu_seqlens_q.dtype == torch.int32

Type guard

def int32_or_none(t) -> bool:
    return t is None or t.dtype == torch.int32

Prevention

When it happens

Trigger: Supplying cu_seqlens_q as Int64 (the default for torch.cumsum output) or Int16/UInt32 to FlashAttentionForward.

Common situations: Computing cu_seqlens with torch.cumsum(seq_lens, dim=0) which yields int64 and passing it directly; converting data between frameworks that default to int64 indices.

Related errors


AI-assisted analysis of sgl-project/sglang@0132848349 (2026-08-28). Data as JSON: /api/errors/28b0986ade595e89. Report an issue: GitHub.