sgl-project/sglang · error · TypeError

seqused_k tensor must be Int32

Error message

seqused_k tensor must be Int32

What it means

The optional seqused_k tensor (per-batch actual key/value sequence lengths for masking) must be Int32, mirroring the seqused_q check. Passing None is allowed.

Source

Thrown at python/sglang/kernels/ops/attention/flash_attn/cute/flash_fwd.py:223

            # SplitKV writes float32 partial outputs; Q/K/V still fp16/bf16.
            if const_expr(not (mQ_type == mK_type == mV_type)):
                raise TypeError("Q/K/V must have the same data type")
            if const_expr(mO_type != Float32):
                raise TypeError("SplitKV partial output (mO) must be Float32")
        elif const_expr(not (mQ_type == mK_type == mV_type == mO_type)):
            raise TypeError("All tensors must have the same data type")
        if const_expr(mQ_type not in [cutlass.Float16, cutlass.BFloat16]):
            raise TypeError("Only Float16 or BFloat16 is supported")
        if const_expr(mLSE_type not in [None, Float32]):
            raise TypeError("LSE tensor must be Float32")
        if const_expr(mCuSeqlensQ_type not in [None, Int32]):
            raise TypeError("cu_seqlens_q tensor must be Int32")
        if const_expr(mCuSeqlensK_type not in [None, Int32]):
            raise TypeError("cu_seqlens_k tensor must be Int32")
        if const_expr(mSeqUsedQ_type not in [None, Int32]):
            raise TypeError("seqused_q tensor must be Int32")
        if const_expr(mSeqUsedK_type not in [None, Int32]):
            raise TypeError("seqused_k tensor must be Int32")
        assert mQ_type == self.dtype

    def _setup_attributes(self):
        # ///////////////////////////////////////////////////////////////////////////////
        # Shared memory layout: Q/K/V
        # ///////////////////////////////////////////////////////////////////////////////
        (
            sQ_layout_atom,
            sK_layout_atom,
            sV_layout_atom,
            sO_layout_atom,
            sP_layout_atom,
        ) = self._get_smem_layout_atom()
        self.sQ_layout = cute.tile_to_shape(
            sQ_layout_atom,
            (self.tile_m, self.tile_hdim),
            (0, 1),
        )

View on GitHub (pinned to 0132848349)

Solutions

  1. Cast: seqused_k = seqused_k.to(torch.int32)
  2. Keep all length metadata tensors int32 consistently in your attention wrapper

Example fix

// before
seqused_k = kv_lens  # int64
// after
seqused_k = kv_lens.to(torch.int32)
Defensive patterns

Strategy: type-guard

Validate before calling

if seqused_k is not None:
    assert seqused_k.dtype == torch.int32

Type guard

def int32_or_none(t) -> bool:
    return t is None or t.dtype == torch.int32

Prevention

When it happens

Trigger: Providing seqused_k with a dtype other than Int32 (typically Int64) to FlashAttentionForward.

Common situations: Reusing an int64 sequence-length table for both seqused_q and seqused_k; masking KV padding with un-cast metadata tensors.

Related errors


AI-assisted analysis of sgl-project/sglang@0132848349 (2026-08-28). Data as JSON: /api/errors/bff7140e177ef8a4. Report an issue: GitHub.