sgl-project/sglang · error · TypeError
Only Float16 or BFloat16 is supported
Error message
Only Float16 or BFloat16 is supported
What it means
The CUTLASS DSL flash-attention forward kernel is only instantiated for Float16 and BFloat16 element types; the Q tensor's dtype failed this check. The kernel templates/gemm configurations do not exist for fp32, fp8, or integer types.
Source
Thrown at python/sglang/kernels/ops/attention/flash_attn/cute/flash_fwd.py:213
mV_type: Type[cutlass.Numeric],
mO_type: Type[cutlass.Numeric],
mLSE_type: Type[cutlass.Numeric] | None,
mCuSeqlensQ_type: Type[cutlass.Numeric] | None,
mCuSeqlensK_type: Type[cutlass.Numeric] | None,
mSeqUsedQ_type: Type[cutlass.Numeric] | None,
mSeqUsedK_type: Type[cutlass.Numeric] | None,
):
# Get the data type and check if it is fp16 or bf16
if const_expr(self.is_split_kv):
# SplitKV writes float32 partial outputs; Q/K/V still fp16/bf16.
if const_expr(not (mQ_type == mK_type == mV_type)):
raise TypeError("Q/K/V must have the same data type")
if const_expr(mO_type != Float32):
raise TypeError("SplitKV partial output (mO) must be Float32")
elif const_expr(not (mQ_type == mK_type == mV_type == mO_type)):
raise TypeError("All tensors must have the same data type")
if const_expr(mQ_type not in [cutlass.Float16, cutlass.BFloat16]):
raise TypeError("Only Float16 or BFloat16 is supported")
if const_expr(mLSE_type not in [None, Float32]):
raise TypeError("LSE tensor must be Float32")
if const_expr(mCuSeqlensQ_type not in [None, Int32]):
raise TypeError("cu_seqlens_q tensor must be Int32")
if const_expr(mCuSeqlensK_type not in [None, Int32]):
raise TypeError("cu_seqlens_k tensor must be Int32")
if const_expr(mSeqUsedQ_type not in [None, Int32]):
raise TypeError("seqused_q tensor must be Int32")
if const_expr(mSeqUsedK_type not in [None, Int32]):
raise TypeError("seqused_k tensor must be Int32")
assert mQ_type == self.dtype
def _setup_attributes(self):
# ///////////////////////////////////////////////////////////////////////////////
# Shared memory layout: Q/K/V
# ///////////////////////////////////////////////////////////////////////////////
(
sQ_layout_atom,View on GitHub (pinned to 0132848349)
Solutions
- Cast inputs: Q = Q.to(torch.bfloat16) (or torch.float16) before calling the op
- Ensure the model runs in half precision (--dtype bfloat16/half or equivalent server config)
- For fp8 attention use the dedicated fp8 attention backend, not this kernel
Example fix
// before out = fa_fwd(Q, K, V) # Q,K,V are float32 // after out = fa_fwd(Q.to(torch.bfloat16), K.to(torch.bfloat16), V.to(torch.bfloat16))
Defensive patterns
Strategy: type-guard
Validate before calling
assert Q.dtype in (torch.float16, torch.bfloat16), f'flash fwd supports fp16/bf16 only, got {Q.dtype}' Type guard
def is_half_dtype(t) -> bool:
return t.dtype in (torch.float16, torch.bfloat16) Prevention
- Cast at the attention-wrapper boundary: Q=Q.to(torch.bfloat16), etc.
- Validate once in the backend-selection code which dtypes each backend supports
When it happens
Trigger: Passing Q (and by extension K/V/O, which must already match) as torch.float32, torch.float8_*, or any non-fp16/bf16 dtype to FlashAttentionForward.
Common situations: Feeding un-cast model weights or activations in fp32 (e.g. a model loaded without half precision); testing with toy float32 tensors; accidentally using fp8 tensors intended for a different attention backend.
Related errors
- SplitKV partial output (mO) must be Float32
- All tensors must have the same data type
- LSE tensor must be Float32
- cu_seqlens_q tensor must be Int32
- cu_seqlens_k tensor must be Int32
AI-assisted analysis of sgl-project/sglang@0132848349 (2026-08-28).
Data as JSON: /api/errors/d42fc53fbaa60061.
Report an issue: GitHub.