{"record":{"id":"25ec53e35014881c","repo":"sgl-project/sglang","slug":"timestep-must-be-a-cuda-bfloat16-tensor","errorCode":null,"errorMessage":"timestep must be a CUDA bfloat16 tensor","messagePattern":"timestep must be a CUDA bfloat16 tensor","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"python/sglang/kernels/ops/diffusion/modulate/ltx2_ada_values_triton.py","lineNumber":145,"sourceCode":"    tl.store(out0_ptr + base, (table0 + temb0).to(tl.bfloat16), mask=mask)\n    tl.store(out1_ptr + base, (table1 + temb1).to(tl.bfloat16), mask=mask)\n    tl.store(out2_ptr + base, (table2 + temb2).to(tl.bfloat16), mask=mask)\n    tl.store(out3_ptr + base, (table3 + temb3).to(tl.bfloat16), mask=mask)\n    tl.store(out4_ptr + base, (table4 + temb4).to(tl.bfloat16), mask=mask)\n    tl.store(out5_ptr + base, (table5 + temb5).to(tl.bfloat16), mask=mask)\n    tl.store(out6_ptr + base, (table6 + temb6).to(tl.bfloat16), mask=mask)\n    tl.store(out7_ptr + base, (table7 + temb7).to(tl.bfloat16), mask=mask)\n    tl.store(out8_ptr + base, (table8 + temb8).to(tl.bfloat16), mask=mask)\n\n\ndef ltx2_ada_values9(\n    scale_shift_table: torch.Tensor,\n    timestep: torch.Tensor,\n) -> tuple[torch.Tensor, ...]:\n    if timestep.ndim != 3:\n        raise ValueError(\"timestep must have shape [B, S, 9 * D]\")\n    if not timestep.is_cuda or timestep.dtype != torch.bfloat16:\n        raise ValueError(\"timestep must be a CUDA bfloat16 tensor\")\n    if not timestep.is_contiguous():\n        raise ValueError(\"timestep must be contiguous\")\n    if scale_shift_table.ndim != 2 or scale_shift_table.shape[0] != 9:\n        raise ValueError(\"scale_shift_table must have shape [9, D]\")\n    if (\n        not scale_shift_table.is_cuda\n        or scale_shift_table.dtype not in (torch.bfloat16, torch.float32)\n        or scale_shift_table.stride(-1) != 1\n    ):\n        raise ValueError(\n            \"scale_shift_table must be CUDA, bf16/fp32, last-dim contiguous\"\n        )\n\n    total_params = int(scale_shift_table.shape[0])\n    hidden = int(scale_shift_table.shape[1])\n    if hidden <= 0 or timestep.shape[-1] != total_params * hidden:\n        raise ValueError(\"timestep last dim must equal 9 * hidden\")\n    if hidden % 256 != 0 or hidden > 8192:","sourceCodeStart":127,"sourceCodeEnd":163,"githubUrl":"https://github.com/sgl-project/sglang/blob/0132848349585cfe6aae51c4941cbae872505f8a/python/sglang/kernels/ops/diffusion/modulate/ltx2_ada_values_triton.py#L127-L163","documentation":"The LTX2 fused AdaLN Triton kernel is hard-coded for bfloat16 timestep inputs on CUDA. Tensors on CPU, other devices, or in other dtypes (fp16/fp32) fail this check.","triggerScenarios":"Passing a float32 timestep embedding (common when the projection layer outputs fp32), a fp16 tensor, or a CPU tensor in tests.","commonSituations":"Models kept in fp32 or fp16 precision; tests that create embeddings without device='cuda' or dtype=torch.bfloat16; upstream components returning fp32 projections that weren't cast.","solutions":["Cast: timestep = timestep.to(torch.bfloat16).cuda()","Align the timestep projection layer to bf16 output","Construct test inputs with device='cuda', dtype=torch.bfloat16","If fp32 is required, use the non-fused reference implementation"],"exampleFix":"# before\nvals = ltx2_ada_values9(table, t_emb_fp32)\n# after\nt_emb = t_emb_fp32.to(device='cuda', dtype=torch.bfloat16)\nvals = ltx2_ada_values9(table, t_emb)","handlingStrategy":"validation","validationCode":"timestep = timestep.to(device='cuda', dtype=torch.bfloat16)\nassert timestep.is_cuda and timestep.dtype == torch.bfloat16","typeGuard":"def t_ok(t: torch.Tensor) -> bool:\n    return t.is_cuda and t.dtype == torch.bfloat16","tryCatchPattern":null,"preventionTips":["Set the timestep projection layer to bf16 output","Create test embeddings with dtype=torch.bfloat16, device='cuda'"],"tags":["dtype","cuda","bfloat16","ltx2"],"backgroundTag":"unsupported-dtype","analyzedSha":"0132848349585cfe6aae51c4941cbae872505f8a","analyzedAt":"2026-08-28T05:10:05.995Z","schemaVersion":2},"datasetVersion":"2026-08-28T06:17:29.519Z"}