sgl-project/sglang · error · ValueError
timestep must be contiguous
Error message
timestep must be contiguous
What it means
The Triton kernel for ltx2_ada_values9 assumes a contiguous timestep tensor so it can compute flat row indices. Non-contiguous inputs (transposed, sliced, or expanded views) are rejected before launch.
Source
Thrown at python/sglang/kernels/ops/diffusion/modulate/ltx2_ada_values_triton.py:147
tl.store(out2_ptr + base, (table2 + temb2).to(tl.bfloat16), mask=mask)
tl.store(out3_ptr + base, (table3 + temb3).to(tl.bfloat16), mask=mask)
tl.store(out4_ptr + base, (table4 + temb4).to(tl.bfloat16), mask=mask)
tl.store(out5_ptr + base, (table5 + temb5).to(tl.bfloat16), mask=mask)
tl.store(out6_ptr + base, (table6 + temb6).to(tl.bfloat16), mask=mask)
tl.store(out7_ptr + base, (table7 + temb7).to(tl.bfloat16), mask=mask)
tl.store(out8_ptr + base, (table8 + temb8).to(tl.bfloat16), mask=mask)
def ltx2_ada_values9(
scale_shift_table: torch.Tensor,
timestep: torch.Tensor,
) -> tuple[torch.Tensor, ...]:
if timestep.ndim != 3:
raise ValueError("timestep must have shape [B, S, 9 * D]")
if not timestep.is_cuda or timestep.dtype != torch.bfloat16:
raise ValueError("timestep must be a CUDA bfloat16 tensor")
if not timestep.is_contiguous():
raise ValueError("timestep must be contiguous")
if scale_shift_table.ndim != 2 or scale_shift_table.shape[0] != 9:
raise ValueError("scale_shift_table must have shape [9, D]")
if (
not scale_shift_table.is_cuda
or scale_shift_table.dtype not in (torch.bfloat16, torch.float32)
or scale_shift_table.stride(-1) != 1
):
raise ValueError(
"scale_shift_table must be CUDA, bf16/fp32, last-dim contiguous"
)
total_params = int(scale_shift_table.shape[0])
hidden = int(scale_shift_table.shape[1])
if hidden <= 0 or timestep.shape[-1] != total_params * hidden:
raise ValueError("timestep last dim must equal 9 * hidden")
if hidden % 256 != 0 or hidden > 8192:
raise ValueError("hidden size is outside the supported LTX2 fast-path range")
View on GitHub (pinned to 0132848349)
Solutions
- Call .contiguous() on timestep before the call
- Materialize expanded views: emb.expand(B, S, D).contiguous()
- Use .reshape instead of views/slices that yield non-contiguous results
Example fix
# before emb = emb[:, None, :].expand(B, S, 9*D) vals = ltx2_ada_values9(table, emb) # after emb = emb[:, None, :].expand(B, S, 9*D).contiguous() vals = ltx2_ada_values9(table, emb)
Defensive patterns
Strategy: validation
Validate before calling
if not timestep.is_contiguous():
timestep = timestep.contiguous() Type guard
def contiguous_or_fix(t):
return t if t.is_contiguous() else t.contiguous() Prevention
- Materialize expanded views before kernel calls
- Prefer .reshape over stride-0 expand for kernel inputs
When it happens
Trigger: Passing an expanded view (timestep[:, None, :].expand(...)) without materializing, a transposed tensor, or a slice along batch/seq that breaks contiguity.
Common situations: Broadcasting a [B, 1, 9*D] embedding across sequence positions with expand (stride-0 seq dim); memory-layout tricks from a checkpoint; sliced batch views in pipeline parallelism.
Related errors
- `mixed_qkv` must be contiguous in the last dim.
- q, k, and v must be contiguous in head_size
- scale_shift_table must be CUDA, bf16/fp32, last-dim contiguo
- hidden size is outside the supported LTX2 fast-path range
- kv-canary: {name} must be contiguous
AI-assisted analysis of sgl-project/sglang@0132848349 (2026-08-28).
Data as JSON: /api/errors/61b1eb93923fef31.
Report an issue: GitHub.