sgl-project/sglang · error · ValueError
`initial_state` must be 4D (got ndim={initial_state.ndim}).
Error message
`initial_state` must be 4D (got ndim={initial_state.ndim}). What it means
The recurrent state for replaySSM decode must be 4D, typically [N, HV, K, V] (slots, heads, key dim, value dim). Passing a 3D or 5D state tensor means the state layout doesn't match the kernel's expectation and the update would read/write out of bounds.
Source
Thrown at python/sglang/kernels/ops/attention/fla/fused_recurrent_linear_replayssm.py:486
(``d_cache`` / ``k_cache`` / ``g_cache``) and the per-decode-row
``write_pos`` cursor. ``initial_state`` is both the checkpoint read (h0)
and the (flush-only) checkpoint write (ht), in place.
Allocates nothing persistent: the caller owns the ring tensors and is
responsible for advancing / resetting ``write_pos`` (e.g. ``(write_pos+1) %
L`` after each step). This is a STANDALONE kernel; the memory-pool / cache
integration is a later phase.
"""
if mixed_qkv.ndim != 2:
raise ValueError(f"`mixed_qkv` must be 2D (got ndim={mixed_qkv.ndim}).")
if mixed_qkv.stride(-1) != 1:
raise ValueError("`mixed_qkv` must be contiguous in the last dim.")
if b.ndim != 2:
raise ValueError(f"`b` must be 2D (got b.ndim={b.ndim}).")
if A_log.ndim != 1:
raise ValueError("`A_log` must be a 1D tensor.")
if initial_state.ndim != 4:
raise ValueError(f"`initial_state` must be 4D (got ndim={initial_state.ndim}).")
if not out.is_contiguous():
raise ValueError("`out` must be contiguous.")
if write_pos.ndim != 1 or write_pos.dtype != torch.int32:
raise ValueError("`write_pos` must be a 1D int32 tensor.")
if force_flush is not None and (
force_flush.ndim != 1 or force_flush.dtype != torch.int32
):
raise ValueError("`force_flush` must be a 1D int32 tensor or None.")
B = mixed_qkv.shape[0]
num_state_slots, HV, V, K = initial_state.shape
qkv_dim = mixed_qkv.shape[1]
q_dim = (qkv_dim - HV * V) // 2
if q_dim <= 0 or q_dim % K != 0:
raise ValueError(
f"Invalid packed `mixed_qkv` last dim={qkv_dim} for HV={HV}, V={V}, K={K}."
)
H = q_dim // KView on GitHub (pinned to 0132848349)
Solutions
- Reshape the state to 4D: state.view(num_slots, HV, K, V) (keep it contiguous)
- Verify the pool's per-slot stride equals HV*K*V elements
- Check ring/write_pos tensors are 1D int32 to satisfy sibling checks
Example fix
// before state = pool_flat[slots] # [N, HV*K*V] // after state = pool_flat.view(num_slots, HV, K, V)[slots] # 4D view
Defensive patterns
Strategy: validation
Validate before calling
assert initial_state.ndim == 4, initial_state.shape initial_state = initial_state.view(-1, HV, K, V) if initial_state.ndim != 4 else initial_state
Type guard
def is_4d_state(s: torch.Tensor) -> bool:
return s.ndim == 4 Prevention
- Store state pools as [slots, HV, K, V] tensors
- Never flatten state rows; index the 4D pool directly
When it happens
Trigger: Passing a flattened state pool [N, HV*K*V], or a per-batch state [B, HV, K*V] with wrong rank.
Common situations: Integrating with a memory pool that stores states as flat rows; converting between chunked-kernel state layout [B, H, K, V] and a slot-indexed pool with extra dims.
Related errors
- `mixed_qkv` must be 2D (got ndim={mixed_qkv.ndim}).
- `b` must be 2D (got b.ndim={b.ndim}).
- `A_log` must be a 1D tensor.
- `dt_bias` must have {HV * K} elements (got {dt_bias.numel()}
- Invalid packed Q size {q_dim}: must be divisible by K={K}. K
AI-assisted analysis of sgl-project/sglang@0132848349 (2026-08-28).
Data as JSON: /api/errors/3e3f60987190a4ed.
Report an issue: GitHub.