sgl-project/sglang

Documented errors, page 4 of 32. Back to sgl-project/sglang

Code / MessageTypeSeverityTags
Action endpoint is not implemented for
validation error action-endpoint, not-implemented, model-support, dispatch
allowed media domains cannot be empty
validation error validation, media, security, config
AttentionOffsetCache should not store data
error_code error mlx, kv-cache, api-misuse
eqlen with B>1 and T %
exception error kda, linear-attention, shape-mismatch, not-implemented
experimental_sgl_marlin LoRA requires…
validation error lora, marlin, virtual-experts, startup-validation, experimental, sglang
f"The class must implement the 'embedding' method, see…
exception error quantization, embedding, not-implemented, constructor
language_model does not support set_embed_and_head().
exception error speculative-decoding, kimi, attribute-error, model-loading
MiniMax H3 SGLang backend only supports…
validation error minimax-h3, video-generation, output-mode, validation
no_rope_layers contains non-binary entries
validation error config-validation, mlx, rope, muse-glimmer
PEFT lora_alpha must be a positive integer
validation error lora, peft, validation, config
thinking.budget_tokens must be >= 1024
validation error anthropic, thinking, validation, request-validation, pydantic
UlyssesAttention's all-to-all spans the combined sequence…
exception critical attention, ring-parallelism, sequence-parallel, distributed, not-implemented
Using a slow tokenizer. This might cause a significant…
console warning tokenizer, performance, huggingface, startup
Download failed for after attempts due to download errors…
panic critical network, download, huggingface, retry-exhausted, weights
Mamba storage zero-copy requires page_first layout, got
validation error sglang, mamba, layout, zero-copy, hierarchical-cache
MiniMax-H3 on MPS requires synchronous layerwise offload for
validation error minimax-h3, mps, apple-silicon, memory-offload, server-args
unsupported input for Sana fused bias-SiLU
exception error sana, diffusion, triton, memory-format, channels-last
Attention backend override
validation error attention, backend-override, configuration, enum-mismatch
Cannot determine attention scale for
error_code error mlx, attention, model-compatibility, patching
denoising_strength must be positive
validation error scheduler, flow-matching, validation, diffusion, config
Expected CHW image tensor, got shape
exception error multimodal, image-processing, tensor-shape, step3-vl
load_path is required for STA_inference mode
validation error sta, attention, sparse-tuning, kwargs-validation
The checkpoint declares quantization, but the model did not…
validation error quantization, model-mismatch, linear-layers, text-encoder, silent-failure-guard
DSV4 ragged verify does not support context parallel (CP)…
exception critical deepseek-v4, ragged-verify, context-parallelism, env-var, speculative-decoding, sglang
Encoder produced tokens, but preprocessor metadata expected
http error multimodal, token-count, encoder, preprocessor, mismatch
language_model does not support get_embed_and_head().
exception error speculative-decoding, kimi, attribute-error, model-loading
_load_function expects 'pkg.module.symbol', got
exception error import, config-validation, debug-utils, sglang
Downloaded model files are still corrupted for
panic critical download, corruption, huggingface, validation, weights
encode metadata not ready
http error timeout, metadata, encoder, disaggregation
Expected a 3D packed tensor for
exception critical lfm2, moe, weight-loading, tensor-shape
expert-pack header coverage is inconsistent
exception critical moe, expert-pack, binary-format, header-validation, sglang
Including the scheme in --host
console info url, deprecation, host-config, networking, tests
Layer-sharded direct HiCache backup only supports…
validation error hicache, direct-io, layout, sglang
--mm-feature-transport=cuda_ipc requires NVIDIA CUDA.
validation error sglang, cuda, multimodal, ipc, hardware-requirement
--speculative-use-rejection-sampling is incompatible with…
validation error speculative-decoding, rejection-sampling, determinism, server-args
Subclasses of BaseDiT must define
validation error sglang, dit, subclass-contract, class-attribute, import-time
The memory capacity is unbalanced. Some GPUs may be…
error_code error distributed, tp, gpu-memory, resource-conflict
Failed to cancel VMM transport slice(s)
exception error cuda, vmm, aggregate-error, dispatch
unknown residency policy
validation error config, validation, layerwise-offload, residency
weight_prefix must be 'w13' or 'w2', got
validation error quantization, moe, npu, modelslim, validation
attn_sink must be float32 with shape
validation error mla, sparse-attention, tensor-validation, dtype
Checkpoint provides gate weight
exception critical laguna, weight-loading, config-mismatch, gating
KV cache dtype mismatch: prefill server has kv_cache_dtype=
exception critical disaggregation, pd-disagg, kv-cache-dtype, config-mismatch, quantization
NIXL KV transfer has no KV memory segments
validation error nixl, disaggregation, hicache, memory-registration, pd-disaggregation
NPU detected, but torchair package is not installed. Please…
exception error npu, ascend, torchair, torch-compile, import-error
queued MiniMax H3 jobs require pre-queue resolved temporal…
validation critical minimax-h3, queue-invariant, frame-count, internal
Unsupported model architecture
exception error registry, alias, architecture, unsupported-architecture
Ascend A2/A3 NPU does not support nvfp4…
validation error ascend, npu, deepep, nvfp4, quantization, hardware-support
canonical request missing
validation error minimax-h3, schema-validation, required-field
Could not access latents of provided encoder_output
validation error vae, encoder-output, attribute-access, diffusers-compat
DFLASH requires draft num_hidden_layers in config. Got…
validation error speculative-decoding, dflash, missing-config-field
Each permutation group must reside on the same gpu
validation error marlin, tensor-parallel, tile-alignment, shape-validation
Expected CHW image tensor with 1 or 3 channels, got shape
exception error multimodal, image-processing, channels, step3-vl
Hunyuan3D requires 'image_path' input.
validation error hunyuan3d, input-validation, multimodal, missing-argument
LPLB fused solver unavailable
exception critical lplb, backend-unavailable, jit, cuda
No model weights found in
exception critical mimo-audio, model-loading, file-not-found, huggingface, weights
Weight output_partition_size =
validation error fp8, quantization, tensor-parallel, shape-mismatch
Z-Image transformer has no `rotary_emb`. It likely loaded…
validation critical z-image, model-loading, fallback, rotary-embeddings, diffusers
Can not import FA3 in sgl_kernel. Please check your…
exception critical sglang, flash-attention, import-error, native-extension, environment
ffprobe returned invalid JSON for MiniMax H3 output
exception error minimax-h3, ffprobe, json-parse, validation
kv-canary: launch_canary_plan_kernels_torch_reference…
validation error kv-cache, missing-argument, ragged-tensor, validation
Serve backend factory returned ; expected…
exception critical sglang, cli, plugin-api-mismatch, type-check
Unsupported params_dtype
validation error quantization, dtype, npu, modelslim, w8a8
Kimi-K3 MLA K projection must remain GGUF Q4_0
exception error gguf, quantization, weight-loading, kimi, mla
move_kv_cache is not yet supported for MiniMaxSparseKVPool…
exception error sglang, kv-cache, not-implemented, speculative-decoding, minimax
--prefill-only-disable-kv-cache currently requires…
validation error sglang, kv-cache, prefill, embedding, server-args
seq_len < used rows
validation error validation, sequence-length, alignment, minimax-h3
Serialized W4A8 layer
validation error quantization, dimension-mismatch, checkpoint
Unsupported image type
validation error multimodal, input-validation, type-error, image-processing
unsupported input for wan_rmsnorm_silu
validation error memory-format, triton, vae, wan, validation
DFlash layer selection requires num_target_layers >= 4. Got…
validation error speculative-decoding, dflash, layer-selection, config-validation
--hicache-host-memory-mode buffer_only requires an SWA host…
validation error hicache, buffer-only, swa, memory-sizing, validation
Hunyuan3D reference attention requires a shared cache.
validation error runtime, diffusion, reference-attention, missing-cache
are not all equal
validation error npu, quant-scale, weight-loading, per-tensor-quant, allclose
The 'enable-mixed-chunk' feature is currently unsupported…
exception error ascend, npu, mixed-chunk, mla, deepseek, not-implemented, huawei
Timesteps must be provided
validation error hunyuan3d, pipeline-order, missing-state, scheduler
Unsupported Kimi-K3 encoder media item
validation error multimodal, kimi-k3, input-validation
cached keyframe preparation disagrees with the resolved plan
exception error minimax-h3, cache-coherence, pipeline, keyframes
Cannot configure checkpoint quantization for
validation error quantization, checkpoint-parsing, config-error, text-encoder
Cannot transition from terminal state to
validation warning disaggregation, request-state, state-machine, terminal-state, race-condition
delta payload size mismatch: expected
error_code critical rocm, allreduce, tensor-parallel
--enable-linear-replayssm is not supported under PD…
validation error sglang, replayssm, pd-disaggregation, unsupported-feature
Humming does not support DeepEP
validation error deepep, humming, moe, dtype, config-validation
mask_search_files_path_pos, mask_search_files_path_neg, and…
validation error sta, attention, sparse-tuning, kwargs-validation, multimodal
Only support per-tensor scaling factor for fp8 KV cache
validation error kv-cache, fp8, scale-format, checkpoint-validation
Please install diso via `pip install diso`, or set mc_algo…
validation error python, import-error, missing-dependency, diso, marching-cubes, hunyuan3d
Rank-local TP shard produced for DTensor parameter
exception error tensor-parallel, fsdp, dtensor, distributed, weight-loading
SANA-WM streaming does not support CFG parallel; run…
exception error sana-wm, streaming, cfg-parallel, notimplementederror, server-args
unknown norm_type
validation error config, validation, diffusion, norm
req_to_token table is empty but gather mask is non-empty.
exception error sglang, speculative-decoding, dflash, kv-cache, memory-pool, internal-invariant
Cosmos3 rollout supports T2V/T2I only; I2V/V2V…
validation error cosmos3, rollout, i2v, sde, rl-sampling
Currently only attention backends are supported for…
validation error sglang, deterministic-inference, mla, attention-backend, deepseek
DeepSeek-V4 GGUF mapping collision
panic error gguf, deepseek, name-collision, weight-mapping
DSA indexer weights_proj LoRA is incompatible with…
exception error dsa, lora, piecewise-cuda-graph, prefill, deepseek, sglang
Expected Hunyuan3D2PipelineConfig, got
validation error hunyuan3d, pipeline-config, type-mismatch, sglang
fl2va keyframe preparation requires cached pre-queue probe…
exception error minimax-h3, fl2va, pipeline-ordering, missing-metadata
Host-pool retraction does not support pure-SWA models.
validation error disaggregation, retraction, sliding-window, swa, unified-cache, value-error
--mamba-max-states-per-path must be -1 (unlimited) or a…
validation error sglang, mamba, server-args, validation, startup
MIXED_PRECISION layer group
exception error quantization, mixed-precision, quark, mxfp4, config-validation, python
NIXL heterogeneous-TP direct-to-host KV transfer is not…
exception error nixl, disaggregation, heterogeneous-tp, hicache, pd-disaggregation