sgl-project/sglang

Documented errors, page 2 of 32. Back to sgl-project/sglang

Code / MessageTypeSeverityTags
resolved target ' ' is not callable
exception error patching, type-check, source-patcher, sglang
trtllm_mla cannot serve decode context parallelism with…
validation critical attention-backend, trtllm-mla, context-parallelism, speculative-decoding, mla, sglang
--disaggregation-decode-retraction-backup=host_pool is only…
validation error sglang, pd-disaggregation, retraction, host-pool, config-validation
Group is destroyed.
validation critical distributed, collective, process-group, weakref
MiniMax H3 requires subblock_sparse_query_block_mask when…
exception error minimax-h3, sparse-attention, dit, mask-required
ReplaySSM inputs must be on the same device.
exception error helion, kda, replayssm, multi-gpu, device-placement
Unknown KV cache quantization method
validation error quantization, kv-cache, config-validation, registry
Error on rank 0
exception critical distributed, multi-gpu, rank-failure, error-propagation
Invalid . Must be > 0.
validation error ltx-2, video-generation, config-validation, sequence-parallelism
LTX-2 conditioning token count mismatch
validation error ltx-2, token-count, conditioning, resolution, shape-mismatch
WindowedAttentionKVCache holds only the trailing window and…
error_code error mlx, kv-cache, sliding-window, attention-mask
keyframe resolved_frame_index values disagree with semantic…
exception error minimax-h3, keyframes, validation, pipeline
model_index.json._minimax_h3.task_aliases must map strings…
validation error minimax-h3, model-index, task-aliases, config-validation
{response.error}
exception error action-inference, scheduler, runtime-error, propagated-error
--enable-linear-replayssm-spec with…
validation error sglang, replayssm, ragged-verify, environment-variable, kda
Transformer has no attribute for cache-dit blocks.
validation error cache-dit, block-adapter, model-internals, integration
LoRA adapter with rank is incompatible with the current…
validation error lora, memory-pool, rank, config, sglang
MXFP8 fused decode prologue requires interleaved K/V scale…
exception error mxfp8, scale-buffers, tensor-shape, decode, inkling
Server was launched with --enable-cfg-parallel but this…
validation error cfg, classifier-free-guidance, cfg-parallel, server-args, request-validation
Kimi K3 additional parameter schema accepts no values
validation error kimi-k3, json-schema, tool-calling, validation
Model path ' ' is already registered for pipeline
validation error registry, pipeline-registration, diffusers, duplicate-entry
Generator must be provided
validation error generator, vae-sampling, determinism, batch-construction
Humming FP8 dispatch requires
validation error deepep, humming, fp8, tensor-parallel, shape-validation
not a complete raw Muse Glimmer HF checkpoint
validation error checkpoint-integrity, missing-keys, weight-loading
SANA-WM refiner requires a string prompt or one prompt per…
validation error sana-wm, refiner, prompt-validation, batch-mismatch, valueerror
This layer norm doesn't support feature dim >= 64KB.
validation error fla, rms-norm, triton, feature-dim-limit
Cosmos3 action prompt list must contain only strings
validation error cosmos3, action-endpoint, prompt-validation, type-validation
MiniMax-H3 quality="high" is validated only for the strict…
validation error minimax-h3, hardware-validation, quality-mode, video-generation
packed seq_len not divisible by the combined…
validation error sequence-parallel, ulysses, ring-attention, divisibility, padding
Post-load processing produced a meta tensor
exception critical meta-tensor, post-load, uninitialized-weights, model-loader, sglang
Error happened when batch testing peer-to-peer access from
exception error cuda, p2p, custom-all-reduce, nccl, subprocess
Expected image placeholder token(s), found .
validation error multimodal, kimi-k3, placeholder-mismatch, validation
get_num_tokens_per_bs_for_target_verify is deprecated; use…
console info speculative-decoding, deprecation, api-rename, python-warnings
A quantized checkpoint requires an in-tree native encoder…
validation error quantization, architecture-unsupported, checkpoint, text-encoder
DFLASH mask_token_id is outside the target vocab size…
validation critical sglang, speculative-decoding, dflash, vocab-size, embedding, model-loading
Head slice size evaluates to zero
validation critical disaggregation, tensor-parallel, kv-cache, integer-division
MXFP8 fused decode prologue requires head_dim-aligned Q/K/V.
exception error mxfp8, quantization, head-dim, decode, inkling
[pred_noise_to_pred_video] Invalid timestep shape
validation error tensor-shape, diffusion, scheduler, validation
Role rank failed to initialize.
exception critical disaggregated-serving, multi-gpu, worker-init, nccl, tensor-parallel
unsupported audio interleave mode
validation error multimodal, config-validation, audio-interleave, dots-note-omni
unsupported mode
validation error deepep, moe, dispatcher, enum, cuda-graph
Comfy NVFP4 layer needs U8 packed weights and FP8 block…
validation error quantization, nvfp4, fp8, dtype-mismatch, checkpoint-validation
fuse_swiglu_interleaved set on an incompatible fused_moe…
validation error moe, triton, swiglu, quantization, dtype, feature-guard
memory_position_mode must be one of
validation error joy-echo, memory, rope, config-validation, multimodal
NIXL transfer encountered ERR room=
exception error nixl, rdma, network, disaggregation, transfer-failure
Weight output_size_per_partition =
validation error marlin, gptq, tensor-parallel, shape-validation
DeepEP is not installed. Please install DeepEP package from…
validation critical deepep, moe, import-error, expert-parallel, distributed
HiSparse device KV transfer requires sgl_kernel.kvcacheio…
exception critical sgl-kernel, cuda, rocm, platform-support, hisparse
No ModelSlim MoE scheme found for layer
validation error modelslim, quantization, moe, ascend-npu, config
Invalid tokens_per_frame=
validation error ltx-2, video-generation, defensive-check, sequence-parallelism
Mamba2AttnBackend's forward is called directly instead of…
exception error mamba, hybrid-linear-attention, interface-contract, not-implemented, sglang
Invalid transition for
validation error disaggregation, request-state, state-machine, invalid-transition
ModelOpt is not available. Please install modelopt.
exception error modelopt, quantization, import-error, dependency, sglang
Mooncake Transfer Engine initialization failed.
exception critical mooncake, rdma, initialization, native-return-code, sglang
Online MXFP4 requantization from compressed-tensors NVFP4…
exception error quantization, mxfp4, nvfp4, compressed-tensors, quark, not-implemented, python
.kind must be a non-empty string
validation error validation, schema, multimodal, minimax-h3
could not import any module prefix of
exception error import, patching, source-patcher, sglang
Expected hidden_size to be
validation error layernorm, shape-mismatch, hidden-size, validation
MXFP8 fused decode prologue requires contiguous interleaved…
exception error mxfp8, scale-buffers, contiguity, decode, inkling
Unsupported KV cache type
validation error flexkv, kv-cache, attributeerror, attention-backend
Kimi-K3 image processor is missing deferred-preprocessing…
validation error multimodal, kimi-k3, config-validation, preprocessing
MiniMax H3 packed sequence alignment
validation critical sequence-parallel, ulysses, ring-attention, alignment, divisibility
Anthropic redacted_thinking history is not supported
http error anthropic, redacted-thinking, conversation-history, request-conversion
--ep-dispatch-algorithm
validation error eplb, dispatch-algorithm, moe, a2a-backend, server-args
qkv_proj scale_inv : shape mismatch vs due to block…
exception error quantization, tensor-parallel, weight-loading, mimo
FlashInfer GDN prefill is not supported with…
validation error gdn, linear-attention, deterministic-inference, flashinfer, triton, config-validation, sglang
Prefill context parallelism with the TRTLLM MHA prefill…
validation error sglang, context-parallel, attention-backend, trtllm, sm100, prefill
unknown absorbed-bmm K variant
validation error q8kv8, triton, variant-selection, internal-invariant
Cannot duplicate `image` of batch size
validation error qwen-image, diffusers, batch-size, latents, image-editing
Cannot parse checkpoint quantization for
validation error quantization, checkpoint-parsing, gguf, text-encoder
dimension ( ) must be divisible by 2
validation error pi05, sinusoidal-embedding, dimension-validation
Online quantization for
validation error quantization, online-quantization, architecture-unsupported, text-encoder
speculative_eagle_topk > 1 with page_size > 1 is only…
validation error speculative-decoding, eagle-topk, page-size, attention-backend, server-args
subgroup missing under
validation error config, namespace, projection-mismatch, sglang
{template_error}
validation error jinja, chat-template, openai-api, bad-request
Unknown forward method
validation error sarvam-moe, attention-backend, config
encounter invalid h_bar
validation error multimodal, image-preprocessing, ernie-4-5-vl, pixel-limits
Initialization failed. Please see the error messages above.
exception critical frontend, server-startup, spawn, oom, sglang
Mixed shared-outer LoRA formats detected across loaded…
exception error lora, moe, shape-mismatch, sglang
W4AFP8 shape_k = must be divisible by 8 for int32…
validation error quantization, w4afp8, humming, tensor-parallel, shape-validation
Comfy W4A4 layer has input size , incompatible with…
validation error quantization, comfy, w4a4, shape-mismatch, checkpoint-validation
Invalid packed `mixed_qkv` last dim=
exception error shape-validation, triton-kernel, gated-delta-rule, packed-decode
Mesh generation failed: surface extraction returned None…
exception error hunyuan3d, mesh-extraction, marching-cubes, degenerate-output
Unknown Pi05 Gemma variant
validation error pi05, gemma, config, invalid-variant
Unsupported ascend_dispatcher_output_dtype
validation error ascend, npu, moe, dispatcher, dtype, quantization
Comfy NVFP4 layer has an incompatible pre_quant_scale
validation error quantization, nvfp4, pre-quant-scale, shape-mismatch, checkpoint-validation
Layer-sharded HiCache backup does not support layout
validation error hicache, mla, layout, context-parallelism, sglang
MiniMax H3 denoise state must be a mapping
validation error minimax-h3, pipeline, batch-state, validation
output_ws should be prepared for cuda-graph mode
exception error sglang, vision-transformer, cuda-graph, kwargs-validation, multimodal
unsupported input for packed fused SiLU-mul
exception error swiglu, silu, triton, strides, bit-exact
DeepEP v2 MoE is not validated for
validation error moe, deepep, a2a-backend, server-args, model-architecture
Eagle3 MLA layer requires q_lora_rank in the draft config
exception error eagle3, speculative-decoding, mla, config-validation, kimi
Found more ' ' placeholders in input prompt than actual…
validation error multimodal, vision, prompt, validation
[internvl][internlm2] image_data provided but no images…
validation error multimodal, internvl, image-placeholders, prompt-validation
not in full attention layers
validation error kv-cache, hybrid-attention, layer-id, mapping, sglang
Sync request failed
exception critical sglang, deep-gemm, compilation, http-500, kernel-build
Cannot find NVIDIA Math-DX (cuBLASDx) headers. Install the…
exception error cuda, jit-kernel, missing-dependency, nvidia, build
f"Transformers-managed
exception error quantization, bitsandbytes, config, model-loading
Failed to import 'set_transfer_engine' from 'mooncake.pg'…
exception error mooncake, version-mismatch, elastic-ep, import-error, sglang
MiniMax H3 latent preparation requires pre-queue…
validation error minimax-h3, geometry, plan-validation, pipeline