sgl-project/sglang
Documented errors, page 2 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| resolved target ' ' is not callable | exception | error | patching, type-check, source-patcher, sglang |
| trtllm_mla cannot serve decode context parallelism with… | validation | critical | attention-backend, trtllm-mla, context-parallelism, speculative-decoding, mla, sglang |
| --disaggregation-decode-retraction-backup=host_pool is only… | validation | error | sglang, pd-disaggregation, retraction, host-pool, config-validation |
| Group is destroyed. | validation | critical | distributed, collective, process-group, weakref |
| MiniMax H3 requires subblock_sparse_query_block_mask when… | exception | error | minimax-h3, sparse-attention, dit, mask-required |
| ReplaySSM inputs must be on the same device. | exception | error | helion, kda, replayssm, multi-gpu, device-placement |
| Unknown KV cache quantization method | validation | error | quantization, kv-cache, config-validation, registry |
| Error on rank 0 | exception | critical | distributed, multi-gpu, rank-failure, error-propagation |
| Invalid . Must be > 0. | validation | error | ltx-2, video-generation, config-validation, sequence-parallelism |
| LTX-2 conditioning token count mismatch | validation | error | ltx-2, token-count, conditioning, resolution, shape-mismatch |
| WindowedAttentionKVCache holds only the trailing window and… | error_code | error | mlx, kv-cache, sliding-window, attention-mask |
| keyframe resolved_frame_index values disagree with semantic… | exception | error | minimax-h3, keyframes, validation, pipeline |
| model_index.json._minimax_h3.task_aliases must map strings… | validation | error | minimax-h3, model-index, task-aliases, config-validation |
| {response.error} | exception | error | action-inference, scheduler, runtime-error, propagated-error |
| --enable-linear-replayssm-spec with… | validation | error | sglang, replayssm, ragged-verify, environment-variable, kda |
| Transformer has no attribute for cache-dit blocks. | validation | error | cache-dit, block-adapter, model-internals, integration |
| LoRA adapter with rank is incompatible with the current… | validation | error | lora, memory-pool, rank, config, sglang |
| MXFP8 fused decode prologue requires interleaved K/V scale… | exception | error | mxfp8, scale-buffers, tensor-shape, decode, inkling |
| Server was launched with --enable-cfg-parallel but this… | validation | error | cfg, classifier-free-guidance, cfg-parallel, server-args, request-validation |
| Kimi K3 additional parameter schema accepts no values | validation | error | kimi-k3, json-schema, tool-calling, validation |
| Model path ' ' is already registered for pipeline | validation | error | registry, pipeline-registration, diffusers, duplicate-entry |
| Generator must be provided | validation | error | generator, vae-sampling, determinism, batch-construction |
| Humming FP8 dispatch requires | validation | error | deepep, humming, fp8, tensor-parallel, shape-validation |
| not a complete raw Muse Glimmer HF checkpoint | validation | error | checkpoint-integrity, missing-keys, weight-loading |
| SANA-WM refiner requires a string prompt or one prompt per… | validation | error | sana-wm, refiner, prompt-validation, batch-mismatch, valueerror |
| This layer norm doesn't support feature dim >= 64KB. | validation | error | fla, rms-norm, triton, feature-dim-limit |
| Cosmos3 action prompt list must contain only strings | validation | error | cosmos3, action-endpoint, prompt-validation, type-validation |
| MiniMax-H3 quality="high" is validated only for the strict… | validation | error | minimax-h3, hardware-validation, quality-mode, video-generation |
| packed seq_len not divisible by the combined… | validation | error | sequence-parallel, ulysses, ring-attention, divisibility, padding |
| Post-load processing produced a meta tensor | exception | critical | meta-tensor, post-load, uninitialized-weights, model-loader, sglang |
| Error happened when batch testing peer-to-peer access from | exception | error | cuda, p2p, custom-all-reduce, nccl, subprocess |
| Expected image placeholder token(s), found . | validation | error | multimodal, kimi-k3, placeholder-mismatch, validation |
| get_num_tokens_per_bs_for_target_verify is deprecated; use… | console | info | speculative-decoding, deprecation, api-rename, python-warnings |
| A quantized checkpoint requires an in-tree native encoder… | validation | error | quantization, architecture-unsupported, checkpoint, text-encoder |
| DFLASH mask_token_id is outside the target vocab size… | validation | critical | sglang, speculative-decoding, dflash, vocab-size, embedding, model-loading |
| Head slice size evaluates to zero | validation | critical | disaggregation, tensor-parallel, kv-cache, integer-division |
| MXFP8 fused decode prologue requires head_dim-aligned Q/K/V. | exception | error | mxfp8, quantization, head-dim, decode, inkling |
| [pred_noise_to_pred_video] Invalid timestep shape | validation | error | tensor-shape, diffusion, scheduler, validation |
| Role rank failed to initialize. | exception | critical | disaggregated-serving, multi-gpu, worker-init, nccl, tensor-parallel |
| unsupported audio interleave mode | validation | error | multimodal, config-validation, audio-interleave, dots-note-omni |
| unsupported mode | validation | error | deepep, moe, dispatcher, enum, cuda-graph |
| Comfy NVFP4 layer needs U8 packed weights and FP8 block… | validation | error | quantization, nvfp4, fp8, dtype-mismatch, checkpoint-validation |
| fuse_swiglu_interleaved set on an incompatible fused_moe… | validation | error | moe, triton, swiglu, quantization, dtype, feature-guard |
| memory_position_mode must be one of | validation | error | joy-echo, memory, rope, config-validation, multimodal |
| NIXL transfer encountered ERR room= | exception | error | nixl, rdma, network, disaggregation, transfer-failure |
| Weight output_size_per_partition = | validation | error | marlin, gptq, tensor-parallel, shape-validation |
| DeepEP is not installed. Please install DeepEP package from… | validation | critical | deepep, moe, import-error, expert-parallel, distributed |
| HiSparse device KV transfer requires sgl_kernel.kvcacheio… | exception | critical | sgl-kernel, cuda, rocm, platform-support, hisparse |
| No ModelSlim MoE scheme found for layer | validation | error | modelslim, quantization, moe, ascend-npu, config |
| Invalid tokens_per_frame= | validation | error | ltx-2, video-generation, defensive-check, sequence-parallelism |
| Mamba2AttnBackend's forward is called directly instead of… | exception | error | mamba, hybrid-linear-attention, interface-contract, not-implemented, sglang |
| Invalid transition for | validation | error | disaggregation, request-state, state-machine, invalid-transition |
| ModelOpt is not available. Please install modelopt. | exception | error | modelopt, quantization, import-error, dependency, sglang |
| Mooncake Transfer Engine initialization failed. | exception | critical | mooncake, rdma, initialization, native-return-code, sglang |
| Online MXFP4 requantization from compressed-tensors NVFP4… | exception | error | quantization, mxfp4, nvfp4, compressed-tensors, quark, not-implemented, python |
| .kind must be a non-empty string | validation | error | validation, schema, multimodal, minimax-h3 |
| could not import any module prefix of | exception | error | import, patching, source-patcher, sglang |
| Expected hidden_size to be | validation | error | layernorm, shape-mismatch, hidden-size, validation |
| MXFP8 fused decode prologue requires contiguous interleaved… | exception | error | mxfp8, scale-buffers, contiguity, decode, inkling |
| Unsupported KV cache type | validation | error | flexkv, kv-cache, attributeerror, attention-backend |
| Kimi-K3 image processor is missing deferred-preprocessing… | validation | error | multimodal, kimi-k3, config-validation, preprocessing |
| MiniMax H3 packed sequence alignment | validation | critical | sequence-parallel, ulysses, ring-attention, alignment, divisibility |
| Anthropic redacted_thinking history is not supported | http | error | anthropic, redacted-thinking, conversation-history, request-conversion |
| --ep-dispatch-algorithm | validation | error | eplb, dispatch-algorithm, moe, a2a-backend, server-args |
| qkv_proj scale_inv : shape mismatch vs due to block… | exception | error | quantization, tensor-parallel, weight-loading, mimo |
| FlashInfer GDN prefill is not supported with… | validation | error | gdn, linear-attention, deterministic-inference, flashinfer, triton, config-validation, sglang |
| Prefill context parallelism with the TRTLLM MHA prefill… | validation | error | sglang, context-parallel, attention-backend, trtllm, sm100, prefill |
| unknown absorbed-bmm K variant | validation | error | q8kv8, triton, variant-selection, internal-invariant |
| Cannot duplicate `image` of batch size | validation | error | qwen-image, diffusers, batch-size, latents, image-editing |
| Cannot parse checkpoint quantization for | validation | error | quantization, checkpoint-parsing, gguf, text-encoder |
| dimension ( ) must be divisible by 2 | validation | error | pi05, sinusoidal-embedding, dimension-validation |
| Online quantization for | validation | error | quantization, online-quantization, architecture-unsupported, text-encoder |
| speculative_eagle_topk > 1 with page_size > 1 is only… | validation | error | speculative-decoding, eagle-topk, page-size, attention-backend, server-args |
| subgroup missing under | validation | error | config, namespace, projection-mismatch, sglang |
| {template_error} | validation | error | jinja, chat-template, openai-api, bad-request |
| Unknown forward method | validation | error | sarvam-moe, attention-backend, config |
| encounter invalid h_bar | validation | error | multimodal, image-preprocessing, ernie-4-5-vl, pixel-limits |
| Initialization failed. Please see the error messages above. | exception | critical | frontend, server-startup, spawn, oom, sglang |
| Mixed shared-outer LoRA formats detected across loaded… | exception | error | lora, moe, shape-mismatch, sglang |
| W4AFP8 shape_k = must be divisible by 8 for int32… | validation | error | quantization, w4afp8, humming, tensor-parallel, shape-validation |
| Comfy W4A4 layer has input size , incompatible with… | validation | error | quantization, comfy, w4a4, shape-mismatch, checkpoint-validation |
| Invalid packed `mixed_qkv` last dim= | exception | error | shape-validation, triton-kernel, gated-delta-rule, packed-decode |
| Mesh generation failed: surface extraction returned None… | exception | error | hunyuan3d, mesh-extraction, marching-cubes, degenerate-output |
| Unknown Pi05 Gemma variant | validation | error | pi05, gemma, config, invalid-variant |
| Unsupported ascend_dispatcher_output_dtype | validation | error | ascend, npu, moe, dispatcher, dtype, quantization |
| Comfy NVFP4 layer has an incompatible pre_quant_scale | validation | error | quantization, nvfp4, pre-quant-scale, shape-mismatch, checkpoint-validation |
| Layer-sharded HiCache backup does not support layout | validation | error | hicache, mla, layout, context-parallelism, sglang |
| MiniMax H3 denoise state must be a mapping | validation | error | minimax-h3, pipeline, batch-state, validation |
| output_ws should be prepared for cuda-graph mode | exception | error | sglang, vision-transformer, cuda-graph, kwargs-validation, multimodal |
| unsupported input for packed fused SiLU-mul | exception | error | swiglu, silu, triton, strides, bit-exact |
| DeepEP v2 MoE is not validated for | validation | error | moe, deepep, a2a-backend, server-args, model-architecture |
| Eagle3 MLA layer requires q_lora_rank in the draft config | exception | error | eagle3, speculative-decoding, mla, config-validation, kimi |
| Found more ' ' placeholders in input prompt than actual… | validation | error | multimodal, vision, prompt, validation |
| [internvl][internlm2] image_data provided but no images… | validation | error | multimodal, internvl, image-placeholders, prompt-validation |
| not in full attention layers | validation | error | kv-cache, hybrid-attention, layer-id, mapping, sglang |
| Sync request failed | exception | critical | sglang, deep-gemm, compilation, http-500, kernel-build |
| Cannot find NVIDIA Math-DX (cuBLASDx) headers. Install the… | exception | error | cuda, jit-kernel, missing-dependency, nvidia, build |
| f"Transformers-managed | exception | error | quantization, bitsandbytes, config, model-loading |
| Failed to import 'set_transfer_engine' from 'mooncake.pg'… | exception | error | mooncake, version-mismatch, elastic-ep, import-error, sglang |
| MiniMax H3 latent preparation requires pre-queue… | validation | error | minimax-h3, geometry, plan-validation, pipeline |