sgl-project/sglang
Documented errors, page 3 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| {name} | exception | error | attribute-access, partial-initialization, cuda-graph, wrapper |
| Sampling mask length | validation | error | disaggregation, sampling, buffer-capacity, pd-disaggregation |
| The hpc_ops attention backend only supports the default… | validation | error | attention-backend, hpc-ops, softmax-scaling, head-dim, sglang |
| video token dim != patch volume * channel for latent_shape=… | validation | error | validation, tensor-shape, unpatchify, minimax-h3 |
| MiMoV2ForCausalLM requires effective attention TP size | validation | critical | sglang, tensor-parallel, model-config, dp-attention, mimo |
| is not a config leaf (no NS namespace) | validation | error | config, typo, field-mapping, sglang |
| rope_pool_fused expects q/k/v to be 3-D | validation | error | shape-validation, rope, metal, sgl-kernel, tensor-dims |
| Weight input_size_per_partition = | validation | error | fp8, quantization, tensor-parallel, block-size |
| Helion KDA decode requires power-of-two key and value head… | validation | error | kda, helion, power-of-two, head-dim, model-config |
| miss_src, miss_dst, and miss_count must be provided… | validation | error | hisparse, argument-validation, all-or-none |
| NVFP4 global scale tensor must already be on the KV tensor… | validation | error | nvfp4, kv-cache, device-mismatch, quantization |
| resolution already failed on this ServerArgs; the handlers… | exception | error | server-args, resolution-pipeline, state-corruption, sglang |
| SGLANG_RUST_SERVER is not supported with the offline Engine… | exception | error | sglang, environment-variable, offline-engine, startup |
| This use case is not supported if api speculative execution… | exception | error | frontend, openai, chat-model, program-structure, sglang |
| Unexpected compressed-MLA dst_kv_ptrs length | exception | error | disaggregation, mla, pipeline-parallel, kv-pointers, internal-invariant |
| Unsupported OTLP protocol | validation | error | opentelemetry, tracing, configuration, env-var |
| Error on rank 0 | exception | critical | distributed, multi-gpu, rank-failure, error-propagation |
| mm_content_hashes has | validation | error | multimodal, content-hash, request-validation, receiver |
| online c128 does not support MTP | validation | error | deepseek-v4, c128, speculative-decoding, mtp, incompatible-features |
| shift must be positive | validation | error | scheduler, config-validation, sigma-shift, flow-matching |
| --speculative-ngram-external-sam-budget must be positive… | validation | error | speculative-decoding, ngram, external-corpus, server-args, validation |
| uniform_samples shape mismatch. Expected | validation | error | sglang, speculative-decoding, dflash, shape-mismatch, rng |
| VerifyCommit committed_tokens must be non-empty: request_id= | validation | error | sglang, speculative-decoding, decoupled, protocol-validation, empty-collection |
| DSV4 draft state transfer expects SWA-only NextN layers | validation | error | disaggregation, deepseek-v4, speculative-decoding, swa |
| InklingMultimodalProcessor | validation | error | multimodal, inkling, placeholder-mismatch, input-validation |
| LoRA on Intern-S2-Mobius model.meta_mlp routed banks is not… | validation | error | lora, intern-s2-mobius, moe, unsupported, sglang |
| LoRA pinned weight cache key collision for | validation | error | lora, cache, shape-mismatch, pin-memory, sglang |
| MXFP8 fused decode prologue requires K/V scale buffers. | exception | error | mxfp8, scale-buffers, decode, kv-cache, inkling |
| Not enough host memory for DSA indexer hierarchical cache… | validation | critical | memory, dsa-hicache, host-memory, startup |
| PEFT adapter_config.json must contain a JSON object | validation | error | lora, peft, json, config |
| QwenImage RoPE text cache overflow before denoising… | validation | error | qwen-image, rope, sequence-length, cache-overflow, multimodal |
| Comfy layer is missing checkpoint tensors | validation | error | quantization, checkpoint, safetensors, validation |
| Dynamic HiCache sidecars require HostPoolGroup. | validation | error | hicache, hybrid-cache, sidecar, type-error |
| ffprobe not found; install ffmpeg | error_code | error | ffmpeg, subprocess, missing-dependency, multimodal, video |
| Insufficient destination DCP pages: required= | validation | error | disaggregation, dcp, destination-pages, allocation |
| kv-canary: launch_canary_plan_kernels_torch_reference… | validation | error | kv-cache, shape-mismatch, validation, speculative-decoding |
| must be rank , got shape= | validation | error | validation, tensor-shape, rank, minimax-h3 |
| NIXL memory registration failed for state tensors | exception | critical | nixl, memory-registration, hybrid-model, startup, disaggregation |
| .apply should not be called. | exception | error | kv-cache, quantization, api-misuse, not-implemented |
| Comfy NVFP4 layer has incompatible weight/scale shapes: and | validation | error | quantization, nvfp4, shape-mismatch, block-scale, checkpoint-validation |
| Invalid deepep_mode | validation | error | deepep, moe, dispatcher, enum, validation, version-skew |
| Missing aux index for last chunk | exception | error | nixl, disaggregation, chunked-transfer, internal-invariant |
| next_token_logits row count mismatch for DFlash verify… | validation | error | speculative-decoding, dflash, shape-mismatch |
| SGLANG_LOG_SCHEDULER_STATUS_TARGET is set but… | validation | error | sglang, scheduler, metrics, env-var, startup-config |
| Unknown value for : . Available | validation | error | cli, preset, validation, whitelist |
| USPAttention's masked path does not support replicated… | exception | error | attention, sequence-parallel, attention-mask, replicated-tokens, not-implemented |
| beam_width must be at least 1, got | validation | error | sampling-params, beam-search, validation, sglang |
| manages its own checkpoint quantization and does not… | validation | error | quantization, online-quantization, text-encoder, unsupported-operation |
| Failed to set up ModelOpt quantization | exception | error | modelopt, quantization, wrapper-exception, chained-exception, sglang |
| Invalid allowed media domain | validation | error | validation, media, security, ssrf |
| Krea-2 sequence parallelism does not support ragged/padded… | validation | error | sglang, krea-2, sequence-parallelism, multi-prompt, batching, diffusion |
| Layer count and attention attribute count differ | validation | error | mlx, kv-cache, layout, validation |
| --dcp-comm-backend fi_a2a delegates the exchange to… | validation | error | sglang, distributed, dcp, flashinfer, mnnvl, hardware-requirement |
| generated MiniMax H3 MP4 frame rate must be | exception | error | minimax-h3, video, fps, mp4, validation |
| Invalid media URL | validation | error | media, url-validation, network, security |
| triton runner was supported but it's temporarily disabled | error_code | error | deepep, deepgemm, triton, moe, not-implemented, feature-flag |
| group_concurrent_contiguous requires equal-length src/dst… | exception | error | disaggregation, kv-transfer, index-mismatch, validation |
| Not support pos_emb_type | exception | error | kimi-k3, vision, config-validation, not-implemented |
| unknown q-prep variant | validation | error | q8kv8, env-var, variant-selection, mla |
| Could not bind port on any configured address family | exception | error | network, socket, bind |
| graph capture input at | exception | error | cuda, vmm, cuda-graph, pointer-range |
| must stay fp32 after load, got . | validation | error | dtype, fp32, weight-loading, adaln |
| No tool call found | exception | error | agent, tool-call, dispatch, no-op |
| Page size mismatch: prefill server has page_size= | exception | critical | disaggregation, pd-disagg, page-size, config-mismatch, kv-cache |
| sglang.srt.layers.attention.nsa.dequant_k_cache is… | console | warning | deprecation, sglang, nsa, import |
| spec_info is unset in TARGET_VERIFY mode; the extend_*… | exception | error | speculative-decoding, target-verify, spec-info, intel-amx, metadata, sglang |
| Failed to parse JSON content from file | validation | error | json, config, mooncake, ib-devices |
| flashinfer_cudnn expects packed indptrs as a torch.Tensor | exception | error | sglang, vision-transformer, flashinfer-cudnn, type-validation, cu-seqlens |
| flashinfer_cutedsl FP4 MoE only supports DeepEP low_latency… | validation | error | moe, flashinfer, cutedsl, fp4, deepep-mode, server-args |
| get_input_embeddings() is not available in encoder-only mode | exception | error | kimi-k3, encoder-only, attribute-error, embeddings |
| Invalid forward mode | validation | error | cuda-graph, forward-mode, hybrid-linear-attention, capture, sglang |
| n must be a positive integer | validation | error | sglang, value-error, config-validation, model-loading, inkling |
| out_channels must be divisible by tp_size for TP-sharded… | exception | error | tp-sharding, divisibility, ltx-2, config-validation, tensor-parallel |
| Packed CUDA VMM features must be reconstructed before… | exception | error | cuda, vmm, packed-tensors, api-misuse |
| pad_masks and att_masks must be [batch, seq] | validation | error | pi05, attention-mask, tensor-rank, input-validation |
| processor image placeholder count mismatch: processor= | validation | error | multimodal, processor-override, tokenization, placeholder-mismatch |
| sequence_lengths should be prepared for vision… | exception | error | sglang, vision-transformer, flashinfer-cudnn, kwargs-validation, multimodal |
| Cannot rewrap thinking history: no reasoning detector is… | exception | error | reasoning, chat-history, server-config |
| Cosmos3 batched action input requires one prompt per image… | validation | error | cosmos3, batching, prompt-validation, cardinality-mismatch |
| CUDA VMM recycler did not stop | exception | error | cuda, vmm, shutdown, threading, timeout |
| kv-canary: SWA slot is outside full_to_swa_index_mapping… | validation | error | kv-cache, index-out-of-range, sliding-window, lut |
| MIXED_PRECISION checkpoint has no NVFP4 layers to… | exception | error | quantization, quark, mxfp4, mixed-precision, nvfp4, python |
| Rank-local FSDP shard produced for non-DTensor parameter | exception | error | fsdp, dtensor, distributed, weight-loading, sharding |
| A quantized checkpoint requires an in-tree native encoder… | validation | error | quantization, architecture-unsupported, text-encoder, tensor-parallel |
| Invalid latent grid for memory RoPE | validation | error | joy-echo, memory, rope, latent-grid, video |
| num_bits must be 4 or 8, got | validation | error | marlin, awq, gptq, bit-width, kernel-support |
| The Triton WNA16 MoE backend only supports symmetric INT4… | validation | error | quantization, moe, triton, int4, compressed-tensors |
| total_pool_size must be positive | validation | error | cuda-ipc, memory-pool, config-validation, multimodal-transport |
| Type must match: != | exception | error | nvfp4, quantization, dtype-mismatch, cutlass, gpu-kernel |
| Unable to infer Z-Image caption length for rotary embeddings | validation | error | z-image, rotary-embeddings, batch-state, multimodal |
| Unsupported nvfp4_gemm_swiglu_nvfp4_quant configuration… | validation | error | nvfp4, cutlass, tiling, shape-alignment, sm100 |
| combined_history=True requires direction=0 (bidi) | validation | error | gdn, linear-attention, triton, argument-validation |
| Component resolved to layerwise-offload, but its loaded… | validation | error | memory, offload, config, component-residency, startup |
| draft_token_num must be positive, got | validation | error | speculative-decoding, dflash, argument-validation |
| `dt_bias` must have elements (got ). | exception | error | pytorch, tensor-shape, kda, linear-attention, validation |
| Invalid Kimi image grid metadata | validation | error | multimodal, kimi, grid-metadata, shape-validation |
| The number of placeholders does not match the number of… | exception | error | multimodal, prompt-template, placeholder-mismatch, step3-vl |
| Unexpected q_proj output shape | error_code | error | mlx, decode, attention, shape-mismatch |
| Unsupported config file | validation | error | hicache, config, file-format, extension |
| Unsupported Kimi-K3 vision attention backend | exception | error | kimi-k3, attention-backend, env-var, startup-validation |