sgl-project/sglang
Documented errors, page 7 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| DSV4 target and draft pools must share SWA ring geometry… | validation | error | disaggregation, deepseek-v4, swa, geometry-mismatch |
| Inkling shared-sink gate-up A and down B must use the same… | validation | error | sglang, lora, shape-validation, consistency-check, moe, inkling |
| Unknown stop criteria | exception | error | configuration, simulation, schedule-simulator, sglang |
| USPAttention masked path supports ring parallelism only for… | exception | error | attention, ring-parallelism, attention-mask, batch-size, fa-backend, not-implemented |
| VLA action expert should not share the prefix TP layout… | validation | error | pi05, vla, parallelism-strategy, tp, config, sglang |
| [weight_cache: ] quantization method is not verified for… | validation | error | quantization, weight-cache, ipc, config |
| --dcp-replicate-q-proj only applies to the a2a/fi_a2a DCP… | validation | error | sglang, distributed, dcp, argument-validation, incompatible-flags |
| DeepSeek-V4 FP4 experts require torch.float4_e2m1fn_x2… | exception | critical | quantization, fp4, pytorch-version, deepseek |
| Error in stream_executor | console | error | interpreter, stream-executor, async, frontend, nested-exception |
| KDA num_heads ( ) must be divisible by shard tp_size ( ) | exception | error | kimi-linear, tensor-parallel, divisibility, startup-validation |
| routed-expert tensors were not loaded (sample: ). Expected… | exception | critical | laguna, moe, weight-loading, missing-weights |
| MiniMax H3 latent preparation requires pre-queue resolved… | validation | error | minimax-h3, temporal-dimensions, plan-validation |
| Mismatched ModelSlim quantization for W13 in layer | validation | error | modelslim, quantization, moe, fused-weights, config-mismatch |
| override_server_args: unknown ServerArgs field(s) | validation | error | config, override, typo, sglang |
| acknowledgements support one consumer or the complete… | validation | error | cuda-ipc, acknowledgement, protocol-validation, multimodal-transport |
| Server failed to start within the timeout period. | exception | error | frontend, timeout, server-startup, spawn, sglang |
| The current scheduler class | validation | error | qwen-image, diffusers, scheduler, timesteps, scheduler-unsupported |
| The original encoder only has | validation | error | siglip, vision-encoder, layer-override |
| uniform_samples_for_final_sampling shape mismatch. Expected | validation | error | sglang, speculative-decoding, dflash, shape-mismatch, rng |
| Unknown request_id | validation | error | disaggregation, request-state, unknown-id, stale-request |
| video rows must be divisible by t*h*w for latent_shape= | validation | error | validation, tensor-shape, unpatchify, row-count, minimax-h3 |
| Waiting for main node timeout! | exception | error | sglang, deep-gemm, timeout, multi-node, compilation |
| Anthropic thinking is not supported for reasoning parser | validation | error | anthropic-api, reasoning, unsupported-feature |
| Comfy W4A8 layer has incompatible weight/scale shapes: , … | validation | error | quantization, shape-mismatch, w4a8, group-size |
| Cosmos3 action batch size | validation | error | cosmos3, batching, capacity-limit, server-args |
| Cosmos3 action endpoint supports action_mode='policy' or… | validation | error | cosmos3, action-mode, enum-validation |
| DeepEP v2 MoE has not validated fused shared experts yet… | validation | error | deepep, shared-experts, fusion, moe, server-args |
| Failed to load LoRA adapter | exception | critical | lora, startup, load-failure, sglang |
| Invalid LoRA merge mode | validation | error | lora, config-validation, diffusion-pipeline |
| repetition_penalty must be in (0, 2] (1.0 = no penalty), got | validation | error | sampling-params, repetition-penalty, validation, sglang |
| token_ids_logprob must be a flat list of integers. | validation | error | sglang, validation, logprob, request-validation |
| attn_res_fused_tma requires SM100+ excluding SM12x; SM | error_code | error | cuda, tma, sm100, architecture-check, jit |
| is not bundled or cached, and Rust extension build mode is… | exception | error | rust-extension, missing-module, build-cache, env-var, sglang |
| failed to move modules to | error_code | critical | cuda-oom, device-movement, rollback, memory-offload |
| hd256 forward varlen expects k rank 3 or 5, got rank | exception | error | cuda, flash-attention, tensor-shape, sm100, cutlass |
| Hunyuan3D Paint expects square latents and a matching view… | validation | error | runtime, shape-mismatch, multiview, diffusion |
| `initial_state` must be a 4D tensor | validation | error | kda, helion, shape-validation, tensor-ndim |
| No call message found for | exception | error | harmony, tool-calls, conversation-history, validation, sglang |
| plucker_emb token count | validation | error | sana-wm, shape-mismatch, camera-embedding, validation |
| Unsupported input shape | validation | error | fla, cumsum, shape-validation, head-first |
| `a`/`b` must have shape [B, HV] with HV= | exception | error | fla, fused-recurrent, head-mismatch, tensor-parallel |
| Cannot put argument inside a f-string. This is not… | validation | error | sglang, f-string, tracer, typeerror |
| dynamic_batch_seeds must be a list with one seed per prompt | validation | error | seed, batch-validation, input-validation |
| expert-pack is not identity triplet layout | exception | critical | moe, expert-pack, binary-format, flags, layout-mismatch |
| .position_ids is required | validation | error | input-validation, position-ids, multimodal, forward |
| LingBot causal sequence sharding currently requires… | exception | error | not-implemented, sequence-parallelism, kv-cache, lingbot |
| MiMoV2 fused qkv_proj checkpoint is TP= | exception | critical | mimo-v2, tensor-parallel, weight-loading, checkpoint-layout |
| Only CUDA and MUSA support GGUF quantization currently. | console | warning | sglang, gguf, rocm, quantization, platform-support |
| Triton is not supported on current platform, roll back to… | console | warning | triton, cuda, device-detection, cpu-fallback, fla |
| Unknown router | exception | error | cli-arguments, simulation, schedule-simulator, sglang |
| Unsupported config option | validation | error | sglang, yaml, config, argparse, unsupported-option |
| Unsupported model type | validation | error | comfyui, diffusion, model-loading, unsupported-architecture |
| Unsupported runner backend | exception | error | quantization, moe, dispatch, version-skew |
| A scheme must be defined for each layer | validation | error | modelslim, quantization, scheme-uninitialized, runtime |
| Cannot use the fast tokenizer in slow tokenizer mode. | validation | error | tokenizer, configuration, sglang |
| Decode out of memory. Try to lower your batch size.\nTry to… | exception | critical | sglang, memory, decode, paged-kv |
| f"Adapter weights at | exception | error | model-loading, state-dict, adapter, weight-mismatch |
| image_mode=' ' is not supported with multiple images (got… | exception | error | multimodal, ocr, multi-image, config-validation |
| Incorrect type of pixel values. Got type | validation | error | multimodal, vision, type-validation, deepseek-ocr |
| Inkling shared-sink LoRA expert count does not match | validation | error | sglang, lora, shape-validation, moe, shared-experts, inkling |
| kernel dispatch requires at least one tensor argument | validation | error | kernel-dispatch, api-misuse, assertion, triton |
| extents exceed BUMPARENA_MAX_EXTENTS ( ) | validation | error | cuda, vmm, memory, capacity-limit |
| LTX-2 token latents seq_len= | validation | error | ltx-2, video-generation, sequence-parallelism, tensor-shape-mismatch |
| Neighborhood attention requires each dim to be at least its… | validation | error | attention, input-shape, video-generation, ltx-2 |
| No DeepSeek-V4 checkpoint mapping for | panic | error | gguf, deepseek, unmapped-tensors, weight-mapping |
| No files found in HF repo | exception | error | huggingface, checksum, model-files |
| `out` must have shape | exception | error | fla, fused-recurrent, output-buffer, shape-validation |
| PD peers must connect matching DCP ranks, got prefill= | exception | critical | disaggregation, dcp, parallelism, bootstrap |
| requires --distilled-lora-path… | validation | error | ltx2, distilled-lora, missing-path, component, sglang |
| SGLANG_USE_MLX requires an available PyTorch MPS device | exception | error | mlx, mps, torch, apple-silicon, device-unavailable |
| Slice size exceeds destination token capacity for TP slice… | validation | critical | disaggregation, tensor-parallel, heterogeneous-tp, kv-cache |
| torch.distributed must be initialised before… | exception | critical | disaggregation, multi-node, torch-distributed, initialization-order |
| Unsupported activation type | validation | error | phi4, activation, glu, config-validation |
| /v1/models | http | error | radix-tree, hicache, not-implemented, mem-cache |
| Cosmos3 policy input requires an observation image | validation | error | cosmos3, policy-mode, missing-image, input-validation |
| --enable-linear-replayssm-spec requires a linear draft chain | validation | error | sglang, replayssm, speculative-decoding, eagle, config-conflict |
| For Fused MoE layers, only | validation | error | quantization, moe, mxint4, compressed-tensors, model-config |
| Invalid spatial patching for packed token latents. Expected… | validation | error | ltx-2, video-generation, resolution-validation, divisibility |
| JoyImage conditioning batch mismatch: hidden_states batch= | exception | error | batch-mismatch, cfg-conditioning, joyimage, multimodal |
| MiniMax H3 TP-local heads | validation | critical | ulysses, sequence-parallel, divisibility, tensor-parallel |
| Ngram speculative decoding only supports CUDA or CPU… | validation | error | speculative-decoding, ngram, device-support, rocm, server-args |
| PD decode DCP requires an MLA or hybrid-MLA KV pool. | exception | critical | disaggregation, dcp, context-parallel, mla, unsupported-feature |
| rope.inv_freq must stay fp32 after load, got | validation | error | dtype, fp32, rope, weight-loading |
| is missing duck-typed methods from SpeculativeAlgorithm: … | validation | error | speculative-decoding, plugin-api, duck-typing |
| The pointers must be multiple of 16 bytes. | validation | error | sglang, cuda-kernel, alignment, silu, shape-validation |
| Unexpected return type from apply_chat_template | error_code | error | cosmos3, tokenizer, apply-chat-template, transformers, type-mismatch |
| attention group range | validation | error | cuda, vmm, parallelism, world-size-mismatch |
| --enable-unified-memory with PD disaggregation does not… | validation | error | unified-memory, pd-disaggregation, hybrid-swa, kv-cache, boot-config |
| f"recurrent_kda state pool breaks the compiled stride… | validation | error | sglang, kda, alignment, memory-layout, flashinfer |
| Failed to load LoRA adapter | validation | error | lora, duplicate, adapter, sglang |
| Got but expected positional dim | validation | error | config, validation, rope, diffusion |
| invalid Kimi-K3 attention-residual target | validation | error | kimi-k3, gguf, weight-conversion, validation |
| Invalid mode: , must be one of 'write', 'read', 'skip | validation | error | sglang, glm-image, kv-cache, mode-validation, enum-value |
| KVTransferError(self.bootstrap_room, failure_reason) | exception | critical | sglang, nixl, kv-transfer, disaggregation, distributed-inference |
| LoRA adapter ' ' contains target modules that are not… | validation | error | lora, target-modules, subset-validation, sglang |
| Mooncake's batch transfer requires mooncake-transfer-engine… | exception | error | mooncake, version-mismatch, batch-transfer, upgrade-required, sglang |
| MXFP8 dense GEMM requested via… | error_code | error | quantization, mxfp8, gemm-backend, hardware-compatibility, flashinfer |
| predict_num_frames supports a single prediction only, got… | validation | error | batching, duration-head, shape-validation, ltx-2 |
| Timeout while waiting for event | error_code | error | timeout, concurrency, async, meta-info |
| Unexpected Ascend TND softmax LSE shape: expected | exception | critical | ascend, npu, flash-attention, lse, shape-mismatch |