sgl-project/sglang
Documented errors, page 25 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| Intern-S2-Mobius requires at least one physical… | exception | critical | model-config, moe, interns2-mobius |
| keyframe_frame_indices must be omitted when keyframe cond… | validation | error | config-validation, minimax-h3, packed-sequence |
| kv-canary: write_offsets_len must be positive, got | validation | error | kv-canary, offsets, argument-validation |
| mask_candidates is required for STA_tuning mode | validation | error | sta, missing-argument, kwargs-validation |
| MXFP8 requires weight_block_size=[1, 32]. | validation | error | quantization, mxfp8, config-validation |
| no safetensors files found in | validation | error | safetensors, checkpoint, file-not-found |
| Not a canonical UMMA_K Layout: Expected MN-size multiple of… | validation | error | cutlass, sm100, layout, alignment |
| NPU packed attention requires q, k, and v in [T, N, D]… | validation | error | npu, ascend, varlen, tensor-layout, shape-validation |
| Only Swizzle<2,5,2> supported for 128B_BASE32B | validation | error | cutlass, sm100, swizzle, shared-memory |
| must be a number | validation | error | minimax-h3, request-validation, float-field, type-check |
| Serialized kitchen_int8 layer | validation | critical | quantization, shape-mismatch, convrot |
| Unsupported weight_bits | validation | error | quantization, auto-round, weight-bits |
| batching config rule cannot set both model and… | validation | error | config, batching, validation, mutually-exclusive |
| Config file not found | console | error | cli, config-file, serve, argument-validation |
| Currently, desc_act (True) is not supported by GPTQ… | validation | error | gptq, npu, ascend, desc-act, act-order |
| gemm_ar: M= outside [1, ] | validation | error | gemm, all-reduce, token-limit, shape-validation |
| --grpc-port is not supported with --encoder-only: encoder… | validation | error | grpc, encoder-only, flag-conflict, config-validation |
| hpc_ops backend does not support speculative decoding for… | validation | error | attention-backend, hpc-ops, speculative-decoding, npu, sglang |
| indices must be int32, got | validation | error | dtype, indices, sparse-attention |
| MiniMax H3 audio decode produced no output payload | error_code | critical | minimax-h3, audio-decode, payload-format |
| MiniMaxH3DiTModel.forward received unexpected kwargs | validation | error | minimax-h3, kwargs-contract, api-signature |
| Missing adaptive runtime state for steps= | exception | error | sglang, speculative-decoding, adaptive-runtime, state-management |
| SANA-Video requires encoder_hidden_states | validation | error | sana-video, encoder-hidden-states, missing-argument |
| scale_shift_table must be CUDA, bf16/fp32, last-dim… | validation | error | device, dtype, contiguity, ltx2 |
| SGLANG_DISAGG_STAGING_BUFFER does not support prefill… | validation | error | sglang, disaggregation, context-parallelism, staging-buffer, config-validation |
| SGLANG_DISAGG_STAGING_BUFFER requires… | validation | error | sglang, pd-disaggregation, staging-buffer, environment-variable, config-validation |
| The MLX tensor bridge supports CPU and MPS tensors, got | validation | error | sglang, mlx, device, torch |
| Unsupported CUTLASS scalar type for accumulator | exception | error | cutlass, sm100, accumulator, dtype |
| VisionFlash4Attention is only available for cuda | exception | error | sglang, vision-transformer, flash-attention-4, platform-support, hardware-compat |
| Weight input_size_per_partition = | validation | error | marlin, tensor-parallel, shape-validation, gptq, awq |
| world_size must be positive and divide global_heads | validation | error | distributed, world-size, divisibility, ulysses |
| Wrong type of stop in sampling parameters. | validation | error | sampling-params, validation, stop-sequences |
| Can't import trtllm_fp8_block_scale_moe from flashinfer… | exception | error | flashinfer, trtllm, moe, fp8, import-error, version-mismatch |
| Comfy INT8 embedding weights support lookup only | validation | error | quantization, int8, embedding, not-implemented |
| comfy_nvfp4 is inferred from per-layer checkpoint metadata… | validation | error | quantization, nvfp4, comfy, api-misuse |
| Expected an index file or a single safetensors shard in | validation | error | modelopt, fp8, safetensors, weight-map |
| f"SGLang KDA state must be [B,H,V,K] with (V,K)= | validation | error | sglang, nvidia, kda, tensor-shape, layout |
| HiSparse destination device indices are not supported by PD… | exception | error | pd-disagg, dcp-relayout, hicache, kv-cache |
| Invalid component residency assignment | validation | error | config, component-residency, type-error, validation |
| min_hits must be positive, got | validation | error | cuda-graph, kimi-k3, config-validation |
| Mode must be one of , got | validation | error | sta, attention, mode-validation, argument-validation |
| Not enough data points for quadratic fitting | validation | error | pipeline-parallel, profiling, chunked-prefill |
| Passing integer indices as timesteps is not supported. | validation | error | scheduler, timestep-type, consistency-model |
| Quanto layer needs a 2D I8 weight, got | validation | error | quantization, quanto, dtype, shape-validation |
| Retry with use_fast=False for | exception | error | tokenizer, transformers, version-mismatch |
| RoPE cos/sin cache is too short for fused KV… | validation | error | rope, context-length, kv-cache |
| Serialized W4A4 checkpoints require CUDA compute capability… | validation | critical | quantization, cuda, compute-capability, gpu-hardware |
| Serialized W4A8 checkpoints require CUDA compute capability… | validation | critical | quantization, cuda, compute-capability, w4a8 |
| Setting SGLANG_LOGGING_CONFIG_PATH from env with | exception | error | logging, configuration, env-var, startup |
| Unregistered UMBP hybrid pool | validation | error | umbp, hybrid-pool, registration |
| Unsupported image processor backend | validation | error | multimodal, image-processor, server-args |
| action string is empty | validation | error | validation, action-string, sana-wm |
| Checkpoint at ' ' is incomplete — the following shard(s)… | exception | critical | safetensors, checkpoint, download, corrupt-checkpoint |
| : cache_head_start required for head slice | validation | error | kv-cache, gqa, head-slicing, argument-validation |
| DeepEP MoE is not supported yet in Step3 model. | validation | error | step3-vl, deepep, moe-backend |
| DeepEP v2 MoE has not implemented the TBO/SBO overlap hooks… | validation | error | deepep, tbo, sbo, overlap, moe, server-args |
| Either num_inference_steps, sigmas, or timesteps must be… | validation | error | scheduler, missing-argument, validation |
| elements must be int or list; got | validation | error | validation, type-error, config |
| = must contain ' }': each rank needs its own path, and a… | validation | error | config, env-var, weight-cache, tensor-parallel |
| GGUF is selected by passing the checkpoint itself, not… | validation | error | gguf, quantization, cli-usage, weights-path |
| --http2-max-concurrent-streams must be between 1 and… | validation | error | http2, server-args, configuration-validation |
| Invalid component residency assignment | validation | error | config, component-residency, type-error, validation |
| Invalid modality ' ' in --limit-mm-data-per-request.Allowed… | validation | error | sglang, multimodal, limit-mm-per-request, server-args, validation |
| Model architectures are not supported for now. Supported… | validation | error | registry, unsupported-architecture, model-resolution |
| No file matching quant type | exception | error | gguf, huggingface, model-loading |
| out must have stride 1 in the last dimension | exception | error | flash-attention, fa4, stride, memory-layout, validation |
| must be a list | validation | error | minimax-h3, conditions, type-mismatch |
| processor_config. must be set for MiMo-V2 | validation | error | config, missing-key, initialization, checkpoint |
| Provide either `prompt` or `prompt_embeds`. Cannot leave… | validation | error | glm-image, prompt-embeds, required-argument, input-validation, valueerror |
| --speculative-draft-window-size must be >=… | validation | error | speculative-decoding, dflash, window-size, argument-validation |
| Unexpected a shape for dense | validation | error | kda, dense-decode, shape-validation, cutedsl |
| Unsupported attention type | validation | critical | attention, config-validation, bailing, hybrid-model |
| Unsupported tcgen05 MMA op kind | exception | error | cutlass, cute, tcgen05, blackwell, mma, unsupported-dtype |
| CuteDSLKDAKernel does not support target_verify | exception | error | sglang, kda, speculative-decoding, not-implemented, linear-attention |
| Dots note omni requires a text prompt for multimodal… | validation | error | multimodal, request-format, typeerror-contract, valueerror |
| is required | validation | error | zmq, endpoint, config, validation |
| flattened_bucket payload must be a dict with… | validation | error | weights-update, flattened-bucket, validation, multimodal |
| indices must have shape | validation | error | shape-mismatch, indices, sparse-attention |
| Inkling only supports group size 16 for NVFP4 | exception | error | quantization, nvfp4, config-mismatch, inkling |
| Intern-S2-Mobius baseline does not support pipeline… | exception | error | pipeline-parallel, unsupported-feature, launch-config |
| kitchen_w4a4 is inferred from per-layer checkpoint… | validation | error | quantization, api-misuse, config |
| kv d_qk must match q d_qk= | validation | error | shape-mismatch, kv-cache, sparse-attention |
| Mamba state layouts differ between prefill and decode | exception | error | pd-disagg, mamba, tp-degree-mismatch, unified-memory |
| MiniMax H3 media material has no positive duration | validation | error | minimax-h3, duration, ffprobe, media |
| must be float32, got | validation | error | mla, quantization, dtype, scale-validation |
| Number of gpus must be positive | console | error | cli, gpu, argument-validation |
| Quanto activation quantization is not supported for | validation | error | quantization, quanto, activations, weight-only |
| tar material URI header must be a JSON object | validation | error | minimax-h3, tar-uri, json, schema, material-io |
| trtllm_mha backend only supports topk = 1 for speculative… | validation | error | speculative-decoding, attention-backend, trtllm-mha, eagle, topk |
| Unsupported Comfy W4A4 format for | validation | critical | quantization, checkpoint, config-validation |
| batchMatch expects state_ids, tokens, and total_lens to… | exception | error | ngram, batch-validation, argument-mismatch |
| BS : candidate_steps must be a list of non-negative ints… | validation | error | sglang, speculative-decoding, config-validation, adaptive |
| cutedsl_bf16_gemm requires an SM10x GPU | error_code | error | gemm, sm100, blackwell, gpu-architecture |
| --disaggregation-decode-enable-radix-cache is incompatible… | validation | error | sglang, pd-disaggregation, radix-cache, fake-backend, config-conflict |
| For INT8 Fused MoE layers, we require channelwise, dynamic… | validation | error | quantization, int8, moe, static-scales |
| H3 conditioning projection W has shape | exception | critical | minimax-h3, conditioning-projection, shape-mismatch, transposed-weight |
| Invalid stacked k_norm_weight shape for fused KV… | validation | error | shape-validation, rmsnorm, speculative-decoding |
| Invalid v_out shape for fused KV materialization: got | validation | error | shape-validation, kv-cache |
| inverse_indices must be | validation | error | minimax-h3, inverse-indices, packed-sequence |
| keyframe denoising requires pixel_frame_indices resolved… | validation | error | minimax-h3, keyframe, index-consistency |