sgl-project/sglang
Documented errors, page 27 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| v shape must match k shape, got | validation | error | shape-validation, rope, kv-projection, metal, sgl-kernel |
| Validate failed: unsupported tensor shape | validation | error | shape, validation, cuda-kernel, diffusion |
| Weight cache daemon for rank | exception | error | daemon, weight-cache, lifecycle, stale-process |
| ComfyFp8Config must be constructed from safetensors layer… | validation | error | quantization, fp8, comfy, api-misuse |
| DFLASH requires --speculative-dflash-block-size to be… | validation | error | speculative-decoding, dflash, argument-validation, block-size |
| --enable-ssl-refresh is not supported with --enable-http2… | validation | error | http2, ssl, certificate-rotation, incompatible-flags |
| Failed to set LoRA adapter | error_code | error | network, http, lora, sgldiffusion |
| hd256 forward non-varlen expects q rank 4 or 5, got rank | exception | error | cuda, attention, tensor-rank, batched, head-dim-256 |
| kv-canary: req_to_token_stride0 must be positive, got | validation | error | kv-canary, stride, argument-validation |
| kv-canary: scatter_req_token_ids offsets must be int64, got | validation | error | kv-cache, dtype-validation, torch |
| No base64 image data found | validation | error | validation, base64, image-decoding, response-schema, sgldiffusion |
| PD decode DCP currently requires chunk cache… | validation | error | sglang, pd-disaggregation, dcp, radix-cache, config-conflict |
| --quantization nvfp4_online is supported only on NVIDIA… | validation | error | sglang, nvfp4, quantization, hardware-gpu, blackwell |
| Quanto auxiliary scale | validation | error | quantization, quanto, scale, scalar-validation |
| Quanto quantization map must be a non-empty object | validation | error | quantization, quanto, json-validation, checkpoint-metadata |
| Resolved GGUF path is not a GGUF file | validation | error | gguf, file-validation, checkpoint |
| RIFE weight file not found | exception | error | rife, weights, file-not-found, postprocess |
| quantization is currently not supported in ROCm. | exception | error | quantization, rocm, amd, hardware-support, sglang |
| `sglang.bench_serving` is deprecated and will be removed in… | console | warning | deprecation, benchmark, serving, future-warning |
| The 'flashkda' KDA prefill backend requires the flash_kda… | exception | error | sglang, flashkda, kda, missing-dependency, pip-install |
| timestep must have shape [B, S, 9 * D] | validation | error | shape, ltx2, adaln, diffusion |
| top_k must be scalar or have one value per row, got | validation | error | sampling, top-k, batch-size-mismatch, validation |
| Unknown schedule_policy | validation | error | startup, config, validation |
| Unsupported data_type | validation | error | quantization, auto-round, data-type |
| Validate failed: not contiguous on dim D. | validation | error | contiguity, stride, validation, cuda-kernel |
| Z-Image batch must contain at least one image latent | validation | error | z-image, empty-batch, validation |
| batching config rule from | validation | error | config, batching, typo, unknown-key |
| Block sparse tensors | validation | error | block-sparse, block-size, alignment |
| Could not resolve a transformer directory from | exception | error | modelopt, fp8, path-resolution |
| CuteDSL masked MoE supports activation 'silu' (gated) or… | validation | error | flashinfer, cutedsl, moe, activation, unsupported-config |
| Diffusion weights must come from a Hugging Face model repo | validation | error | huggingface, weights, wrong-repo-type |
| DSpark with dp attention requires --enable-dp-lm-head. | validation | error | speculative-decoding, dspark, dp-attention, lm-head |
| KV memory descriptors are empty on prefill side | exception | critical | mori, kv-cache, descriptor, initialization-order |
| kv_scales must be fp32, got | validation | error | fp8, kv-cache, dtype, scales |
| logit_bias must has keys in | validation | error | sampling-params, logit-bias, vocab-size, validation, sglang |
| LPLB fused solver requires CUDA tensors; got A on | validation | error | device-mismatch, cuda, lplb |
| max_new_tokens must be non-negative, got | validation | error | multimodal, sampling-params, valueerror |
| must be a 1D int32 or int64 tensor | validation | error | npu, ascend, varlen, dtype-validation, tensor-shape |
| num_attention_heads must be divisible by attention TP | exception | error | tensor-parallel, attention, launch-config |
| O tensor must have 3 or 4 dimensions: (batch, seqlen… | validation | error | flash-attention, shape-mismatch, validation |
| rollout_sde_type must be one of | validation | error | rl-rollout, sde, enum-validation, config |
| subblock_sparse_query_block_mask must be a tensor | validation | error | minimax-h3, sparse-attention, type-validation |
| UMBPHostMemAllocator.alloc | exception | critical | umbp, hugepages, numa, out-of-memory |
| Unsupported IO backend for models with head_dim !=… | validation | error | sglang, mla, hicache, io-backend |
| unused image_token_count entries | validation | error | minimax-h3, ref2va, alignment |
| BF16 fallback patterns are enabled, but… | validation | error | modelopt, fp8, cli, missing-argument |
| cache_dit_params['secondary'] must be a dict, got | validation | error | cache-dit, secondary-cache, type-error, request-validation |
| Comfy NVFP4 layer must request full_precision_matrix_mult | validation | error | quantization, nvfp4, comfy, metadata-validation |
| cu_seqlens_q tensor must be Int32 | validation | error | cuda, dtype, flash-attention, varlen, int32 |
| Expected experts in , got | exception | error | weight-loading, expert-count-mismatch, mobius, moe |
| Failed to setup Mooncake store, error code | exception | critical | mooncake, setup, transfer-engine, network |
| Generation is still in-progress | http | warning | http, async-job, polling, mesh-generation |
| H3 conditioning projection must output width | exception | critical | minimax-h3, conditioning-projection, width-mismatch, config-validation |
| hidden size is outside the supported LTX2 fast-path range | validation | error | hidden-size, kernel-limits, ltx2, triton |
| Invalid attention backend | validation | error | attention-backend, enum-validation, config |
| Invalid port number | validation | error | network, parsing, address |
| Invalid --warmup-mode | validation | error | config, cli, warmup, validation |
| kv must be on q's device | validation | error | device-validation, multi-gpu, tensor-parallel, sparse-mla |
| num nextn_predict_layers is not in the config | validation | critical | mtp, speculative-decoding, bailing, weight-loading |
| reference video has no frames | validation | error | minimax-h3, video, empty-media, ffmpeg |
| SGLANG_INKLING_DEFAULT_REASONING_EFFORT must be in [0.0… | exception | error | environment-variable, inkling, reasoning, server-config, sglang |
| unknown Inkling special token | validation | error | inkling, special-tokens, tokenizer, key-error |
| Unsupported modulate_scale_shift dtype | exception | error | dtype, jit, modulate, cuda |
| Unsupported swizzle triple for UMMA smem descriptor | validation | error | cutlass, sm100, swizzle, shared-memory |
| Video placeholder count does not match video_data | validation | error | multimodal, video, placeholder-mismatch, valueerror |
| video_vae became unavailable during decode | error_code | critical | minimax-h3, vae, race-condition, component-registry |
| auxiliary PP output must contain at least one tensor | exception | error | pipeline-parallel, sampling, empty-payload, validation |
| auxiliary PP output is not a tensor | exception | error | pipeline-parallel, sampling, torch, validation |
| Component residency must use COMPONENT=MODE, got | validation | error | config, component-residency, parsing |
| Config file must contain a dictionary at root level | validation | error | sglang, yaml, config, validation |
| --disaggregation-decode-enable-radix-cache is incompatible… | validation | error | sglang, pd-disaggregation, radix-cache, hisparse, config-conflict |
| Dots note omni expanded prompt is too long | validation | error | multimodal, context-length, prompt-too-long, valueerror |
| expected beta shape | validation | error | kda, mtp, shape-validation, beta-decay |
| extra_key should be a list or a string. | validation | error | sglang, extra-key, type-error, input-validation |
| f"No safetensors files found in | exception | error | model-loading, safetensors, missing-weights |
| failed to build with Cargo | exception | error | rust, cargo, build-failure, compile |
| Humming expected DeepEP FP8 hidden states and group-128… | validation | error | sglang, moe, humming, deepep, fp8, dtype-mismatch, distributed |
| Invalid type for --lora-paths | validation | error | sglang, lora, cli-parsing, type-validation |
| kv-canary: req_to_token_stride0= | validation | error | kv-canary, stride, layout |
| MiniMax H3 material localization completed without cached… | validation | error | internal-state, minimax-h3, probe |
| Regular expression is not supported in the Anthropic… | console | warning | regex, anthropic, structured-output, constraint-dropped, warning |
| renorm kernels require a CUDA/HIP tensor | validation | error | sampling, top-p, cuda, device-placement |
| setup_metal.py only supports macOS (Apple Silicon). | console | error | metal, build, platform-check, apple-silicon |
| --sidecar-args requires --sidecar. | validation | error | grpc, sidecar, config-validation, server-args |
| Unsupported activation type | validation | error | realesrgan, activation, config, postprocess |
| Unsupported Comfy W4A8 format for | validation | critical | quantization, checkpoint, validation, w4a8 |
| --asr-max-concurrent-sessions must be positive | validation | error | sglang, asr, concurrency, server-args, validation |
| Block sparsity + sheared bias is not supported on SM90 | exception | error | flash-attention, block-sparse, bias, sm90, unsupported-feature |
| Cannot find corresponding multimodal processor registered… | validation | error | multimodal, llava, model-loading, unsupported-model |
| conditions[ ].frame_index must be -1 or in | validation | error | minimax-h3, frame-index, out-of-range |
| dim must be divisible by head_dim . | validation | error | attention, model-config, divisibility, ltx-2 |
| Explicit vision TP cannot be combined with data parallel | validation | error | qwen3-vl, vision, data-parallel, tensor-parallel |
| External ngram corpus exceeds the remaining token budget | validation | error | resource-limit, ngram, corpus-loading |
| extra_config['ssd_durability_mode'] must be one of: strict… | validation | error | umbp, config-validation, durability |
| FA4 CuTe FP8 backward is not supported yet (forward-only). | exception | error | flash-attention, fp8, autograd, not-implemented |
| Fitted quadratic coefficient a= | validation | error | pipeline-parallel, profiling, data-quality |
| Kimi expert-pack header is truncated | validation | critical | kimi, moe, expert-pack, truncated-file |
| kv-canary: kv_canary must be one of none/log/raise, got | exception | error | kv-canary, config, cli-args, enum-validation |
| LSE tensor must have 2 or 3 dimensions: (batch, seqlen… | validation | error | flash-attention, shape-mismatch, validation |
| MiniMax H3 does not support enable_upscaling: the accepted… | validation | error | minimax-h3, upscaling, video-generation, validation |