sgl-project/sglang
Documented errors, page 13 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| Kimi expert-pack role sizes do not match object bytes | validation | critical | kimi, moe, expert-pack, size-consistency |
| MiniMax H3 final output ffprobe timed out after 30 seconds | exception | error | minimax-h3, ffprobe, timeout, subprocess |
| name must be provided | validation | error | sglang, reasoning, naming, validation |
| No cached files for match | exception | error | offline-cache, snapshot-download, huggingface, modelscope |
| op has multiple backends usable on device ( ); pass… | validation | error | kernels, backend-selection, ambiguity, sglang |
| has invalid port number . Valid TCP port range is 0- . | validation | error | network, port, validation |
| --prefill-only-disable-kv-cache requires… | validation | error | sglang, chunked-prefill, kv-cache, flashattention, server-args |
| thinking.budget_tokens is not allowed when thinking.type is… | validation | error | anthropic, thinking, adaptive, validation, request-validation |
| topk_ids must be int32, got | exception | error | moe, dtype, int32, triton |
| varlen KDA requires batch size 1 | exception | error | kda, helion, varlen, batch-shape |
| --enable-session-radix-cache requires UnifiedRadixCache… | validation | error | session-cache, unified-cache, server-args, value-error |
| expected a tensor with at least one dimension | validation | error | nvfp4, weight-loading, shape-validation |
| expert-pack direct I/O is unavailable on this platform | exception | error | moe, expert-pack, direct-io, platform-support, configuration |
| k_pool has incompatible shape | validation | error | shape-validation, kv-cache, gqa, metal, sgl-kernel |
| kv_scales supplied but unified_kv is | validation | error | attention, fp8, quantization, kv-cache, triton |
| MiniMax H3 audio decode failed on rank 0 | error_code | critical | minimax-h3, audio-decode, distributed, rank0 |
| out is only supported for forward-only inference | validation | error | autograd, attention, inference, flash-attention |
| replica broadcast of batch.extra | error_code | error | minimax-h3, replica-broadcast, distributed, distributed-communication |
| speculative miss_src/miss_dst must have shape [batch, >=… | validation | error | hisparse, shape-validation, miss-plan |
| {template_error}{suffix} | validation | error | jinja, embeddings, chat-template |
| The hpc_ops MoE runner backend does not support fused… | validation | error | sglang, moe, hpc-ops, shared-experts, config-validation |
| Unexpected arguments: `**rope_kwargs` and `config` are… | exception | error | mimo-audio, rope, api-misuse, argument-validation |
| unsupported input for modulate_scale_shift CUDA | exception | error | cuda, modulate, fallback, input-validation |
| Unsupported PD DCP topology | exception | critical | disaggregation, dcp, topology, parallelism |
| Varlen USPAttention does not support ring parallelism yet. | exception | error | attention, varlen, ring-parallelism, sequence-parallel, not-implemented |
| adaln out_features mismatch | validation | error | minimax-h3, adaln, shape-mismatch, config-validation |
| All tensors must have the same data type | validation | error | cuda, dtype, flash-attention, cutlass, type-mismatch |
| conditions for task must include one or two ordered… | validation | error | minimax-h3, keyframes, fl2va, ref2va, request-validation |
| Dynamic LoRA currently supports only one adapter per… | validation | error | lora, merge-mode, multi-adapter |
| every tool call must be a JSON object with a 'name' | exception | error | tool-calls, json-parsing, validation |
| f"Cannot parse checkpoint quantization for | exception | error | quantization, bitsandbytes, config, model-loading |
| Failed to get model info | error_code | error | network, http-client, model-info, sglang-server |
| Invalid gate_up_proj shape for | exception | critical | lfm2, moe, weight-loading, shape-validation |
| Kimi K3 uses its model-native structural tag implementation | exception | error | kimi-k3, function-call, not-implemented, structural-tag |
| must start with 0 and contain at least one sequence | validation | error | npu, ascend, varlen, packed-sequences, validation |
| No compatible attention backend is available | validation | critical | attention-backend, no-backend-available, multimodal, sglang |
| positions/slots must have one entry per token | validation | error | shape-validation, positions, kv-cache-slots, rope, sgl-kernel |
| QKV last dimensions must be contiguous | validation | error | triton, rope, tensor-contiguity, hunyuan |
| Quest query hidden size | validation | error | quest, sparse-attention, head-dim, shape-mismatch, value-error |
| SANA-WM does not support temporal sequence parallelism yet… | validation | error | sana-wm, sequence-parallelism, unsupported-feature, sglang |
| Speculative algorithm | validation | error | speculative-decoding, overlap-scheduling, configuration |
| task is not served by MiniMax H3 partition ; supported tasks | validation | critical | minimax-h3, model-index, config-validation |
| tool call function arguments must decode to an object | validation | error | inkling, tool-calls, json-decode, arguments |
| Unknown CacheAware Policy | validation | error | scheduling, config, enum-validation |
| Unsupported LTX-2 RoPE type | validation | error | sana-wm, ltx2, rope, attention-config, valueerror |
| unsupported MiniMax H3 task | validation | error | minimax-h3, task-validation, unknown-task |
| Use get_key_buffer instead. | exception | error | deepseek-v4, memory-pool, api-misuse, not-implemented |
| attn_sink must be contiguous | validation | error | mla, sparse-attention, tensor-validation, contiguity |
| Cannot repeat tensor with batch= | exception | error | ltx-2, batch-dim, cfg-guidance, shape-validation |
| expert-pack alignment is invalid | exception | critical | moe, expert-pack, binary-format, alignment, corruption |
| fastokens failed to load tokenizer for | exception | error | tokenizer, sglang, backend |
| Found safetensors files in and no index to disambiguate… | validation | error | safetensors, sharding, ambiguous-checkpoint |
| Height and width must be provided | validation | error | latent-preparation, height-width, missing-dimensions, validation |
| Host-pool retraction does not support Mamba models. | validation | error | disaggregation, retraction, mamba, hybrid-ssm, unified-cache, value-error |
| InklingNvfp4MoEMethod is the dense shared-expert method… | exception | error | quantization, nvfp4, moe, not-implemented, inkling |
| Invalid STA_param | validation | error | sliding-tile-attention, metadata, index-out-of-range, prefix-parsing |
| Kimi K3 tool parameters 'additionalProperties' must be a… | validation | error | kimi-k3, json-schema, additional-properties, validation |
| kv-canary: cuda_graph_max_bs must be non-negative, got | exception | error | kv-canary, validation, cuda-graph, value-error |
| Mismatched batch sizes: mixed_qkv.shape[0]= | exception | error | fla, fused-recurrent, batch-mismatch |
| No embedding available for Mooncake GPU-direct transfer | http | critical | mooncake, disaggregation, encoder, gpu-direct |
| nvImageCodec returned an invalid JPEG tensor: shape= | exception | error | nvjpeg, image-decoding, tensor-shape, multimodal, sglang |
| Previous frame size does not match current delta payload | error_code | error | rocm, allreduce, deterministic, alignment, float32 |
| Shard id with multiple indices is not supported in… | validation | error | weight-loading, shard-id, merged-column, api-version |
| The number of image placeholders does not match… | validation | error | kimi, multimodal, placeholder-count, validation |
| The package `amd-quark` is required to use MX-FP4 models… | error_code | error | mxfp4, amd, quark, missing-dependency, quantization |
| Unsupported DisaggregationMode | exception | critical | ascend, disaggregation, config, init |
| All inputs must be on the same device. | exception | error | fla, fused-recurrent, device-mismatch, multi-gpu |
| alloc_req_slots runs out of memory. Please set a smaller… | exception | critical | sglang, memory, request-pool, capacity |
| Block sparse tensors | validation | error | block-sparse, block-size, configuration, attention |
| Comfy W4A8 layer needs an F32[16] codebook | validation | error | quantization, codebook, dtype-mismatch, w4a8 |
| CUDA error | exception | critical | cuda, gpu, memory-allocation, driver |
| Failed to generate MiniMax-H3 video | exception | error | comfyui, sgldiffusion, rpc, video-generation |
| Failed to load HiCache native hash extension | exception | error | native-extension, build-failure, openssl, toolchain, cpp-extension |
| --hicache-host-memory-mode buffer_only is only implemented… | validation | error | hicache, buffer-only, unified-cache, server-args, value-error |
| Kimi-K3 processor feature length does not match image grids | validation | error | multimodal, kimi-k3, processor-mismatch, shape-mismatch |
| kv-canary: RealKvSource.tensor dim-1 byte width must be a… | validation | error | kv-cache, alignment, strides, validation |
| layer_id= does not have an index V cache (either dense, or… | validation | error | kv-cache, sparse-attention, layer-id, mapping, minimax, sglang |
| MiniMaxH3TextEncodingStage requires the pipeline processor… | validation | error | minimax-h3, pipeline-components, model-index, init-validation |
| MP mode requires --lmcache-config-file (the YAML supplies… | validation | error | lmcache, mp-mode, missing-config, server-args |
| multiple Cargo packages under | validation | error | rust, discovery, ambiguity, metadata |
| and its host copy must have the same length | validation | error | npu, ascend, varlen, host-device-sync, validation |
| The size of ( ) is ( ), but you haven't specified the order… | exception | critical | distributed, parallelism, config |
| Unknown input type | exception | error | harmony, input-validation, discriminator, sglang |
| unsupported dtype for causal Conv3D cat/pad | exception | error | cuda, dtype, conv3d, diffusion |
| Unsupported dtype in flattened_bucket metadata | validation | error | weights-update, dtype, flattened-bucket, multimodal |
| [weight_cache] of model weights is not supported while the… | error_code | error | weight-cache, cuda-ipc, memory-management, feature-incompatibility |
| world_size ( ) is less than tensor_parallel_degree ( ) x… | exception | critical | distributed, parallelism, configuration, startup |
| Z-Image caption tensor must have rank 2 or 3 | validation | error | tensor-shape, cuda-graph, z-image, multimodal, padding |
| `a` and `b` must be 2D tensors | validation | error | fla, fused-recurrent, shape-validation |
| A dict of processors was passed, but the number of… | validation | error | attention, processor, autoencoder, value-error |
| [Elastic EP] WORLD MLP sync dp_size exceeds WORLD size… | panic | critical | distributed, elastic-ep, dp-size, topology |
| H3 conditioning projection contains unsupported tensors | validation | critical | minimax-h3, conditioning-projection, unsupported-tensors, checkpoint |
| JoyEchoPipeline requires JoyEchoPipelineConfig, got | validation | error | joyecho, pipeline-stages, type-mismatch, sglang |
| contains occurrences of embed_override_token_id= , but… | validation | error | embeddings, overrides, count-mismatch, validation |
| --load-diffusion-decoder was requested, but this checkpoint… | validation | error | ltx2, diffusion-decoder, checkpoint-manifest, sglang |
| Out of memory. Try to lower your batch size.\nTry to… | exception | critical | sglang, memory, kv-cache, allocation |
| Output N= must be divisible by sf_vec_size= | validation | error | nvfp4, shape-alignment, scale-factor, quantization |
| quantize_and_serve functionality is currently disabled due… | exception | error | quantization, modelopt, not-implemented, feature-disabled, sglang |
| Real-ESRGAN weight file | exception | error | realesrgan, checkpoint, architecture, postprocess |
| rope_pool_fused expects positions/slots to be 1-D | validation | error | shape-validation, rope, metal, positions, sgl-kernel |