sgl-project/sglang
Documented errors, page 22 of 32. Back to sgl-project/sglang
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| DeepSeek-V4 flashmla_sparse_q8 prefill requires SM90 CUDA… | exception | error | deepseek-v4, flashmla, q8-kv, sm90, gpu-compatibility, sglang |
| Expected > 0 for packed token latents. | validation | error | sglang, ltx-2, sequence-length, latent-packing, sequence-parallelism, video |
| External ngram corpus max tokens must be positive. | validation | error | sglang, ngram, validation, invalid-argument |
| Fused MoE layer ' ' requires consistent quant config for… | validation | error | quantization, auto-round, moe, fused-layer |
| layer_types contains unknown entries | validation | error | config-validation, layer-types, enum-values |
| max_inflight_slices must be positive | validation | error | memory-pool, config-validation, cuda-ipc, multimodal-transport |
| MiniMaxH3AudioEncodingStage direct audio tokenizer encode… | exception | error | minimax-h3, audio-encoding, legacy-api, not-implemented |
| MiniMaxH3DiTModel.forward requires kwarg | exception | error | missing-argument, forward-contract, minimax-h3, kwargs-validation |
| missing `this_sample` as a required keyword argument | validation | error | scheduler, diffusion, api-misuse, unipc |
| model_index.json._minimax_h3.schema_version must be 1 | validation | critical | minimax-h3, model-index, schema-version |
| n must be a positive power of 2, got | exception | error | ascend, dsv4, hadamard, validation, npu, sglang |
| out_tokens buffer too small | exception | error | ngram, buffer-overflow, ffi, tensor-shape |
| Port out of range (0-65535) | validation | error | network, port, parsing |
| return_meta_info is not supported with streaming. Please… | validation | error | streaming, meta-info, openai-api, sglang |
| SANA-WM plucker_embedder is not initialized. | validation | error | sana-wm, plucker-embedder, camera-conditioning, uninitialized-module |
| Sparse Video Gen 2 attention backend requires svg package… | exception | critical | missing-dependency, installation, attention-backend, optional-package |
| The quantization block size of weight must have 2… | validation | error | quantization, fp8, config-validation, shape-validation |
| tokenspeed_mla backend is only supported on Blackwell GPUs… | validation | error | sglang, tokenspeed, mla, hardware-gpu, blackwell |
| topk kernels only support streaming implementation | exception | error | moe, topk, not-implemented, feature-flag |
| Unsupported transfer backend | validation | error | disaggregation, transfer-backend, invalid-argument, configuration |
| Unsupported usp_merge_heads dtype | exception | error | dtype, jit, ulysses, cuda |
| AITer backend requires num_heads | validation | error | aiter, gqa, attention-backend, rocm, shape-validation |
| Assistant tool call function.arguments must be valid JSON. | validation | error | tool-calling, json, chat-template, openai-api, sglang |
| attn_sink must be a CUDA tensor | validation | error | mla, sparse-attention, tensor-validation, cuda-device |
| base64 URI header is too large | validation | error | minimax-h3, base64-uri, header-limit, material-io |
| Batch tokenization is not needed for pre-tokenized… | validation | error | sglang, config, pretokenized, batch-encode |
| content part must be mapping, got | validation | error | inkling, type-error, content-parts |
| Cosmos3CrossAttention requires num_attention_heads… | validation | error | sglang, cosmos3, tensor-parallel, cross-attention |
| cta_n= invalid for use_2cta= : bf16 K-major mma requires N… | validation | error | gemm, blackwell, cutedsl, tile-config, mma |
| --cuda-graph-config[ ].backend= not allowed; allowed | validation | error | cuda-graph, config-validation, backend-selection |
| Expected decoded audio with 1, 2, or 3 dims, got shape= | validation | error | joy-echo, audio, waveform, tensor-shape |
| For INT8 Fused MoE layers, we require channelwise, dynamic… | validation | error | quantization, int8, moe, strategy |
| height/width must be divisible by… | validation | error | ideogram, image-resolution, divisibility, validation |
| HIP does not support fused_marlin_moe currently. | console | warning | sglang, awq, marlin, rocm, moe, platform-support |
| Kimi-K3 image feature must be a torch.Tensor, got | exception | error | kimi-k3, type-error, feature-tensor, multimodal |
| Kimi manifest expert-pack path does not match pack_path | validation | critical | kimi, moe, expert-pack, path-mismatch |
| KV4 is not tested on non-CUDA platforms. | validation | error | sglang, kv4, platform-support, rocm, non-cuda, kv-cache-dtype |
| Not a canonical UMMA_K Layout: Expected stride failure. | validation | error | cutlass, sm100, layout, stride |
| NPU packed attention requires q, k, and v with the same… | validation | error | npu, ascend, dtype, mixed-precision, validation |
| prediction_type given as | validation | error | scheduler, diffusion, prediction-type, unipc |
| --prefill-decode-interval must be non-negative. | validation | error | sglang, scheduler, prefill-decode, server-args, validation |
| received truncated fd header | exception | error | network, unix-socket, protocol-mismatch, fd-passing |
| RoPE config mismatch across layers for fused KV path… | validation | error | config-validation, rope |
| shard_offset and shard_size must be provided | exception | error | weight-loading, column-parallel, shard, checkpoint |
| TGV cute_ext backend supports | validation | error | gemm, tgv, dtype, cutlass |
| topk kernels only support k <= 32 | exception | error | moe, topk, triton, capacity-limit |
| TP size > num_experts . | exception | error | laguna, moe, tensor-parallel, startup-validation |
| Unknown quantization method | exception | error | quantization, typo, unsupported-method, sglang |
| Unsupported query shape for Quest | validation | error | quest, sparse-attention, tensor-rank, shape-mismatch, value-error |
| Validate failed: unsupported dtype | validation | error | dtype, validation, cuda-kernel, diffusion |
| action_mode is set but the loaded Cosmos3 checkpoint has no… | validation | error | cosmos3, action-generation, checkpoint-capability, validation |
| Apple toolchain not found. Install the Xcode Command Line… | console | error | metal, build, xcode, missing-toolchain, macos |
| Blocked unsafe class loading | exception | error | security, pickle, cve, deserialization |
| .frame_index must be -1 or in | validation | error | minimax-h3, frame-index, bounds-check, frame-alignment |
| Expected x.shape[-1] to be even for split rotary, got | validation | error | rope, shape-validation, ltx-2 |
| Failed to load config from | exception | error | config, json, file-io, mooncake |
| File not found | exception | error | mistral, config, file-not-found |
| GGUF diffusion checkpoints require CUDA; the GGML kernels… | validation | error | gguf, cuda, platform-unsupported |
| H3 conditioning projection has neither W nor an MLP | exception | critical | minimax-h3, conditioning-projection, empty-checkpoint |
| mask_block_cnt and mask_block_idx must be provided for… | validation | error | block-sparse, missing-argument, validation |
| MiniMax H3 Qwen3-VL language-layer configuration is… | exception | critical | minimax-h3, config, layer-count, inconsistent-config |
| MiniMax H3 ring parallelism requires the FlashAttention… | exception | error | ring-parallelism, attention-backend, flashattention, minimax-h3, not-implemented |
| missing `last_sample` as a required keyword argument | validation | error | scheduler, diffusion, api-misuse, unipc |
| Model path ' ' is already registered | validation | error | registry, duplicate, value-error, model-path, hf-hub |
| Number of labels must match the number of tag keys. Expected | validation | error | observability, ray, metrics, labels |
| Quanto quantization map entries must be named objects | validation | error | quantization, quanto, json-validation, checkpoint-metadata |
| TP-local heads not divisible by Ulysses world size (total… | validation | error | minimax-h3, ulysses, attention-heads, tensor-parallel |
| Unexpected A_log shape | validation | error | kda, linear-attention, shape-validation, cutedsl |
| unified_kv dtype mismatch: kv= | validation | error | attention, dtype, kv-cache, triton |
| Unsupported CUTLASS scalar type for A/B | exception | error | cutlass, sm100, dtype, tensor-core |
| vis_freqs_cis must be a 2D cos_sin_cache tensor | validation | error | runtime, rope, shape-mismatch, diffusion |
| allow_neg_eigval=True requires 2*sigmoid(beta), which is… | exception | error | kda, beta, flag-conflict, not-implemented |
| Block sparse tensors | validation | error | block-sparse, shape-mismatch, rank |
| Cosmos3CrossAttention requires num_key_value_heads… | validation | error | sglang, cosmos3, tensor-parallel, cross-attention, kv-heads |
| --disaggregation-decode-enable-radix-cache is incompatible… | validation | error | sglang, pd-disaggregation, radix-cache, speculative-decoding, config-conflict |
| hf3fs_fuse.io is not available. Please install the… | exception | critical | hf3fs, importerror, missing-dependency, installation |
| Hunyuan3D SD2.1 checkpoints require linear projection. | validation | error | stable-diffusion, unet, config-validation, hunyuan3d |
| Invalid arch format | validation | error | gpu-arch, validation, configuration, cuda |
| Invalid attention metadata values.Sparsity should be in | validation | error | block-sparse, attention, metadata, validation, value-out-of-range |
| Invalid disaggregation_mode= | validation | error | sglang, disaggregation, pd-disaggregation, argument-validation |
| Invalid image data | validation | error | multimodal, llava, image-input, type-validation |
| Only Float16 or BFloat16 is supported | validation | error | cuda, dtype, flash-attention, cutlass, unsupported-dtype |
| reference audio duration bound must be positive | validation | error | minimax-h3, audio, duration-validation |
| Sparse Video Gen 2 attention does not support causal… | validation | error | causal-mask, attention-backend, config |
| startExternalCorpusLoad called while another load is in… | exception | error | ngram, corpus-loading, concurrency, state-machine |
| task does not belong to partition | validation | error | minimax-h3, partition, task-routing, model-index |
| The input_size of down's weight = | validation | error | quantization, fp8, moe, block-quantization, tensor-parallel, shape-mismatch |
| Type mismatch: != | validation | error | flash-attention, dtype-mismatch, fp8, sm100 |
| Unknown time_shift_type | validation | error | scheduler, invalid-enum-value, config-mutation |
| Unsupported Comfy NVFP4 companion for | validation | error | quantization, nvfp4, comfy, mixed-precision |
| Unsupported Kimi-K3 image channel count | validation | error | multimodal, image-processing, numpy, kimi-k3 |
| Unsupported layout for models with head_dim != v_head_dim… | validation | error | sglang, mla, hicache, io-backend, layout |
| Unsupported output_s dtype | validation | error | quantization, fp8, scale-factor, dtype-validation |
| Weight URL pins revision | validation | error | huggingface, weights, revision-conflict, configuration |
| Cosmos3 observation image arrays must have shape [H, W] or… | validation | error | cosmos3, numpy, shape, image-input |
| CUDA coredump env var | console | info | cuda-coredump, environment-variables, debugging, configuration |
| Currently, gptq_v2 is not supported on CPU with AMX. | validation | error | gptq, checkpoint-format, gptq-v2, cpu, amx |
| expected v/g shape | validation | error | kda, mtp, shape-validation, dspark |
| f"Unknown feature map | validation | error | config, feature-map, validation, constructor |
| must be a JSON object | validation | error | api, json, type-validation, http-400 |