sgl-project/sglang

Documented errors, page 13 of 32. Back to sgl-project/sglang

Code / MessageTypeSeverityTags
Kimi expert-pack role sizes do not match object bytes
validation critical kimi, moe, expert-pack, size-consistency
MiniMax H3 final output ffprobe timed out after 30 seconds
exception error minimax-h3, ffprobe, timeout, subprocess
name must be provided
validation error sglang, reasoning, naming, validation
No cached files for match
exception error offline-cache, snapshot-download, huggingface, modelscope
op has multiple backends usable on device ( ); pass…
validation error kernels, backend-selection, ambiguity, sglang
has invalid port number . Valid TCP port range is 0- .
validation error network, port, validation
--prefill-only-disable-kv-cache requires…
validation error sglang, chunked-prefill, kv-cache, flashattention, server-args
thinking.budget_tokens is not allowed when thinking.type is…
validation error anthropic, thinking, adaptive, validation, request-validation
topk_ids must be int32, got
exception error moe, dtype, int32, triton
varlen KDA requires batch size 1
exception error kda, helion, varlen, batch-shape
--enable-session-radix-cache requires UnifiedRadixCache…
validation error session-cache, unified-cache, server-args, value-error
expected a tensor with at least one dimension
validation error nvfp4, weight-loading, shape-validation
expert-pack direct I/O is unavailable on this platform
exception error moe, expert-pack, direct-io, platform-support, configuration
k_pool has incompatible shape
validation error shape-validation, kv-cache, gqa, metal, sgl-kernel
kv_scales supplied but unified_kv is
validation error attention, fp8, quantization, kv-cache, triton
MiniMax H3 audio decode failed on rank 0
error_code critical minimax-h3, audio-decode, distributed, rank0
out is only supported for forward-only inference
validation error autograd, attention, inference, flash-attention
replica broadcast of batch.extra
error_code error minimax-h3, replica-broadcast, distributed, distributed-communication
speculative miss_src/miss_dst must have shape [batch, >=…
validation error hisparse, shape-validation, miss-plan
{template_error}{suffix}
validation error jinja, embeddings, chat-template
The hpc_ops MoE runner backend does not support fused…
validation error sglang, moe, hpc-ops, shared-experts, config-validation
Unexpected arguments: `**rope_kwargs` and `config` are…
exception error mimo-audio, rope, api-misuse, argument-validation
unsupported input for modulate_scale_shift CUDA
exception error cuda, modulate, fallback, input-validation
Unsupported PD DCP topology
exception critical disaggregation, dcp, topology, parallelism
Varlen USPAttention does not support ring parallelism yet.
exception error attention, varlen, ring-parallelism, sequence-parallel, not-implemented
adaln out_features mismatch
validation error minimax-h3, adaln, shape-mismatch, config-validation
All tensors must have the same data type
validation error cuda, dtype, flash-attention, cutlass, type-mismatch
conditions for task must include one or two ordered…
validation error minimax-h3, keyframes, fl2va, ref2va, request-validation
Dynamic LoRA currently supports only one adapter per…
validation error lora, merge-mode, multi-adapter
every tool call must be a JSON object with a 'name'
exception error tool-calls, json-parsing, validation
f"Cannot parse checkpoint quantization for
exception error quantization, bitsandbytes, config, model-loading
Failed to get model info
error_code error network, http-client, model-info, sglang-server
Invalid gate_up_proj shape for
exception critical lfm2, moe, weight-loading, shape-validation
Kimi K3 uses its model-native structural tag implementation
exception error kimi-k3, function-call, not-implemented, structural-tag
must start with 0 and contain at least one sequence
validation error npu, ascend, varlen, packed-sequences, validation
No compatible attention backend is available
validation critical attention-backend, no-backend-available, multimodal, sglang
positions/slots must have one entry per token
validation error shape-validation, positions, kv-cache-slots, rope, sgl-kernel
QKV last dimensions must be contiguous
validation error triton, rope, tensor-contiguity, hunyuan
Quest query hidden size
validation error quest, sparse-attention, head-dim, shape-mismatch, value-error
SANA-WM does not support temporal sequence parallelism yet…
validation error sana-wm, sequence-parallelism, unsupported-feature, sglang
Speculative algorithm
validation error speculative-decoding, overlap-scheduling, configuration
task is not served by MiniMax H3 partition ; supported tasks
validation critical minimax-h3, model-index, config-validation
tool call function arguments must decode to an object
validation error inkling, tool-calls, json-decode, arguments
Unknown CacheAware Policy
validation error scheduling, config, enum-validation
Unsupported LTX-2 RoPE type
validation error sana-wm, ltx2, rope, attention-config, valueerror
unsupported MiniMax H3 task
validation error minimax-h3, task-validation, unknown-task
Use get_key_buffer instead.
exception error deepseek-v4, memory-pool, api-misuse, not-implemented
attn_sink must be contiguous
validation error mla, sparse-attention, tensor-validation, contiguity
Cannot repeat tensor with batch=
exception error ltx-2, batch-dim, cfg-guidance, shape-validation
expert-pack alignment is invalid
exception critical moe, expert-pack, binary-format, alignment, corruption
fastokens failed to load tokenizer for
exception error tokenizer, sglang, backend
Found safetensors files in and no index to disambiguate…
validation error safetensors, sharding, ambiguous-checkpoint
Height and width must be provided
validation error latent-preparation, height-width, missing-dimensions, validation
Host-pool retraction does not support Mamba models.
validation error disaggregation, retraction, mamba, hybrid-ssm, unified-cache, value-error
InklingNvfp4MoEMethod is the dense shared-expert method…
exception error quantization, nvfp4, moe, not-implemented, inkling
Invalid STA_param
validation error sliding-tile-attention, metadata, index-out-of-range, prefix-parsing
Kimi K3 tool parameters 'additionalProperties' must be a…
validation error kimi-k3, json-schema, additional-properties, validation
kv-canary: cuda_graph_max_bs must be non-negative, got
exception error kv-canary, validation, cuda-graph, value-error
Mismatched batch sizes: mixed_qkv.shape[0]=
exception error fla, fused-recurrent, batch-mismatch
No embedding available for Mooncake GPU-direct transfer
http critical mooncake, disaggregation, encoder, gpu-direct
nvImageCodec returned an invalid JPEG tensor: shape=
exception error nvjpeg, image-decoding, tensor-shape, multimodal, sglang
Previous frame size does not match current delta payload
error_code error rocm, allreduce, deterministic, alignment, float32
Shard id with multiple indices is not supported in…
validation error weight-loading, shard-id, merged-column, api-version
The number of image placeholders does not match…
validation error kimi, multimodal, placeholder-count, validation
The package `amd-quark` is required to use MX-FP4 models…
error_code error mxfp4, amd, quark, missing-dependency, quantization
Unsupported DisaggregationMode
exception critical ascend, disaggregation, config, init
All inputs must be on the same device.
exception error fla, fused-recurrent, device-mismatch, multi-gpu
alloc_req_slots runs out of memory. Please set a smaller…
exception critical sglang, memory, request-pool, capacity
Block sparse tensors
validation error block-sparse, block-size, configuration, attention
Comfy W4A8 layer needs an F32[16] codebook
validation error quantization, codebook, dtype-mismatch, w4a8
CUDA error
exception critical cuda, gpu, memory-allocation, driver
Failed to generate MiniMax-H3 video
exception error comfyui, sgldiffusion, rpc, video-generation
Failed to load HiCache native hash extension
exception error native-extension, build-failure, openssl, toolchain, cpp-extension
--hicache-host-memory-mode buffer_only is only implemented…
validation error hicache, buffer-only, unified-cache, server-args, value-error
Kimi-K3 processor feature length does not match image grids
validation error multimodal, kimi-k3, processor-mismatch, shape-mismatch
kv-canary: RealKvSource.tensor dim-1 byte width must be a…
validation error kv-cache, alignment, strides, validation
layer_id= does not have an index V cache (either dense, or…
validation error kv-cache, sparse-attention, layer-id, mapping, minimax, sglang
MiniMaxH3TextEncodingStage requires the pipeline processor…
validation error minimax-h3, pipeline-components, model-index, init-validation
MP mode requires --lmcache-config-file (the YAML supplies…
validation error lmcache, mp-mode, missing-config, server-args
multiple Cargo packages under
validation error rust, discovery, ambiguity, metadata
and its host copy must have the same length
validation error npu, ascend, varlen, host-device-sync, validation
The size of ( ) is ( ), but you haven't specified the order…
exception critical distributed, parallelism, config
Unknown input type
exception error harmony, input-validation, discriminator, sglang
unsupported dtype for causal Conv3D cat/pad
exception error cuda, dtype, conv3d, diffusion
Unsupported dtype in flattened_bucket metadata
validation error weights-update, dtype, flattened-bucket, multimodal
[weight_cache] of model weights is not supported while the…
error_code error weight-cache, cuda-ipc, memory-management, feature-incompatibility
world_size ( ) is less than tensor_parallel_degree ( ) x…
exception critical distributed, parallelism, configuration, startup
Z-Image caption tensor must have rank 2 or 3
validation error tensor-shape, cuda-graph, z-image, multimodal, padding
`a` and `b` must be 2D tensors
validation error fla, fused-recurrent, shape-validation
A dict of processors was passed, but the number of…
validation error attention, processor, autoencoder, value-error
[Elastic EP] WORLD MLP sync dp_size exceeds WORLD size…
panic critical distributed, elastic-ep, dp-size, topology
H3 conditioning projection contains unsupported tensors
validation critical minimax-h3, conditioning-projection, unsupported-tensors, checkpoint
JoyEchoPipeline requires JoyEchoPipelineConfig, got
validation error joyecho, pipeline-stages, type-mismatch, sglang
contains occurrences of embed_override_token_id= , but…
validation error embeddings, overrides, count-mismatch, validation
--load-diffusion-decoder was requested, but this checkpoint…
validation error ltx2, diffusion-decoder, checkpoint-manifest, sglang
Out of memory. Try to lower your batch size.\nTry to…
exception critical sglang, memory, kv-cache, allocation
Output N= must be divisible by sf_vec_size=
validation error nvfp4, shape-alignment, scale-factor, quantization
quantize_and_serve functionality is currently disabled due…
exception error quantization, modelopt, not-implemented, feature-disabled, sglang
Real-ESRGAN weight file
exception error realesrgan, checkpoint, architecture, postprocess
rope_pool_fused expects positions/slots to be 1-D
validation error shape-validation, rope, metal, positions, sgl-kernel