vllm-project/vllm
Documented errors, page 2 of 6. Back to vllm-project/vllm
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| invalid repetition detection params | validation | error | rust, validation, repetition-detection, vllm |
| num_speculative_tokens must be provided with speculative… | exception | error | speculative-decoding, num-speculative-tokens, required-field, pydantic |
| Worker with dp_rank= | http | error | distributed, kv-transfer, mooncake, registration, http-400 |
| Actor not found in communicator group | validation | error | ray, pipeline-parallel, distributed, actor-management |
| Only readers can dequeue | exception | error | distributed, shared-memory, api-misuse |
| Hf3fsClient.check Failed | exception | error | kv-transfer, hf3fs, validation, bounds-check, vllm |
| Failed to get key locations | http | error | hf3fs, metadata-server, http-500, kv-transfer |
| failed to initialize | exception | error | rust, parser, initialization, wrapped-error |
| No valid cudagraph sizes after rounding to multiple of | validation | error | |
| DeepEPv2 requires NCCL GIN (GPU-Initiated Networking). This… | exception | critical | vllm, deepep, nccl, ibgda, infiniband, distributed |
| MoRIIO heterogeneous TP requires replicated KV heads on… | exception | error | tensor-parallel, kv-transfer, gqa, heterogeneous-tp, model-config |
| num_speculative_tokens | exception | error | speculative-decoding, mtp, num-speculative-tokens, validation |
| seq_pooling_type is not set; it should be resolved by… | validation | error | pooling, internal-contract, api-misuse |
| token_id(s) in are out of vocabulary. Vocabulary size | validation | error | rust, validation, token-ids, tokenizer, vllm |
| Invalid keys format | http | error | hf3fs, metadata-server, http-400, request-validation |
| NCCL error | exception | critical | nccl, distributed, network, multi-node |
| Please install mooncake by following the instructions at… | exception | critical | mooncake, import-error, installation, kv-transfer |
| Eager MoRIIO handshake failed for | exception | critical | handshake, tensor-parallel, network, kv-transfer, distributed |
| Group count mismatch: tracker has | exception | error | mooncake, kv-transfer, internal-invariant, scheduler |
| quantization_config is only supported when quantization is… | validation | error | quantization, cli, config, validation |
| Unsupported dynamic dimensions | exception | error | torch-compile, dynamic-shapes, type-validation, vllm |
| Flashinfer allreduce quantization fusion is not supported… | validation | error | distributed, flashinfer, allreduce, quantization, multi-node |
| Unsupported connector type | validation | error | configuration, ec-transfer, factory, unknown-class |
| Unsupported object type | validation | error | serialization, version-mismatch, shared-memory |
| DeepEPv2 communicator properties query failed; networking… | exception | critical | vllm, deepep, nccl, networking, distributed |
| dspark_draft_topk must be between 1 and the draft… | exception | error | speculative-decoding, dspark, topk, range-validation |
| Invalid "device" in mm_processor_kwargs | validation | error | vllm, config, multimodal, device, torch, validation |
| ReasoningConfig: failed to tokenize reasoning strings… | exception | error | reasoning, tokenizer, config, validation |
| multimodal preprocessing error | exception | error | eplb, model-config, moe, vllm |
| online shorthand does not define a spec | validation | error | quantization, config, validation |
| shape_id=' ' requires PyTorch >= 2.11.0 | exception | error | torch-compile, version-compat, dynamic-shapes, vllm |
| Key existence check failed | http | error | hf3fs, metadata-server, http-500, kv-transfer |
| Nested BreakableCUDAGraphCapture is not supported. | exception | error | cuda-graph, compilation, runtime, nested-context |
| Initialize MooncakeDistributedStore failed. | exception | critical | mooncake, rdma, initialization, network, native-error |
| startup handshake timed out while waiting for | exception | error | rust, startup, handshake, timeout, configuration |
| Attempted to free more buffers than allocated | exception | error | hf3fs, buffer-pool, double-free, kv-transfer, resource-management |
| MLA DSpark does not currently support decode context… | exception | error | speculative-decoding, dspark, context-parallelism, config |
| Parameter `normalize` was removed; use `use_activation`… | validation | error | pooling, reward-model, migration, removed-parameter |
| text request ` ` must contain at least one prompt token ID | validation | error | request-validation, prompt, token-ids, api, rust, vllm |
| use_communication_streams is not supported | exception | error | ray, pipeline-parallel, distributed, not-implemented |
| Input to maybe_inplace node is used again after the node… | exception | error | compilation, inplace-ops, custom-ops, vllm |
| To load a model from object storage (S3/GCS/Azure)… | validation | error | model-loading, object-storage, s3, load-format |
| ec_transfer_config must be set for ECConnectorBase | validation | error | configuration, ec-transfer, api-misuse |
| This recipe is a multi-process deployment and cannot be… | exception | error | vllm-recipes, multi-process, prefill-decode, deploy-type |
| messagepack decode failed for | exception | error | rust, deserialization, messagepack, version-skew, protocol |
| Connector uses deprecated 2-argument constructor signature… | validation | error | kv-transfer, api-change, external-connector, vllm |
| failed to read chat template file | exception | error | rust, io, chat-template, file-permissions, configuration |
| transport error | exception | error | rust, zeromq, network, transport |
| allreduce is not supported | exception | error | ray, pipeline-parallel, allreduce, not-implemented |
| Unexpected positional/short argument | exception | error | vllm-recipes, argv-parsing, cli, long-form-options |
| Worker has been garbage collected | exception | critical | elastic-ep, lifetime-management, garbage-collection, race-condition |
| Allocation failed | http | error | kv-transfer, hf3fs, metadata-server, http-500, vllm |
| Invalid value of `_api_process_rank`. Expected to be `-1` or | validation | error | api-server, configuration, internal, parallelism |
| layerwise MLA connector is not supported yet | validation | error | lmcache, mla, layerwise, config, kv-transfer |
| nnodes > 1 can only be set when distributed executor… | validation | critical | parallelism, multi-node, ray, configuration |
| max_num_scheduled_tokens is set to | validation | error | scheduler, speculative-decoding, batched-tokens |
| caching is not supported | exception | error | compilation, cache, custom-backend, api-contract |
| decorated class should have a forward method. | exception | error | torch-compile, decorators, vllm, model-porting |
| io error | exception | error | rust, io, error-handling, diagnostics |
| rejection_sample_method='synthetic' requires exactly one of… | exception | error | speculative-decoding, config, validation |
| --ssl-certfile is required to enable TLS… | validation | error | configuration, tls, ssl, security, rust, vllm, startup |
| messagepack ext value decode failed | exception | error | rust, messagepack, logprobs, deserialization, version-skew |
| --use-replayssm is incompatible with KV connectors (P/D… | validation | error | vllm, config, replayssm, kv-transfer, disaggregation, mamba |
| Attribute not exists in the runnable of cudagraph wrapper | exception | error | cudagraph, attribute-access, wrapper, vllm |
| block_size ( ) must be a multiple of hash_block_size ( ) | validation | error | mooncake, kv-transfer, block-size, config |
| customized max_cudagraph_capture_size | validation | error | cuda-graphs, compilation-config, startup-config |
| failed to build structural tag | exception | error | rust, structural-tag, guided-decoding, grammar, validation |
| A speculative model was provided, but… | exception | error | speculative-decoding, config, num-speculative-tokens, required-field |
| ec_cpu_bytes must be specified in ec_connector_extra_config | validation | error | configuration, ec-transfer, setup |
| The quantization method | validation | error | quantization, deprecation, startup, config |
| chat request stream ` ` closed before terminal output | exception | error | rust, streaming, sse, lifecycle, chat |
| cudagraph_capture_sizes not supported in compile_sizes.This… | exception | error | compilation, piecewise-backend, cudagraph, vllm |
| Error happened when batch testing peer-to-peer access from | exception | error | vllm, cuda, p2p, multi-gpu, diagnostics |
| managed frontend engine count | validation | error | configuration, data-parallel, transport, handshake, rust, vllm, startup |
| proton_profiler_dir must be a local directory | validation | error | profiling, proton, path, configuration |
| external coordinator mode is not implemented yet | exception | error | rust, coordinator, not-implemented, configuration |
| max_logprobs must be non-negative or -1 | validation | error | configuration, logprobs, render, validation, rust, vllm, startup |
| suffix_decoding_max_tree_depth= | exception | error | speculative-decoding, suffix-decoding, range-validation, config |
| Could not collect pip list output (pip or uv module not… | exception | warning | diagnostics, environment, pip, uv, packaging |
| did not return a JSON object. | exception | error | vllm-recipes, api-discovery, json-shape |
| Recipe deploy_type= is not a single-node deployment. A… | exception | error | vllm-recipes, deploy-type, multi-node |
| Unknown ECConnectorRole | validation | error | api-misuse, ec-transfer, enum |
| gpt_oss uses native Harmony output parsing; generic | exception | error | rust, gpt-oss, harmony, parser, configuration |
| request_id does not embed a peer zmq_address and… | exception | error | kv-transfer, config, validation, disaggregated-prefill |
| --use-replayssm supports prefix caching only in align mode… | validation | error | vllm, config, mamba, replayssm, prefix-caching |
| utility call ` ` returned inconsistent results across… | exception | warning | data-parallel, utility-call, consistency, rust |
| Connector does not support HMA but HMA is enabled. Please… | validation | error | kv-transfer, hma, config, vllm |
| harmony output parsing failed | exception | error | rust, gpt-oss, harmony, parsing, streaming |
| 'mm_encoder_fp8_scale_save_path' cannot be used with… | validation | error | vllm, config, multimodal, fp8, mutually-exclusive, validation |
| text request ` ` stop strings cannot be empty | validation | error | request-validation, stop-strings, sampling, api, rust, vllm |
| tool response messages require a tool_call_id; use… | panic | critical | rust, panic, chat-message, tool-response, api-misuse |
| Output structure mismatch | exception | error | python, benchmark, helion, kernels, testing |
| cannot be larger than | exception | error | speculative-decoding, max-model-len, draft-model, validation |
| --use-replayssm is only supported for Nemotron-H models | validation | error | mamba, replayssm, nemotron, model-support |
| Argument not found in the forward method of | exception | error | torch-compile, decorators, dynamic-shapes, vllm |
| To allow overriding this maximum, set the env var… | validation | error | max-model-len, context-length, config, startup |
| unexpected non-control output on coordinator path | exception | error | rust, coordinator, protocol, routing |
| must be a finite number, got | validation | error | rust, validation, sampling-params, vllm |
| RayPPCommunicator has been destroyed. | exception | critical | ray, pipeline-parallel, lifecycle, distributed |
| chat template looks like a file path but does not exist | exception | error | rust, chat-template, file-not-found, configuration, docker |