hiyouga/LlamaFactory
Documented errors, page 2 of 4. Back to hiyouga/LlamaFactory
| Code / Message | Type | Severity | Tags |
|---|---|---|---|
| `reward_model_type` cannot be oft for Freeze/Full PPO… | validation | error | config, ppo, oft, reward-model |
| FSDPTurbo EP mesh is not initialized. | exception | critical | distributed, fsdp, expert-parallelism, device-mesh |
| Neat packing is not supported for gemma4, gpt_oss models… | exception | error | neat-packing, flash-attention, gemma, gpt-oss, training |
| Total Megatron Bridge parallel size | validation | error | megatron, distributed, parallelism, config, validation |
| Cannot merge adapters to a quantized model. | exception | error | export, quantization, lora, merge |
| DPO training dataset is empty | validation | error | data, dataset, dpo, validation |
| Per-layer GaLore does not support gradient accumulation. | exception | error | galore, layerwise, gradient-accumulation, config |
| Bitsandbytes only accepts 4-bit or 8-bit quantization. | validation | error | quantization, bitsandbytes, configuration |
| Cannot find sufficient samples, consider increasing dataset… | exception | error | dataset, pretrain, empty-dataset, preprocessing |
| KTransformers is incompatible with DeepSpeed ZeRO-3. | validation | error | ktransformers, deepspeed, zero3, training-config |
| PPO only accepts wandb, tensorboard, or trackio logger. | validation | error | ppo, rlhf, logging, wandb, tensorboard, config-validation |
| This model does not support image input. Please check… | exception | error | multimodal, image, template, config |
| Current model is not supported by mixture-of-depth. | exception | error | mixture-of-depths, architecture, model-support |
| FSDPTurbo EFSDP mesh is not initialized. | exception | critical | distributed, fsdp, expert-parallelism, device-mesh, config |
| Invalid URL | http | error | security, ssrf, url-validation, http-400 |
| Output directory already exists and is not empty. Please… | validation | error | output-dir, checkpoint, resume, training-config |
| RM training requires pair-format samples containing… | validation | error | data, dataset, reward-model, preference-data, validation |
| compute_dtype= is not a torch dtype name. | validation | error | quantization, bitsandbytes, dtype, configuration |
| Currently merge and export model function is only supported… | validation | error | peft, export, lora, configuration |
| `quantization_bit` cannot be combined with KT weight caches. | exception | error | fp8, kt-kernel, quantization, config-conflict |
| `recompute_num_layers` must be >= 1 when set. | validation | error | megatron, activation-checkpointing, config-validation |
| All 'dcp_path', 'hf_path', and 'config_path' are required. | exception | error | version, transformers, internlm3, dependency |
| LLaMA-Factory `kt_config` must be a flat mapping. | validation | error | ktransformers, yaml, type-error, config-validation |
| Module is not found, please choose from | validation | error | freeze-tuning, module-names, architecture, config |
| `predict_with_generate` cannot be set as True except SFT. | validation | error | config, validation, evaluation, stage, sft |
| Unknown logging level | validation | error | logging, environment, config, startup |
| bf16 and fp16 cannot be both True. | exception | error | dpo, loss-function, reference-model, config |
| Device not supported | exception | error | version, transformers, kto, dependency |
| `num_layers` should be divisible by `num_layer_trainable` . | exception | error | lora, num-layer-trainable, config-validation |
| The dataset is not applicable in the current training stage. | exception | error | dataset, training-stage, ranking, config |
| tools is not valid JSON | validation | error | tools, json, function-calling, rendering |
| `train_on_prompt` or `mask_history` cannot be set as True… | validation | error | config, validation, data-mask, stage, sft |
| Invalid input type . | http | error | api, multimodal, validation, http-400, chat |
| Megatron Bridge cannot be used together with MCA or… | validation | error | megatron, distributed, config, backend-conflict, validation |
| `image_max_pixels` cannot be smaller than… | validation | error | hparams, multimodal, vision, config-validation |
| Invalid role | http | error | ppo, rlhf, eval-dataset, config, not-implemented |
| megatron-bridge is required when USE_MEGATRON_BRIDGE=1… | exception | error | megatron, dependency, environment, import-error |
| All datasets must be streaming or non-streaming. | validation | error | v1, datasets, streaming, configuration |
| DeepSpeed ZeRO-3 or FSDP is incompatible with PTQ-quantized… | exception | error | quantization, gptq, deepspeed, fsdp, zero3, training-config |
| More tags than provided media files. | validation | error | data, multimodal, dataset-conversion, validation |
| Radio-based BAdam does not yet support distributed… | validation | error | badam, optimizer, distributed, layerwise, incompatible-flags |
| Expected a string, got | exception | error | formatter, dataset, data-quality, template |
| `mask_history` is incompatible with `train_on_prompt`. | validation | error | config, multi-turn, loss-masking, data |
| Drop last must be True. | validation | error | batching, config, dataloader, training |
| The number of images does not match the number of | exception | error | multimodal, image, data-format, placeholders |
| Allowed file types: . | exception | error | dataset, file-format, config |
| FSDPTurbo parallel state is already initialized with | exception | error | fsdpturbo, distributed, singleton, lifecycle |
| LLaMA-Factory YAML and Accelerate config cannot define… | validation | error | ktransformers, accelerate, config-conflict |
| Sequence parallel is not supported for qwen3.5 model due to… | exception | error | v1, context-parallelism, qwen3, model-support |
| Unknown sample backend | exception | error | v1, sampling, backend, configuration |
| Model was not supported. | exception | error | lora, num-layer-trainable, config, custom-model |
| `max_samples` is incompatible with `streaming`. | validation | error | config, streaming, data, sampling |
| tool_call must be a JSON object with 'name' and 'arguments'… | validation | error | tool-calls, json, schema, multimodal, rendering |
| Unexpected role | exception | error | data, roles, sharegpt, template |
| FP8 training is not compatible with quantization. Please… | validation | error | fp8, quantization, precision, incompatible-flags, config-validation |
| Other sequence parallel modes are to be implemented. | exception | error | not-implemented, sequence-parallel, attention, distributed |
| Please provide `model_name_or_path`. | validation | critical | hparams, required-field, config-validation, quickstart |
| tool_call value is not valid JSON | validation | error | tool-calls, json, multimodal, data-format, rendering |
| Iterable dataset is not supported yet. | exception | error | dataset, iterable-dataset, streaming, not-implemented, v1 |
| Megatron Bridge arguments are missing. Please set… | validation | error | megatron, config, environment, validation |
| MOSS-VL video frame tokens do not match the processed video… | exception | error | multimodal, moss-vl, video, frames, truncation |
| world_size ( ) must be divisible by mp_replicate_size ( ). | exception | error | distributed-training, v1, configuration, device-mesh |
| GaLore and APOLLO are incompatible with DeepSpeed yet. | validation | error | galore, apollo, deepspeed, optimizer, incompatible-flags |
| `kt_config` requires `use_kt: true`. | validation | error | ktransformers, config, llamafactory |
| Pipeline parallel size should be smaller than the number of… | exception | error | dependency, megatron, mcore-adapter, installation |
| Quantized models can only be used for the LoRA or OFT… | validation | error | qlora, quantization, finetuning-type, full-tuning |
| File not found. | exception | error | dataset, file-not-found, path, config |
| Unknown task: . | exception | error | config, stage, dispatch, typo |
| Please use scripts/pissa_init.py to initialize PiSSA for a… | validation | error | pissa, quantization, config, llamafactory |
| Unable to process key | exception | error | model, llava, multimodal, checkpoint-format, config |
| Video processor was not found, please check and update your… | exception | error | multimodal, video-processor, model-files, transformers |
| When `adapter_name_or_path` is provided for training, only… | validation | error | configuration, lora, peft, training |
| SGLang only supports n=1. | exception | error | sglang, sampling, not-implemented, inference |
| chunk_size must be one of | validation | error | configuration, validation, kernel-plugin |
| `dynamic_batching` requires `max_steps` because it is… | exception | error | v1, configuration, batching, training-args |
| Both 'hf_path' and 'dcp_path' are required. | exception | error | version, transformers, lfm2-vl, dependency |
| Unknown load type: . | exception | error | dataset, dataset-info, config |
| Access to private or reserved IP addresses is not allowed. | http | error | security, ssrf, network, http-403 |
| Cannot stream function calls. | http | error | api, streaming, tools, http-400, function-calling |
| `interleave_probs` is only valid for interleaved mixing. | exception | error | config, data-args, mixing |
| Streaming dataset does not support index access. | validation | error | v1, datasets, streaming, api-misuse |
| GLM-4 does not support parallel functions. | exception | error | tools, glm4, data |
| No valid RM pairs found in this micro-batch. This is… | validation | error | reward-model, truncation, cutoff-len, data |
| The installed Transformers-KT does not provide… | exception | critical | ktransformers, dependencies, version-mismatch, llamafactory |
| training render expects the last message to be the… | validation | error | training-data, chat-template, supervision, rendering |
| Global batch size must be divisible by DP size and micro… | validation | error | batching, distributed, config, training |
| mcore_adapter is required when USE_MCA=1. Please install… | exception | critical | megatron, mca, environment, import-error, dependency |
| Context parallelism currently requires `dist_config.name… | exception | error | v1, context-parallelism, deepspeed, fsdp2, configuration |
| Per-layer APOLLO does not support gradient accumulation. | exception | error | apollo, layerwise, gradient-accumulation, config |
| Empty formatter should not contain any placeholder. | exception | error | formatter, template, config, custom-template |
| megatron-bridge is not installed. Please install it with… | exception | error | megatron, optional-dependency, installation, distributed |
| Unknown Liger op(s) for model_type= . Valid | validation | error | kernels, liger, config, validation |
| `video_max_pixels` cannot be smaller than… | validation | error | hparams, multimodal, video, config-validation |
| Please launch distributed training with `llamafactory-cli`… | validation | error | launcher, distributed, torchrun, config-validation |
| RM and PPO stages do not support `load_best_model_at_end`. | validation | error | config, validation, reward-model, ppo, checkpoints |
| Adapter is only valid for the LoRA method. | validation | error | lora, adapter, config, llamafactory |
| Invalid image found in video frames. | exception | error | multimodal, video, frames, path-resolution |
| Audio feature extractor was not found, please check and… | exception | error | multimodal, audio, feature-extractor, model-files |
| `compute_accuracy` is not supported in KTransformers SFT… | exception | error | ktransformers, sft, evaluation, accuracy |
| Unknown pref_loss | validation | error | dpo, config, loss-function, validation |