hiyouga/LlamaFactory · error · ValueError
Please specify dataset for training.
Error message
Please specify dataset for training.
What it means
Raised in parser.py:478 when training_args.do_train is true but data_args.dataset is None. LlamaFactory trains purely from dataset names registered in data/dataset_info.json, so an empty dataset list leaves the trainer with nothing to consume.
Source
Thrown at src/llamafactory/hparams/parser.py:478
if finetuning_args.reward_model_type == "lora" and model_args.use_unsloth:
raise ValueError("Unsloth does not support lora reward model.")
if training_args.report_to and any(
logger not in ("wandb", "tensorboard", "trackio", "none") for logger in training_args.report_to
):
raise ValueError("PPO only accepts wandb, tensorboard, or trackio logger.")
if not model_args.use_kt and training_args.parallel_mode == ParallelMode.NOT_DISTRIBUTED:
raise ValueError("Please launch distributed training with `llamafactory-cli` or `torchrun`.")
if training_args.deepspeed and training_args.parallel_mode != ParallelMode.DISTRIBUTED:
raise ValueError("Please use `FORCE_TORCHRUN=1` to launch DeepSpeed training.")
if training_args.max_steps == -1 and data_args.streaming:
raise ValueError("Please specify `max_steps` in streaming mode.")
if training_args.do_train and data_args.dataset is None:
raise ValueError("Please specify dataset for training.")
if (training_args.do_eval or training_args.do_predict or training_args.predict_with_generate) and (
data_args.eval_dataset is None and data_args.val_size < 1e-6
):
raise ValueError("Please make sure eval_dataset be provided or val_size >1e-6")
if training_args.predict_with_generate:
if is_deepspeed_zero3_enabled():
raise ValueError("`predict_with_generate` is incompatible with DeepSpeed ZeRO-3.")
if finetuning_args.compute_accuracy:
raise ValueError("Cannot use `predict_with_generate` and `compute_accuracy` together.")
if training_args.do_train and model_args.quantization_device_map == "auto":
raise ValueError("Cannot use device map for quantized models in training.")
if finetuning_args.pissa_init and is_deepspeed_zero3_enabled():
raise ValueError("Please use scripts/pissa_init.py to initialize PiSSA in DeepSpeed ZeRO-3.")View on GitHub (pinned to f28afaf635)
Solutions
- Add `dataset: <name>` matching an entry in data/dataset_info.json to the config
- Check spelling of the key — it must be exactly `dataset` (list or string)
- If the dataset is not registered yet, add its definition to dataset_info.json first, then reference it
Example fix
# before (YAML) ### dataset # do_train: true, no dataset key # after do_train: true dataset: alpaca_gpt4_zh
Defensive patterns
Strategy: validation
Validate before calling
if config.get("do_train") and not config.get("dataset"):
raise SystemExit("do_train: true requires dataset: <name from dataset_info.json>") Prevention
- Assert required keys (dataset, do_train pairs) right after loading the YAML
- Register datasets in dataset_info.json before referencing them
- Fail fast on empty template variables (e.g. ${DATASET}) in generated configs
When it happens
Trigger: `llamafactory-cli train` with a config missing the `dataset:` key (or dataset: null) while do_train is set; also calling get_train_args on a dict without "dataset" then running run_sft.
Common situations: Typos like `datasets:` or `dataset_name:` instead of `dataset:`; a WebUI-generated YAML where the dataset field was left blank; template variables (e.g. ${DATASET}) expanding to empty in CI.
Related errors
- Unsupported model type: {getattr(config, 'model_type')}.
- Cannot specify `val_size` if `eval_dataset` is not None.
- YAML config must be a dictionary mapping tokens to descripti
- LLaMA-Factory `kt_config` must be a flat mapping.
- Please specify `max_steps` in streaming mode.
AI-assisted analysis of hiyouga/LlamaFactory@f28afaf635 (2026-08-14).
Data as JSON: /api/errors/2680050f76fdf2cd.
Report an issue: GitHub.