hiyouga/LlamaFactory · error · ValueError
Please use `FORCE_TORCHRUN=1` to launch DeepSpeed training.
Error message
Please use `FORCE_TORCHRUN=1` to launch DeepSpeed training.
What it means
Raised in parser.py:472 when training_args.deepspeed is set but parallel_mode is not DISTRIBUTED. DeepSpeed integration in LlamaFactory depends on torchrun's distributed environment (RANK, LOCAL_RANK, WORLD_SIZE); without it, DeepSpeed hooks cannot initialize.
Source
Thrown at src/llamafactory/hparams/parser.py:472
if model_args.shift_attn:
raise ValueError("PPO training is incompatible with S^2-Attn.")
if finetuning_args.reward_model_type == "lora" and model_args.use_kt:
raise ValueError("KTransformers does not support lora reward model.")
if finetuning_args.reward_model_type == "lora" and model_args.use_unsloth:
raise ValueError("Unsloth does not support lora reward model.")
if training_args.report_to and any(
logger not in ("wandb", "tensorboard", "trackio", "none") for logger in training_args.report_to
):
raise ValueError("PPO only accepts wandb, tensorboard, or trackio logger.")
if not model_args.use_kt and training_args.parallel_mode == ParallelMode.NOT_DISTRIBUTED:
raise ValueError("Please launch distributed training with `llamafactory-cli` or `torchrun`.")
if training_args.deepspeed and training_args.parallel_mode != ParallelMode.DISTRIBUTED:
raise ValueError("Please use `FORCE_TORCHRUN=1` to launch DeepSpeed training.")
if training_args.max_steps == -1 and data_args.streaming:
raise ValueError("Please specify `max_steps` in streaming mode.")
if training_args.do_train and data_args.dataset is None:
raise ValueError("Please specify dataset for training.")
if (training_args.do_eval or training_args.do_predict or training_args.predict_with_generate) and (
data_args.eval_dataset is None and data_args.val_size < 1e-6
):
raise ValueError("Please make sure eval_dataset be provided or val_size >1e-6")
if training_args.predict_with_generate:
if is_deepspeed_zero3_enabled():
raise ValueError("`predict_with_generate` is incompatible with DeepSpeed ZeRO-3.")
if finetuning_args.compute_accuracy:
raise ValueError("Cannot use `predict_with_generate` and `compute_accuracy` together.")View on GitHub (pinned to f28afaf635)
Solutions
- Relaunch with `FORCE_TORCHRUN=1 llamafactory-cli train config.yaml`
- Ensure you go through `llamafactory-cli` / `lmf` so torchrun wraps the process when deepspeed is configured
- Verify no wrapper strips distributed env vars (e.g. a custom docker entrypoint or python -m path)
Example fix
# before llamafactory-cli train my_deepspeed.yaml # missing FORCE_TORCHRUN # after FORCE_TORCHRUN=1 llamafactory-cli train my_deepspeed.yaml
Defensive patterns
Strategy: validation
Validate before calling
if config.get("deepspeed") and not os.environ.get("FORCE_TORCHRUN"):
raise SystemExit("DeepSpeed configs require FORCE_TORCHRUN=1 llamafactory-cli train") Prevention
- Standardize on FORCE_TORCHRUN=1 for every DeepSpeed run
- Never call tuner functions directly when deepspeed is set in args
- Wrap launches in a make target or shell script that sets the env var
When it happens
Trigger: Passing `--deepspeed ds_z2_config.json` (or the YAML equivalent) while launching the process so that ParallelMode stays NOT_DISTRIBUTED — typically running via plain python, or forgetting FORCE_TORCHRUN=1 which suppresses the torchrun wrapper.
Common situations: Adding a deepspeed section to the YAML but launching with a plain python entry point; CI scripts that call the tuner API directly; single-GPU users who assume DeepSpeed ZeRO works without a distributed launcher.
Related errors
- Please launch distributed training with `llamafactory-cli` o
- Megatron Bridge is incompatible with DeepSpeed.
- Layer-wise BAdam only supports DeepSpeed ZeRO-3 training.
- DeepSpeed only supports bf16 mixed precision for now, fp16 i
- DeepSpeed config_file is required in dist_config
AI-assisted analysis of hiyouga/LlamaFactory@f28afaf635 (2026-08-14).
Data as JSON: /api/errors/fc7090405c16c3b4.
Report an issue: GitHub.