{"record":{"id":"b6dea30a38cffc0c","repo":"hiyouga/LlamaFactory","slug":"please-specify-max-steps-in-streaming-mode","errorCode":null,"errorMessage":"Please specify `max_steps` in streaming mode.","messagePattern":"Please specify `max_steps` in streaming mode\\.","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"src/llamafactory/hparams/parser.py","lineNumber":475,"sourceCode":"        if finetuning_args.reward_model_type == \"lora\" and model_args.use_kt:\n            raise ValueError(\"KTransformers does not support lora reward model.\")\n\n        if finetuning_args.reward_model_type == \"lora\" and model_args.use_unsloth:\n            raise ValueError(\"Unsloth does not support lora reward model.\")\n\n        if training_args.report_to and any(\n            logger not in (\"wandb\", \"tensorboard\", \"trackio\", \"none\") for logger in training_args.report_to\n        ):\n            raise ValueError(\"PPO only accepts wandb, tensorboard, or trackio logger.\")\n\n    if not model_args.use_kt and training_args.parallel_mode == ParallelMode.NOT_DISTRIBUTED:\n        raise ValueError(\"Please launch distributed training with `llamafactory-cli` or `torchrun`.\")\n\n    if training_args.deepspeed and training_args.parallel_mode != ParallelMode.DISTRIBUTED:\n        raise ValueError(\"Please use `FORCE_TORCHRUN=1` to launch DeepSpeed training.\")\n\n    if training_args.max_steps == -1 and data_args.streaming:\n        raise ValueError(\"Please specify `max_steps` in streaming mode.\")\n\n    if training_args.do_train and data_args.dataset is None:\n        raise ValueError(\"Please specify dataset for training.\")\n\n    if (training_args.do_eval or training_args.do_predict or training_args.predict_with_generate) and (\n        data_args.eval_dataset is None and data_args.val_size < 1e-6\n    ):\n        raise ValueError(\"Please make sure eval_dataset be provided or val_size >1e-6\")\n\n    if training_args.predict_with_generate:\n        if is_deepspeed_zero3_enabled():\n            raise ValueError(\"`predict_with_generate` is incompatible with DeepSpeed ZeRO-3.\")\n\n        if finetuning_args.compute_accuracy:\n            raise ValueError(\"Cannot use `predict_with_generate` and `compute_accuracy` together.\")\n\n    if training_args.do_train and model_args.quantization_device_map == \"auto\":\n        raise ValueError(\"Cannot use device map for quantized models in training.\")","sourceCodeStart":457,"sourceCodeEnd":493,"githubUrl":"https://github.com/hiyouga/LlamaFactory/blob/f28afaf6355af515454dfb16c97d728307c93897/src/llamafactory/hparams/parser.py#L457-L493","documentation":"Raised in parser.py:475 when training_args.max_steps == -1 (the default: rely on epochs) and data_args.streaming is true. Streaming datasets have no known length, so epoch-based training loops cannot compute an end condition.","triggerScenarios":"A YAML/JSON config containing `streaming: true` without a `max_steps` entry, so max_steps keeps its default -1; equally `--streaming` on the CLI with no --max_steps.","commonSituations":"Switching a large pretrain/SFT dataset to streaming mode to avoid disk usage and forgetting to add max_steps; copying a non-streaming example config and flipping only the streaming flag.","solutions":["Add an explicit `max_steps: 1000` (or your budget) to the training config","Alternatively remove `streaming: true` if the dataset fits on disk, then use num_train_epochs"],"exampleFix":"# before (YAML)\nstreaming: true\n# max_steps not set (defaults to -1)\n\n# after\nstreaming: true\nmax_steps: 1000","handlingStrategy":"validation","validationCode":"if config.get(\"streaming\") and config.get(\"max_steps\", -1) == -1:\n    raise SystemExit(\"streaming: true requires an explicit max_steps\")","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Whenever flipping streaming: true, immediately add max_steps in the same edit","Pre-flight check YAML pairs (streaming, max_steps) in CI before GPU submission"],"tags":["streaming","dataset","max-steps","config-validation"],"backgroundTag":null,"analyzedSha":"f28afaf6355af515454dfb16c97d728307c93897","analyzedAt":"2026-08-14T21:57:28.298Z","schemaVersion":2},"datasetVersion":"2026-08-15T22:17:37.221Z"}