{"record":{"id":"ac5fb997441c906e","repo":"hiyouga/LlamaFactory","slug":"distributed-training-does-not-support-layer-wise-a","errorCode":null,"errorMessage":"Distributed training does not support layer-wise APOLLO.","messagePattern":"Distributed training does not support layer-wise APOLLO\\.","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"src/llamafactory/hparams/parser.py","lineNumber":510,"sourceCode":"    if training_args.do_train and model_args.quantization_device_map == \"auto\":\n        raise ValueError(\"Cannot use device map for quantized models in training.\")\n\n    if finetuning_args.pissa_init and is_deepspeed_zero3_enabled():\n        raise ValueError(\"Please use scripts/pissa_init.py to initialize PiSSA in DeepSpeed ZeRO-3.\")\n\n    if finetuning_args.pure_bf16:\n        if not (is_torch_bf16_gpu_available() or (is_torch_npu_available() and torch.npu.is_bf16_supported())):\n            raise ValueError(\"This device does not support `pure_bf16`.\")\n\n        if is_deepspeed_zero3_enabled():\n            raise ValueError(\"`pure_bf16` is incompatible with DeepSpeed ZeRO-3.\")\n\n    if training_args.parallel_mode == ParallelMode.DISTRIBUTED:\n        if finetuning_args.use_galore and finetuning_args.galore_layerwise:\n            raise ValueError(\"Distributed training does not support layer-wise GaLore.\")\n\n        if finetuning_args.use_apollo and finetuning_args.apollo_layerwise:\n            raise ValueError(\"Distributed training does not support layer-wise APOLLO.\")\n\n        if finetuning_args.use_badam:\n            if finetuning_args.badam_mode == \"ratio\":\n                raise ValueError(\"Radio-based BAdam does not yet support distributed training, use layer-wise BAdam.\")\n            elif not is_deepspeed_zero3_enabled():\n                raise ValueError(\"Layer-wise BAdam only supports DeepSpeed ZeRO-3 training.\")\n\n    if training_args.deepspeed is not None and (finetuning_args.use_galore or finetuning_args.use_apollo):\n        raise ValueError(\"GaLore and APOLLO are incompatible with DeepSpeed yet.\")\n\n    if (\n        not finetuning_args.use_mca\n        and not finetuning_args.use_megatron_bridge\n        and training_args.fp8\n        and model_args.quantization_bit is not None\n    ):\n        raise ValueError(\"FP8 training is not compatible with quantization. Please disable one of them.\")\n","sourceCodeStart":492,"sourceCodeEnd":528,"githubUrl":"https://github.com/hiyouga/LlamaFactory/blob/f28afaf6355af515454dfb16c97d728307c93897/src/llamafactory/hparams/parser.py#L492-L528","documentation":"Raised in parser.py:510 inside the distributed block when use_apollo and apollo_layerwise are both true. APOLLO is a GaLore-like low-rank projector optimizer; its layer-wise variant has the same cross-rank synchronization problem as layer-wise GaLore, so it is rejected in distributed runs.","triggerScenarios":"torchrun/llamafactory-cli multi-process launch with `use_apollo: true` and `apollo_layerwise: true` in the config.","commonSituations":"Scaling a single-GPU APOLLO memory-saving recipe to multiple GPUs; mixing APOLLO layerwise with DeepSpeed or DDP setups.","solutions":["Remove `apollo_layerwise: true` for distributed training","If layer-wise is required, run on a single process/GPU without torchrun"],"exampleFix":"# before (YAML)\nuse_apollo: true\napollo_layerwise: true  # multi-GPU launch\n\n# after\nuse_apollo: true\n# apollo_layerwise removed","handlingStrategy":"validation","validationCode":"if world_size > 1 and config.get(\"use_apollo\") and config.get(\"apollo_layerwise\"):\n    raise SystemExit(\"apollo_layerwise is single-process only\")","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Apply the same distributed lint to APOLLO flags as to GaLore","Avoid porting single-GPU layerwise recipes to torchrun unchanged"],"tags":["apollo","optimizer","distributed","layerwise","incompatible-flags"],"backgroundTag":null,"analyzedSha":"f28afaf6355af515454dfb16c97d728307c93897","analyzedAt":"2026-08-14T21:57:28.298Z","schemaVersion":2},"datasetVersion":"2026-08-15T17:31:12.345Z"}