{"record":{"id":"1233af61e079224b","repo":"hiyouga/LlamaFactory","slug":"radio-based-badam-does-not-yet-support-distributed","errorCode":null,"errorMessage":"Radio-based BAdam does not yet support distributed training, use layer-wise BAdam.","messagePattern":"Radio-based BAdam does not yet support distributed training, use layer-wise BAdam\\.","errorType":"validation","errorClass":"ValueError","httpStatus":null,"severity":"error","filePath":"src/llamafactory/hparams/parser.py","lineNumber":514,"sourceCode":"        raise ValueError(\"Please use scripts/pissa_init.py to initialize PiSSA in DeepSpeed ZeRO-3.\")\n\n    if finetuning_args.pure_bf16:\n        if not (is_torch_bf16_gpu_available() or (is_torch_npu_available() and torch.npu.is_bf16_supported())):\n            raise ValueError(\"This device does not support `pure_bf16`.\")\n\n        if is_deepspeed_zero3_enabled():\n            raise ValueError(\"`pure_bf16` is incompatible with DeepSpeed ZeRO-3.\")\n\n    if training_args.parallel_mode == ParallelMode.DISTRIBUTED:\n        if finetuning_args.use_galore and finetuning_args.galore_layerwise:\n            raise ValueError(\"Distributed training does not support layer-wise GaLore.\")\n\n        if finetuning_args.use_apollo and finetuning_args.apollo_layerwise:\n            raise ValueError(\"Distributed training does not support layer-wise APOLLO.\")\n\n        if finetuning_args.use_badam:\n            if finetuning_args.badam_mode == \"ratio\":\n                raise ValueError(\"Radio-based BAdam does not yet support distributed training, use layer-wise BAdam.\")\n            elif not is_deepspeed_zero3_enabled():\n                raise ValueError(\"Layer-wise BAdam only supports DeepSpeed ZeRO-3 training.\")\n\n    if training_args.deepspeed is not None and (finetuning_args.use_galore or finetuning_args.use_apollo):\n        raise ValueError(\"GaLore and APOLLO are incompatible with DeepSpeed yet.\")\n\n    if (\n        not finetuning_args.use_mca\n        and not finetuning_args.use_megatron_bridge\n        and training_args.fp8\n        and model_args.quantization_bit is not None\n    ):\n        raise ValueError(\"FP8 training is not compatible with quantization. Please disable one of them.\")\n\n    if model_args.infer_backend != EngineName.HF:\n        raise ValueError(\"vLLM/SGLang backend is only available for API, CLI and Web.\")\n\n    if model_args.use_unsloth and is_deepspeed_zero3_enabled():","sourceCodeStart":496,"sourceCodeEnd":532,"githubUrl":"https://github.com/hiyouga/LlamaFactory/blob/f28afaf6355af515454dfb16c97d728307c93897/src/llamafactory/hparams/parser.py#L496-L532","documentation":"Raised in parser.py:514 inside the distributed block when use_badam is true and badam_mode == \"ratio\" (note: the message's \"Radio\" is a typo for ratio-based). Ratio-based BAdam updates a random fraction of blocks per step; the random choice diverges across ranks in distributed training, so only layer-wise BAdam is supported there.","triggerScenarios":"Multi-GPU launch with `use_badam: true` and `badam_mode: ratio` in the finetuning args.","commonSituations":"Using the BAdam blockwise optimizer to cut optimizer memory on a multi-GPU finetune while keeping the default ratio mode; copying BAdam docs/examples that default to ratio.","solutions":["Set `badam_mode: layer-wise` for distributed runs","Or drop use_badam entirely and use a supported distributed optimizer (adamw_torch, GaLore non-layerwise, or DeepSpeed)","Run single-GPU if ratio-based BAdam is specifically needed"],"exampleFix":"# before (YAML)\nuse_badam: true\nbadam_mode: ratio  # multi-GPU launch\n\n# after\nuse_badam: true\nbadam_mode: layer-wise","handlingStrategy":"validation","validationCode":"if world_size > 1 and config.get(\"use_badam\") and config.get(\"badam_mode\", \"ratio\") == \"ratio\":\n    raise SystemExit(\"Set badam_mode: layer-wise for distributed runs\")","typeGuard":null,"tryCatchPattern":null,"preventionTips":["Remember ratio is the default badam_mode — set it explicitly in distributed configs","Note the upstream typo (\"Radio\") so log-greps don't miss it"],"tags":["badam","optimizer","distributed","layerwise","incompatible-flags"],"backgroundTag":null,"analyzedSha":"f28afaf6355af515454dfb16c97d728307c93897","analyzedAt":"2026-08-14T21:57:28.298Z","schemaVersion":2},"datasetVersion":"2026-08-15T22:17:37.221Z"}