deepspeedai / deepspeedai/DeepSpeed
LLama factory, quantized model and deepspeed compatibility
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
hello,
May I ask at this moment if deepspeed is compatible with 4-bit quantized model at ZeRo-3(multi-GPUs)?
I downloaded a Deepseek-32B-4bit model and try to use LLama factory to launch Lora finetuning, and was prompted with the following error:
main/src/llamafactory/model/model_utils/quantization.py", line 114, in configure_quantization
[rank1]: raise ValueError("DeepSpeed ZeRO-3 or FSDP is incompatible with PTQ-quantized models.")
[rank1]: ValueError: DeepSpeed ZeRO-3 or FSDP is incompatible with PTQ-quantized models.
If LLama factory is incompatible, can I use python scripts to launch finetuning with deepspeed Zero-3?
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with main/src/llamafactory/model/model_utils/quantization.py at line 114 and trace the reported ZeRO-3 or FSDP compatibility check. Compare the Deepseek-32B 4-bit LoRA fine-tuning setup with the requested multi-GPU workflow; done would require a documented compatibility answer or a clearly supported launch path.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100