Why use DeepSpeed's implementation instead of torch's FSDP?
Open
- Dominant language
- Python
- Stars
- 3.4k
- Forks
- 312
- Avg merge
- 1h 2m
- Merged PRs (30d)
- 2
Description
It's really weird to see that a training framework in 2025 is still using DeepSpeed.
DeepSpeed's performance is terrible, and they nearly give up maintaining the repo themselves.
I can't see why you guys are still using DeepSpeed as the backend for ZeRO.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by locating the ZeRO backend and comparing its DeepSpeed integration with torch's FSDP. A complete outcome would document the performance and maintenance rationale and, if a change is agreed, define the migration scope and validation criteria.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100