deepspeedai / deepspeedai/DeepSpeed
Pipeline comparison with Megatron
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
I realized that Megatron now also provides pipeline parallelism implementation, do you have any comparison with their implementation? For example, any benchmark, or design difference? Also, if it is possible, can you also comment on the implementation on FairScale? How would you evaluate the extendibility of these three repos on further researching on the direction on degree of parallelism? Thanks!
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by comparing DeepSpeed's pipeline-parallelism implementation with the corresponding Megatron and FairScale implementations. Investigate benchmark results, design differences, and extensibility for different parallelism degrees; done would be a documented comparison with conclusions on future research directions.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100