deepspeedai / deepspeedai/DeepSpeed
Two PipelineModule within a Model will block the deepspeed initialization
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 43.1k
- Forks
- 5k
- Avg merge
- 4d 15h
- Merged PRs (30d)
- 112
Description
Describe the bug
I am training a text-protein model, and I need to split both text (LLaMA) and protein (ESM2) models into multiple GPUs.
I add PipelineModule in the text model and protein model separately; then, the code is blocked at the deepspeed initialization.
Does deepspeed not support this manner? Any suggestions, please?
I already make sure that using PipelineModule either in text or protein models individually is feasible.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the PipelineModule initialization path in DeepSpeed and reproduce the reported setup using separate pipeline modules for the text and protein models. The issue names no files or tests, so first determine where initialization blocks and whether multiple PipelineModule instances are supported. Done means identifying the cause and documenting or validating a supported resolution.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- distributed-systems, machine-learning
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100