Lightning-AI / Lightning-AI/pytorch-lightning
CUDA Streams for parallel sub-module training
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 31.4k
- Forks
- 3.8k
- Avg merge
- 6d 7h
- Merged PRs (30d)
- 6
Description
### Description & Motivation
I have a simple lightning model that is composed of a DAG of nodes/sub-models represented by feed forward networks. At training time, each of these nodes can be trained in parallel. However, I don't see any documentation on how to use [torch.cuda.Stream](https://pytorch.org/docs/stable/generated/torch.cuda.Stream.html) with lightning.
### Pitch
For models that have small sub-models, this allows parallelization within a single gpu.
### Alternatives
The alternative is to loop over every sub module.
### Additional context
_No response_
cc @borda
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
No files, tests, or entry points are identified in the issue. Start by locating Lightning's documentation and training-loop integration points, then determine how torch.cuda.Stream could be used for parallel sub-module training; done means documenting a supported, usable approach and its constraints.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- documentation, machine-learning
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 32/100