bigscience-workshop / bigscience-workshop/Megatron-DeepSpeed
Can we also train BLOOM model using tensor using tensor-Parallelism and efficient fused CUDA kernels
- Dominant language
- Python
- Stars
- 1.4k
- Forks
- 226
- PR merge metrics
- No merged PRs in 30d
Description
Hi,
Thanks for the great work. I was able to do Inference on BLOOM 7.1 model on 24 GB GPU memory. Can we train the BLOOM models using tensor-Parallelism and efficient fused CUDA kernels? As I don't have access to high memory.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing the repository's existing BLOOM inference and training support, then determine how tensor parallelism and efficient fused CUDA kernels would apply to training on a 24 GB GPU. Done would mean a documented, working way to train a BLOOM model under the stated memory constraint, with appropriate validation of the training result.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100