bigscience-workshop / bigscience-workshop/Megatron-DeepSpeed

Can we also train BLOOM model using tensor using tensor-Parallelism and efficient fused CUDA kernels

Open
#334 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.4k
Forks
226
PR merge metrics
No merged PRs in 30d

Description

Hi,

Thanks for the great work. I was able to do Inference on BLOOM 7.1 model on 24 GB GPU memory. Can we train the BLOOM models using tensor-Parallelism and efficient fused CUDA kernels? As I don't have access to high memory.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the repository's existing BLOOM inference and training support, then determine how tensor parallelism and efficient fused CUDA kernels would apply to training on a 24 GB GPU. Done would mean a documented, working way to train a BLOOM model under the stated memory constraint, with appropriate validation of the training result.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.