deepspeedai / deepspeedai/DeepSpeed

[REQUEST] Allow pre-sharding of models using a single process

Open
#3,562 2 comments 2 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

Is your feature request related to a problem? Please describe.
When using multiple GPUs for tensor parallel inference, DeepSpeed loads the model into memory N times, where N is the number of GPUs. This can quickly exhaust system memory when the model is large. For example, I cannot load opt-66b onto 8 GPUs with tp_size = 8 on a DGXA100 system with 1TiB of system memory.

A solution is to pre-shard the model so that each process doesn't need to load the entire model into memory, but in order to do this, one STILL has to load the model into memory N times.

Describe the solution you'd like
Allow pre-sharding of a model across tp_size devices using a SINGLE process, so as to avoid loading the model redundantly into memory N times.

Describe alternatives you've considered
A clear and concise description of any alternative solutions or features you've considered.

Additional context
If you try to do this today, with this kind of setup:

deepspeed --num_gpus 1 SCRIPT.py --tp_size 8 --checkpoint_path <PATH>

tp_config = deepspeed.inference.config.DeepSpeedTPConfig(
    enabled = True,
    tp_size = args.tp_size,
)
deepspeed.init_inference(
    tensor_parallel = tp_config,
    save_mp_checkpoint_path = args.checkpoint_path,
    ....
)

You'll get this error:

RuntimeError: the new group's world size should be less or equal to the world size set by init_process_group

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start with DeepSpeedTPConfig and deepspeed.init_inference, then trace the process-group path that produces the world-size RuntimeError when --num_gpus 1 is used with tp_size 8. Done means a single process can pre-shard the model across tp_size devices without redundant full-model loads, while preserving the save_mp_checkpoint_path workflow.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.