deepspeedai / deepspeedai/DeepSpeed

[REQUEST] Allow fallback to torch.distributed.all_reduce for multi-node inference.

Open
#7,625 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

Is your feature request related to a problem? Please describe.
At present, DeepSpeed’s inference communication backend defaults to using the SHM-based operation torch.ops.deepspeed.inference_all_reduce_ for performing all_reduce operations when the shared memory (SHM) op is available. This default behavior, while effective for intra-node communication (e.g., dual-socket CPU machines), becomes problematic in multi-node inference scenarios.

The SHM implementation is currently tailored with x86-specific intrinsics, and as such, it’s incompatible with alternative CPU architectures like Arm. Moreover, based on my exploration, there’s no fallback mechanism to explicitly select a different communication backend (like torch.distributed.all_reduce) even when SHM is present. This limits flexibility for multi-node or non-x86 deployments where SHM may not work or be optimal.

Describe the solution you'd like
We propose modifying the inference_all_reduce method in deepspeed/comm/torch.py to introduce a runtime-aware mechanism that conditionally chooses between the SHM and PyTorch’s default distributed backends based on node locality.

A sample implementation might look like this:

def inference_all_reduce(self, tensor, op, group=None):
    all_local_ranks = True
    if group != None:
        world_size = torch.distributed.get_world_size(group=group)
        if (world_size > int(os.getenv("LOCAL_SIZE"))):
            all_local_ranks = False
             
    if not hasattr(torch.ops, 'deepspeed') or not hasattr(torch.ops.deepspeed, 'inference_all_reduce_') or not all_local_ranks:
        op = self._reduce_op(op)
        return torch.distributed.all_reduce(tensor=tensor, op=op, group=group, async_op=False)
    else:
        return torch.ops.deepspeed.inference_all_reduce_(tensor)

Describe alternatives you've considered - but not tested

  • Add a runtime flag or env var (e.g., DS_SHM_COMM_ALL_REDUCE_OFF) to disable SHM-based collectives and fallback to torch.distributed ops.
  • Introduce a parameter to deepspeed.init_inference() to control SHM usage, useful for hybrid systems (e.g., dual-socket or mixed-arch setups).
  • Skip SHM build in setup.py for unsupported architectures like arm to avoid x86-intrinsic issues.

Additional Context
We are working on enabling multi-node inference for large language models on CPU-only clusters, including support for Arm processors. The current behavior of using SHM collectives creates runtime issues. Adding flexibility in collective selection or building would improve support.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start in deepspeed/comm/torch.py at inference_all_reduce and trace how the SHM operation and torch.distributed backend are selected. Compare the proposed node-locality behavior with the runtime flag and init_inference alternatives, then validate the chosen approach for multi-node CPU inference and Arm support.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
42/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.