deepspeedai / deepspeedai/DeepSpeed

Consolidate norm computation logic

Open
#1,839 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement
Dominant language
Python
Stars
43.1k
Forks
5k
Avg merge
4d 15h
Merged PRs (30d)
112

Description

I wonder if it's possible to consolidate some of the implementations of global norm? I count at least seven implementations in DeepSpeed: five in this file and one each in stage1/2 and stage 3. I'm

Summarizing implementations in utils.py:

clip_grad_norm_: has additional MOE code, but does not handle model parallelism correctly (uses get_model_parallel_rank() instead of bwc_tensor_model_parallel_rank()
get_grad_norm: does not include the MOE considerations, handles model parallelism
get_weight_norm: handles model parallelism, but if the list of input grads are already extracted from parameters then we will not have the model parallel attributes.
get_global_norm_of_tensors: does not consider model parallelism. I see it's used in conjuction with filtering tensors based on MP attributes. I like this refactor :)
get_global_norm: used to get the total of a list of existing partial norms. No parallelism considered. FP32, fused FP16, and unfused FP16 all use this combined with split_params_grads... and get_weight_norm, which I believe means we will not account for any model parallelism.

Credit to @ShadenSmith for raising this issue while reviewing #1801.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by comparing the five norm implementations named in utils.py with the implementations in stage1/2 and stage3. Trace how MOE and model-parallel attributes flow through clip_grad_norm_, get_grad_norm, get_weight_norm, get_global_norm_of_tensors, and get_global_norm. Done means the implementations are consolidated without losing the required MOE or model-parallel behavior.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
distributed-systems, machine-learning
Issue type
Refactor
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.