NVIDIA / NVIDIA/Megatron-LM

[QUESTION] Why BF16 has FP32 for param gradient at the first time unlike FP16 has FP16 for param gradient at the first time and switches them to FP32 when update param?

Open
#1,335 1 comment 0 reactions 0 assignees View on GitHub
question
Dominant language
Python
Stars
17.9k
Forks
4.5k
Avg merge
4d 6h
Merged PRs (30d)
271

Description

Hi, All

I saw that Megatron-LM supports the configuration of **FP16** and **BF16** mixed-precision data types, and I found that these two different data type configurations correspond to different parameter types and gradient types:

| | **FP16 Training** | **BF16 Training** |
|------ | ----------- | ----------- |
| Weight | FP16| BF16 |
| Gradient | FP16 | **FP32** |

- When training with FP16, a copy of the FP32 gradient is required before updating the parameters
- When training with BF16, no additional copies are required because the gradient itself is FP32

So, why BF16 has FP32 for param gradient at the first time unlike FP16 has FP16 for param gradient at the first time and switches them to FP32 when update param?

Anyone's reply will be helpful to me, thanks

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.