NVIDIA / NVIDIA/TransformerEngine

Questions about accuracy alignment between BF16 and FP8

Open
#1,419 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

question
Dominant language
Python
Stars
3.5k
Forks
831
Avg merge
3d 11h
Merged PRs (30d)
65

Description

Hello developers,

Thanks for introducing such a great library that demonstrate the power of FP8 training.

But when I tried to integrate FP8 training into my training framework. I found it is hard to make the acc/loss aligned between BF16 and FP8. Actually, in my experiments, the difference is somehow not negligible. I thinks the ideal difference should be e-2~e-3, which is the negligible random error.

Also, I cannot found some resources that show the accuracy alignment of BF16 and FP8.

So, could anyone give some hint about this?

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

The report names no files, tests, model, or reproducible configuration. Start by collecting a minimal BF16/FP8 comparison with the training framework and recording the expected tolerance and environment; done means the discrepancy is reproducible and its accuracy impact is explained.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.