Use fp16 performance decrease
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 9k
- Forks
- 1.5k
- Avg merge
- 2d 4h
- Merged PRs (30d)
- 3
Description
Hi, I use apex op_level O2, and do some cast in some places, but the performance decrease.

Any way to find where the precision decrease caused by fp16 (or by casting float to float16)?
Besides, will FFNLayer(3 * hidden_size, hidden_size, 2, dropout_prob)only accept float16 input when using amp?

Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start with the Apex AMP op_level O2 configuration and the FFNLayer(3 * hidden_size, hidden_size, 2, dropout_prob) call shown in the report. Reproduce the fp16 performance decrease while checking the reported casts and input precision; done means identifying whether the slowdown comes from fp16 or casting and clarifying the accepted input type.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python, pytorch
- Domain
- machine-learning, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100