google / google/highway

Should FMA optimizations be implemented for SCALAR/EMU128 on PPC/RISC-V/GPU?

Open
#2,542 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
C++
Stars
5.8k
Forks
471
Avg merge
1d 6h
Merged PRs (30d)
81

Description

There are some ISA's that have hardware FMA instructions that are guaranteed to be available if hardware floating-point support is present, including PPC, RISC-V (on CPU's implementing the "F" and "D" extension), NVidia GPU's, and AMD GPU's.

GCC/Clang also has the __builtin_fma and __builtin_fmaf builtins that are guaranteed to be compiled down to a single FMA instruction on ISA's with hardware floating point that can carry out FMA using a single instruction, even with optimizations disabled (-O0).

There are use cases for implementing MulAdd/NegMulAdd/MulSub/NegMulSub using __builtin_fma and __builtin_fmaf for SCALAR/EMU128 on RISC-V CPU's that have the RISC-V F and D extensions (but might not necessarily support the "V" SIMD extension), NVidia GPU's, and AMD GPU's.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.