llvm / llvm/llvm-project

Suboptimal lowering of float bitwise ops on targets without hardware support

Open
#180,630 2 comments 0 reactions 1 assignee Claimed by @bala-bhargav View on GitHub
llvm:codegen missed-optimization
Dominant language
LLVM
Stars
40.5k
Forks
18.7k
PR merge metrics
PR metrics pending

Description

Demo: https://llvm.godbolt.org/z/7E6xxn47b. Input:

```llvm
define zeroext i1 @foo(half %x) unnamed_addr {
start:
%i = bitcast half %x to i16
%masked = and i16 %i, 32767
%r = icmp eq i16 %masked, 0
ret i1 %r
}
```

Instcombine turns this into:

```llvm
define zeroext i1 @foo(half %x) unnamed_addr {
start:
%r = fcmp oeq half %x, 0xH0000
ret i1 %r
}
```

Then on x86, the following is generated:

```asm
foo:
push rax
call __extendhfsf2@PLT
xorps xmm1, xmm1
cmpeqss xmm1, xmm0
movd eax, xmm1
and eax, 1
pop rcx
ret
```

The bitwise ops would be ~4 instructions. The generated code is significantly worse given the cost of calling `__extendhfsf2`.

This shows up for all float types where there isn't hardware support. For example, https://rust.godbolt.org/z/oYWYnja4q on a soft float arm target has a libcall for all of `half`, `float`, `double`, `fp128` that shouldn't be needed.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.