[X86] scalar fccmp/fcmp chain should prefer keep in FPU/mask domain
- Dominant language
- LLVM
- Stars
- 40.5k
- Forks
- 18.7k
- PR merge metrics
- PR metrics pending
Description
```c++
bool foo(float a, float b, float c, float d, float e, float f){
return a > b & a > c & a > d & a > e & a > f;
}
bool foo2(float a, float b, float c, float d, float e, float f){
return a > b & a < c & a == d & a != e & a >= f;
}
bool foo3(float a, float b, float c, float d, float e, float f){
return a > b || a < c || a == d || a != e || a >= f;
}
```
```asm
Iterations: 100
Instructions: 5100
Total Cycles: 1511
Total uOps: 6000
Dispatch Width: 6
uOps Per Cycle: 3.97
IPC: 3.38
Block RThroughput: 15.0
Instruction Info:
[1]: #uOps
[2]: Latency
[3]: RThroughput
[4]: MayLoad
[5]: MayStore
[6]: HasSideEffects (U)
[1] [2] [3] [4] [5] [6] Instructions:
2 4 1.00 vucomiss xmm0, xmm3
1 2 1.00 vcmpltss k0, xmm2, xmm0
1 2 1.00 vcmpltss k1, xmm1, xmm0
1 1 0.50 kandw k0, k1, k0
1 1 0.50 kmovd eax, k0
1 1 1.00 seta cl
1 1 0.25 and cl, al
2 4 1.00 vucomiss xmm0, xmm4
1 1 1.00 seta dl
2 4 1.00 vucomiss xmm0, xmm5
1 1 1.00 seta al
1 1 0.25 and al, dl
1 1 0.25 and al, cl
1 5 0.50 U ret
2 4 1.00 vucomiss xmm0, xmm3
1 2 1.00 vcmpltss k0, xmm0, xmm2
1 2 1.00 vcmpltss k1, xmm1, xmm0
1 1 0.50 kandw k0, k1, k0
1 1 0.50 kmovd eax, k0
1 1 1.00 setnp cl
1 1 1.00 sete dl
1 1 0.25 and dl, cl
1 1 0.25 and dl, al
2 4 1.00 vucomiss xmm0, xmm4
1 1 1.00 setp al
1 1 1.00 setne cl
1 1 0.25 or cl, al
2 4 1.00 vucomiss xmm0, xmm5
1 1 1.00 setae al
1 1 0.25 and al, cl
1 1 0.25 and al, dl
1 5 0.50 U ret
2 4 1.00 vucomiss xmm0, xmm3
1 2 1.00 vcmpltss k0, xmm0, xmm2
1 2 1.00 vcmpltss k1, xmm1, xmm0
1 1 0.50 korw k0, k1, k0
1 1 0.50 kmovd eax, k0
1 1 1.00 setnp cl
1 1 1.00 sete dl
1 1 0.25 and dl, cl
1 1 0.25 or dl, al
2 4 1.00 vucomiss xmm0, xmm4
1 1 1.00 setp al
1 1 1.00 setne cl
1 1 0.25 or cl, al
2 4 1.00 vucomiss xmm0, xmm5
1 1 1.00 setae al
1 1 0.25 or al, cl
1 1 0.25 or al, dl
1 1 1.00 and al, 1
1 5 0.50 U ret
```
should be like this:
```c++
typedef float f32 [[clang::ext_vector_type(2)]];
bool bar(f32 a, f32 b, f32 c, f32 d, f32 e, f32 f){
return (a > b & a > c & a > d & a > e & a > f)[0];
}
bool bar2(f32 a, f32 b, f32 c, f32 d, f32 e, f32 f){
return (a > b & a < c & a == d & a != e & a >= f)[0];
}
bool bar3(f32 a, f32 b, f32 c, f32 d, f32 e, f32 f){
return (a > b || a < c || a == d || a != e || a >= f)[0];
}
```
```asm
Iterations: 100
Instructions: 3600
Total Cycles: 1507
Total uOps: 3600
Dispatch Width: 6
uOps Per Cycle: 2.39
IPC: 2.39
Block RThroughput: 15.0
Instruction Info:
[1]: #uOps
[2]: Latency
[3]: RThroughput
[4]: MayLoad
[5]: MayStore
[6]: HasSideEffects (U)
[1] [2] [3] [4] [5] [6] Instructions:
1 2 1.00 vcmpltss k0, xmm1, xmm0
1 2 1.00 vcmpltss k1, xmm2, xmm0
1 2 1.00 vcmpltss k5, xmm3, xmm0
1 2 1.00 vcmpltss k6, xmm4, xmm0
1 2 1.00 vcmpltss k7, xmm5, xmm0
1 1 0.50 kandw k0, k0, k1
1 1 0.50 kandw k0, k0, k5
1 1 0.50 kandw k0, k0, k6
1 1 0.50 kandw k0, k0, k7
1 1 0.50 kmovd eax, k0
1 1 1.00 and al, 1
1 5 0.50 U ret
1 2 1.00 vcmpltss k0, xmm1, xmm0
1 2 1.00 vcmpltss k1, xmm0, xmm2
1 2 1.00 vcmpeqss k5, xmm0, xmm3
1 2 1.00 vcmpneqss k6, xmm0, xmm4
1 2 1.00 vcmpless k7, xmm5, xmm0
1 1 0.50 kandw k0, k0, k1
1 1 0.50 kandw k0, k0, k5
1 1 0.50 kandw k0, k0, k6
1 1 0.50 kandw k0, k0, k7
1 1 0.50 kmovd eax, k0
1 1 1.00 and al, 1
1 5 0.50 U ret
1 2 1.00 vcmpltss k0, xmm1, xmm0
1 2 1.00 vcmpltss k1, xmm0, xmm2
1 2 1.00 vcmpeqss k5, xmm0, xmm3
1 2 1.00 vcmpneqss k6, xmm0, xmm4
1 2 1.00 vcmpless k7, xmm5, xmm0
1 1 0.50 korw k0, k0, k1
1 1 0.50 korw k0, k0, k5
1 1 0.50 korw k0, k0, k6
1 1 0.50 korw k0, k0, k7
1 1 0.50 kmovd eax, k0
1 1 1.00 and al, 1
1 5 0.50 U ret
```
https://godbolt.org/z/qMeMTnWeh
Contributor guide
Research direction
Start with the scalar foo, foo2, and foo3 reproducers and compare their generated x86 assembly with the vector bar variants using the linked Godbolt example. Trace the LLVM x86 code-generation path responsible for floating-point comparison chains; done means scalar chains keep comparisons and boolean combinations in the FPU/mask domain without the extra setcc and integer-register operations shown.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- compilers, performance
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Active
- Clarity
- Mostly clear
- Newbie friendliness
- 48/100