zero_call_used_regs("all") zeroes XMM0-15 three times
- Dominant language
- LLVM
- Stars
- 40.5k
- Forks
- 18.7k
- PR merge metrics
- PR metrics pending
Description
Testcase:
```c++
[[gnu::zero_call_used_regs("all")]]
int f()
{
return 1;
}
```
This produces:
```asm
f():
movl $1, %eax
fldz
fldz
fldz
fldz
fldz
fldz
fldz
fldz
fstp %st(0)
fstp %st(0)
fstp %st(0)
fstp %st(0)
fstp %st(0)
fstp %st(0)
fstp %st(0)
fstp %st(0)
xorl %ecx, %ecx
xorl %edi, %edi
xorl %edx, %edx
xorl %esi, %esi
xorl %r8d, %r8d
xorl %r9d, %r9d
xorl %r10d, %r10d
xorl %r11d, %r11d
xorl %r16d, %r16d
xorl %r17d, %r17d
xorl %r18d, %r18d
xorl %r19d, %r19d
xorl %r20d, %r20d
xorl %r21d, %r21d
xorl %r22d, %r22d
xorl %r23d, %r23d
xorl %r24d, %r24d
xorl %r25d, %r25d
xorl %r26d, %r26d
xorl %r27d, %r27d
xorl %r28d, %r28d
xorl %r29d, %r29d
xorl %r30d, %r30d
xorl %r31d, %r31d
vxorps %xmm0, %xmm0, %xmm0
vxorps %xmm1, %xmm1, %xmm1
vxorps %xmm2, %xmm2, %xmm2
vxorps %xmm3, %xmm3, %xmm3
vxorps %xmm4, %xmm4, %xmm4
vxorps %xmm5, %xmm5, %xmm5
vxorps %xmm6, %xmm6, %xmm6
vxorps %xmm7, %xmm7, %xmm7
vxorps %xmm8, %xmm8, %xmm8
vxorps %xmm9, %xmm9, %xmm9
vxorps %xmm10, %xmm10, %xmm10
vxorps %xmm11, %xmm11, %xmm11
vxorps %xmm12, %xmm12, %xmm12
vxorps %xmm13, %xmm13, %xmm13
vxorps %xmm14, %xmm14, %xmm14
vxorps %xmm15, %xmm15, %xmm15
vxorps %xmm0, %xmm0, %xmm0
vxorps %xmm1, %xmm1, %xmm1
vxorps %xmm2, %xmm2, %xmm2
vxorps %xmm3, %xmm3, %xmm3
vxorps %xmm4, %xmm4, %xmm4
vxorps %xmm5, %xmm5, %xmm5
vxorps %xmm6, %xmm6, %xmm6
vxorps %xmm7, %xmm7, %xmm7
vxorps %xmm8, %xmm8, %xmm8
vxorps %xmm9, %xmm9, %xmm9
vxorps %xmm10, %xmm10, %xmm10
vxorps %xmm11, %xmm11, %xmm11
vxorps %xmm12, %xmm12, %xmm12
vxorps %xmm13, %xmm13, %xmm13
vxorps %xmm14, %xmm14, %xmm14
vxorps %xmm15, %xmm15, %xmm15
kxorq %k0, %k0, %k0
kxorq %k0, %k0, %k1
kxorq %k0, %k0, %k2
kxorq %k0, %k0, %k3
kxorq %k0, %k0, %k4
kxorq %k0, %k0, %k5
kxorq %k0, %k0, %k6
kxorq %k0, %k0, %k7
vxorps %xmm0, %xmm0, %xmm0
vxorps %xmm1, %xmm1, %xmm1
vxorps %xmm2, %xmm2, %xmm2
vxorps %xmm3, %xmm3, %xmm3
vxorps %xmm4, %xmm4, %xmm4
vxorps %xmm5, %xmm5, %xmm5
vxorps %xmm6, %xmm6, %xmm6
vxorps %xmm7, %xmm7, %xmm7
vxorps %xmm8, %xmm8, %xmm8
vxorps %xmm9, %xmm9, %xmm9
vxorps %xmm10, %xmm10, %xmm10
vxorps %xmm11, %xmm11, %xmm11
vxorps %xmm12, %xmm12, %xmm12
vxorps %xmm13, %xmm13, %xmm13
vxorps %xmm14, %xmm14, %xmm14
vxorps %xmm15, %xmm15, %xmm15
vxorps %xmm16, %xmm16, %xmm16
vxorps %xmm17, %xmm17, %xmm17
vxorps %xmm18, %xmm18, %xmm18
vxorps %xmm19, %xmm19, %xmm19
vxorps %xmm20, %xmm20, %xmm20
vxorps %xmm21, %xmm21, %xmm21
vxorps %xmm22, %xmm22, %xmm22
vxorps %xmm23, %xmm23, %xmm23
vxorps %xmm24, %xmm24, %xmm24
vxorps %xmm25, %xmm25, %xmm25
vxorps %xmm26, %xmm26, %xmm26
vxorps %xmm27, %xmm27, %xmm27
vxorps %xmm28, %xmm28, %xmm28
vxorps %xmm29, %xmm29, %xmm29
vxorps %xmm30, %xmm30, %xmm30
vxorps %xmm31, %xmm31, %xmm31
vzeroupper
retq
```
This is a minor performance issue.
I guess what happened is that Clang is emitting zeroing of XMM0-15 (SSE), then zeroing of YMM0-15 (AVX), then zeroing of ZMM0-31 (AVX512/AVX10), but it emitted assembly for XMM-sized registers because they have the same effect. And it emitted a useless VZEROUPPER at the end because the function used YMM/ZMM registers... to zero them.
Bonus: optimal code would do VZEROALL + VXORPS for XMM16-31.
Contributor guide
Assessment
This issue has not been assessed yet.