llvm / llvm/llvm-project

Use `pmovmskb` instead of `@llvm.vector.reduce.or.v16i8` where possible

Open
#174,504 4 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

backend:X86 missed-optimization
Dominant language
LLVM
Stars
40.5k
Forks
18.7k
PR merge metrics
PR metrics pending

Description

Same as https://github.com/llvm/llvm-project/issues/174500, but x86 doesn't have horizontal umax either. However, we can use pmovmskb to extract sign bit from each byte:

auto src16(u8* p) {
    auto ret = true;
    for (usize i = 0; i < 16; i++) {
        ret &= p[i] < 0x80;
    }
    return ret;
}

auto tgt16(i8x16* p) {
    i8x16 v0 = p[0];
    return _mm_movemask_epi8(v0) == 0;
}
define dso_local noundef zeroext i1 @src16(unsigned char*)(ptr noundef readonly captures(none) %0) local_unnamed_addr #0 {
  %2 = load <16 x i8>, ptr %0, align 1
  %3 = tail call i8 @llvm.vector.reduce.or.v16i8(<16 x i8> %2)
  %4 = icmp sgt i8 %3, -1
  ret i1 %4
}

define dso_local noundef zeroext i1 @tgt16(signed char vector[16]*)(ptr noundef readonly captures(none) %0) local_unnamed_addr #1 {
  %2 = load <16 x i8>, ptr %0, align 16
  %3 = icmp slt <16 x i8> %2, zeroinitializer
  %4 = bitcast <16 x i1> %3 to i16
  %5 = icmp eq i16 %4, 0
  ret i1 %5
}
src16(unsigned char*):
        movdqu  xmm0, xmmword ptr [rdi]
        pshufd  xmm1, xmm0, 238
        por     xmm1, xmm0
        pshufd  xmm0, xmm1, 85
        por     xmm0, xmm1
        movdqa  xmm1, xmm0
        psrld   xmm1, 16
        por     xmm1, xmm0
        movdqa  xmm0, xmm1
        psrlw   xmm0, 8
        por     xmm0, xmm1
        movd    eax, xmm0
        not     al
        shr     al, 7
        ret

tgt16(signed char vector[16]*):
        movdqa  xmm0, xmmword ptr [rdi]
        pmovmskb        eax, xmm0
        test    eax, eax
        sete    al
        ret

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by comparing the shown src16 and tgt16 LLVM IR and x86 assembly, then trace the lowering of llvm.vector.reduce.or.v16i8 for x86. Done means applicable cases use pmovmskb while preserving the boolean result shown in the examples; the issue names no files or tests.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
compilers, performance
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.