nfrechette / nfrechette/rtm

Optimize vector_dup_* on NEON and SSE

Open
#3 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

enhancement good first issue help wanted
Dominant language
C++
Stars
799
Forks
51
PR merge metrics
No merged PRs in 30d

Description

To improve inlinability, the SSE version should be explicit and use the shuffle_ps intrinsic directly instead of relying on vector_mix.

The NEON implementation also needs to be added. Even though vector_mix isn't optimized on NEON yet, the compiler has no trouble figuring out what you mean with all 4 components are identical at least with the matrix_mul generated assembly which is quite clean.

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Locate the SSE vector_dup_* implementation and the related vector_mix and matrix_mul code, then inspect the generated assembly for the existing behavior. Add the explicit SSE shuffle_ps path and the NEON implementation; done means both targets optimize vector_dup_* as described without relying on vector_mix for SSE.

Written by the indexing model from the issue text.

Assessment

Tech stack
cpp
Domain
performance
Issue type
Refactor
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
35/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.