Optimize vector_dup_* on NEON and SSE
Nobody has claimed this yet.
- Dominant language
- C++
- Stars
- 799
- Forks
- 51
- PR merge metrics
- No merged PRs in 30d
Description
To improve inlinability, the SSE version should be explicit and use the shuffle_ps intrinsic directly instead of relying on vector_mix.
The NEON implementation also needs to be added. Even though vector_mix isn't optimized on NEON yet, the compiler has no trouble figuring out what you mean with all 4 components are identical at least with the matrix_mul generated assembly which is quite clean.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Locate the SSE vector_dup_* implementation and the related vector_mix and matrix_mul code, then inspect the generated assembly for the existing behavior. Add the explicit SSE shuffle_ps path and the NEON implementation; done means both targets optimize vector_dup_* as described without relying on vector_mix for SSE.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- cpp
- Domain
- performance
- Issue type
- Refactor
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 35/100