AdamNiederer / AdamNiederer/faster

Run-time feature detection

Open
#2 7 comments 9 reactions 0 assignees View on GitHub
Dominant language
Rust
Stars
1.6k
Forks
52
PR merge metrics
No merged PRs in 30d

Description

Currently the vector size and the SIMD instructions to be used are fixed at compile-time. This is a good first step.

It allows me to write an algorithm once, and by changing a compiler-flag generate different versions of this algorithm for different target architectures (with different vector sizes, SIMD instructions, etc.).

I really like to be able to do this at compile-time within the same binary as well, so that I can write:

```rust
// a generic algorithm
fn my_generic_algorithm(x: T, y: T) {
// generic simd operations
}
```
and monomorphize it for different instruction sets:

```
let x: VecAvx;
my_generic_algorithm(x, x); // AVX version
let y: VecSSE42;
my_generic_algorithm(y, y); // SSE42 version
```

That way `input.simd_iter().map(my_generic_algorithm)` could monomorphize `my_generic_algorithm` for `SSE`, `SSE42`, `AVX`, and `AVX2`, and do run-time feature detection to detect the best that a given CPU supports at run-time, and then dispatch to that one.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.