AdamNiederer / AdamNiederer/faster
Run-time feature detection
- Langage dominant
- Rust
- Étoiles
- 1.6k
- Forks
- 52
- Métriques de merge des PR
- Aucune PR mergée en 30 j
Description
Currently the vector size and the SIMD instructions to be used are fixed at compile-time. This is a good first step.
It allows me to write an algorithm once, and by changing a compiler-flag generate different versions of this algorithm for different target architectures (with different vector sizes, SIMD instructions, etc.).
I really like to be able to do this at compile-time within the same binary as well, so that I can write:
```rust
// a generic algorithm
fn my_generic_algorithm(x: T, y: T) {
// generic simd operations
}
```
and monomorphize it for different instruction sets:
```
let x: VecAvx;
my_generic_algorithm(x, x); // AVX version
let y: VecSSE42;
my_generic_algorithm(y, y); // SSE42 version
```
That way `input.simd_iter().map(my_generic_algorithm)` could monomorphize `my_generic_algorithm` for `SSE`, `SSE42`, `AVX`, and `AVX2`, and do run-time feature detection to detect the best that a given CPU supports at run-time, and then dispatch to that one.
Guide de contribution
Aucun guide de contribution indexé pour ce dépôt
Évaluation
Cette issue n'a pas encore été évaluée.