AdamNiederer / AdamNiederer/faster
Run-time feature detection
- Lenguaje dominante
- Rust
- Estrellas
- 1.6k
- Forks
- 52
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
Currently the vector size and the SIMD instructions to be used are fixed at compile-time. This is a good first step.
It allows me to write an algorithm once, and by changing a compiler-flag generate different versions of this algorithm for different target architectures (with different vector sizes, SIMD instructions, etc.).
I really like to be able to do this at compile-time within the same binary as well, so that I can write:
```rust
// a generic algorithm
fn my_generic_algorithm(x: T, y: T) {
// generic simd operations
}
```
and monomorphize it for different instruction sets:
```
let x: VecAvx;
my_generic_algorithm(x, x); // AVX version
let y: VecSSE42;
my_generic_algorithm(y, y); // SSE42 version
```
That way `input.simd_iter().map(my_generic_algorithm)` could monomorphize `my_generic_algorithm` for `SSE`, `SSE42`, `AVX`, and `AVX2`, and do run-time feature detection to detect the best that a given CPU supports at run-time, and then dispatch to that one.
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Evaluación
Este issue todavía no se ha evaluado.