AccelerateHS / AccelerateHS/accelerate

Performance of highly skewed multidimensional reductions

Abierto
#140 7 comentarios 0 reacciones 0 asignados Ver en GitHub
good first issue help wanted llvm-ptx
Lenguaje dominante
Haskell
Estrellas
1k
Forks
135
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

Performance of multidimensional reductions is not good when the array is highly skewed. For example, a `fold` where the number of columns is (innermost dimension) is very small. See also this thread:

https://groups.google.com/forum/#!topic/accelerate-haskell/KAFYUz4Sjsk

Multidimensional reduction uses one thread block per reduction; so an `(Z :. m :. n)` sized matrix uses `m` thread blocks. If `n` is very small, then many threads in the block sit idle. We could change this to a warp-per-reduction style, which is actually the strategy segmented fold uses. This will likely have a negative impact if `m` is small and `n` large.

It would be possible to generate both variants and choose dynamically which to execute. That implies compiling four kernels per reduction (because fusion; initial vs. recursive step).

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.