arrayfire / arrayfire/arrayfire-python
Poor slicing performance compared to NumPy
- Lenguaje dominante
- Python
- Estrellas
- 422
- Forks
- 63
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
reported by @floopcz on over here: https://github.com/arrayfire/arrayfire/issues/1428
ArrayFire slicing seems to suffer from a performance issue. Consider the following python code, that:
- calculates the dot product of two matrices, first using NumPy, than ArrayFire
- calculates each column/row of the dot product separately by slicing a single column/row from one of the matrices
``` python
#!/usr/bin/env python3
from time import time
import arrayfire as af
import numpy as np
af.set_backend('cpu')
af.info()
iters = 1000
n = 512
af_A = af.randu(n, n)
af_B = af.randu(n, n)
np_A = np.random.rand(n, n).astype(np.float32)
np_B = np.random.rand(n, n).astype(np.float32)
start = time()
for t in range(iters):
np_C = np.dot(np_A, np_B)
print('numpy - dot: {}'.format(time() - start))
af.sync()
start = time()
for t in range(iters):
af_C = af.matmul(af_A, af_B)
af.sync()
print('arrayfire - matmul: {}'.format(time() - start))
start = time()
for t in range(iters):
for i in range(np_B.shape[1]):
np_C = np.dot(np_A, np_B[:, i])
print('numpy - sliced dot - column major: {}'.format(time() - start))
af.sync()
start = time()
for t in range(iters):
for i in range(af_B.shape[1]):
af_C = af.matmul(af_A, af_B[:, i])
af.sync()
print('arrayfire - sliced matmul - column major: {}'.format(time() - start))
start = time()
for t in range(iters):
for i in range(np_B.shape[0]):
np_C = np.dot(np_B[i, :], np_A)
print('numpy - sliced dot - row major: {}'.format(time() - start))
af.sync()
start = time()
for t in range(iters):
for i in range(af_B.shape[0]):
af_C = af.matmul(af_B[i, :], af_A)
af.sync()
print('arrayfire - sliced matmul - row major: {}'.format(time() - start))
```
The results are following:
```
ArrayFire v3.3.2 (CPU, 64-bit Linux, build f65dd97)
[0] Unknown: Unknown, 15880 MB, Max threads(1)
numpy - dot: 1.3848536014556885
arrayfire - matmul: 1.325775146484375
numpy - sliced dot - column major: 7.156768798828125
arrayfire - sliced matmul - column major: 38.87605834007263
numpy - sliced dot - row major: 7.6784679889678955
arrayfire - sliced matmul - row major: 41.27544379234314
```
The results suggest that with slicing, arrayfire performance is significantly degraded compared to NumPy. I have achieved similarly distributed results also with the GPU backend. Both numpy and arrayfire are linked against Intel MKL.
Am I doing something "illegal" or is it an inefficiency of the library? Thanks.
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Línea de trabajo
Start by running the supplied Python benchmark with the CPU and GPU backends, focusing on af_B[:, i], af_B[i, :], and af.matmul. No source file or test is named in the issue, so trace those entry points through the Python bindings and compare them with unsliced matmul. Done means reproducing the regression, identifying its cause, and adding a regression test or benchmark that demonstrates the improvement.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- python
- Área
- performance
- Tipo de issue
- Error
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Estado de actividad
- Estancado
- Claridad
- Necesita aclaración
- Aptitud para principiantes
- 25/100