arrayfire / arrayfire/arrayfire-python

Poor slicing performance compared to NumPy

Aperta
#84 10 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
422
Fork
63
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

reported by @floopcz on over here: https://github.com/arrayfire/arrayfire/issues/1428

ArrayFire slicing seems to suffer from a performance issue. Consider the following python code, that:
- calculates the dot product of two matrices, first using NumPy, than ArrayFire
- calculates each column/row of the dot product separately by slicing a single column/row from one of the matrices

``` python
#!/usr/bin/env python3
from time import time
import arrayfire as af
import numpy as np

af.set_backend('cpu')
af.info()

iters = 1000
n = 512

af_A = af.randu(n, n)
af_B = af.randu(n, n)

np_A = np.random.rand(n, n).astype(np.float32)
np_B = np.random.rand(n, n).astype(np.float32)

start = time()
for t in range(iters):
np_C = np.dot(np_A, np_B)
print('numpy - dot: {}'.format(time() - start))

af.sync()
start = time()
for t in range(iters):
af_C = af.matmul(af_A, af_B)
af.sync()
print('arrayfire - matmul: {}'.format(time() - start))

start = time()
for t in range(iters):
for i in range(np_B.shape[1]):
np_C = np.dot(np_A, np_B[:, i])
print('numpy - sliced dot - column major: {}'.format(time() - start))

af.sync()
start = time()
for t in range(iters):
for i in range(af_B.shape[1]):
af_C = af.matmul(af_A, af_B[:, i])
af.sync()
print('arrayfire - sliced matmul - column major: {}'.format(time() - start))

start = time()
for t in range(iters):
for i in range(np_B.shape[0]):
np_C = np.dot(np_B[i, :], np_A)
print('numpy - sliced dot - row major: {}'.format(time() - start))

af.sync()
start = time()
for t in range(iters):
for i in range(af_B.shape[0]):
af_C = af.matmul(af_B[i, :], af_A)
af.sync()
print('arrayfire - sliced matmul - row major: {}'.format(time() - start))
```

The results are following:

```
ArrayFire v3.3.2 (CPU, 64-bit Linux, build f65dd97)
[0] Unknown: Unknown, 15880 MB, Max threads(1)
numpy - dot: 1.3848536014556885
arrayfire - matmul: 1.325775146484375
numpy - sliced dot - column major: 7.156768798828125
arrayfire - sliced matmul - column major: 38.87605834007263
numpy - sliced dot - row major: 7.6784679889678955
arrayfire - sliced matmul - row major: 41.27544379234314
```

The results suggest that with slicing, arrayfire performance is significantly degraded compared to NumPy. I have achieved similarly distributed results also with the GPU backend. Both numpy and arrayfire are linked against Intel MKL.

Am I doing something "illegal" or is it an inefficiency of the library? Thanks.

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.