AccelerateHS / AccelerateHS/accelerate

Performance issues

Đang mở
#305 14 bình luận 0 reaction 0 người được giao Xem trên GitHub
llvm backend
Ngôn ngữ chính
Haskell
Star
1k
Fork
135
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

I'm trying to use accelerate to get better-than-hmatrix performance on a program similar to this one:

``` haskell
module Main where

import Data.IORef
import Data.Array.Accelerate
import Data.Array.Accelerate.CUDA
--import Data.Array.Accelerate.Interpreter

import Prelude hiding (zipWith, replicate, (++), sum, map)

type Matrix e = Array DIM2 e

dot :: (Elt a, IsNum a) => Acc (Vector a) -> Acc (Vector a) -> Acc (Scalar a)
dot xs ys = fold1 (+) (zipWith (*) xs ys)

norm xs = map sqrt (dot xs xs)

mult
:: (Elt a, IsNum a) =>
Acc (Matrix a) -> Acc (Vector a) -> Acc (Vector a)

mult m v = fold1 (+) (zipWith (*) m vs)
where
n = indexHead . indexTail $ shape m
vs = replicate (lift $ Z :. n :. All) v

main = do
let n = 1000 :: Int

let m = fill (constant $ Z :. n :. n) (constant (1 :: Float))
vRef <- newIORef $ fill (constant $ Z :. n) (constant 1)

let iterate = do
v <- readIORef vRef
let mv = mult m v
vv = norm v
v' = map (/ the vv) mv
--writeIORef vRef (compute v')
writeIORef vRef v'

traverse (const iterate) [1..1000]

v <- readIORef vRef
print $ run (norm v)
```

Here's the output of profiling:

```
Tue Feb 9 14:27 2016 Time and Allocation Profiling Report (Final)

TestAccelerate +RTS -p -h -s -RTS

total time = 4.03 secs (4032 ticks @ 1000 us, 1 processor)
total alloc = 4,114,779,024 bytes (excludes profiling overheads)

COST CENTRE MODULE %time %alloc

compile.code Data.Array.Accelerate.CUDA.Compile 15.1 33.1
ppr Data.Array.Accelerate.CUDA.CodeGen.Base 10.1 11.1
hashWithSalt Data.Array.Accelerate.CUDA.CodeGen.Monad 2.8 5.3
lookup/go Data.HashTable.ST.Basic 1.6 0.1
compile.key Data.Array.Accelerate.CUDA.Compile 1.5 3.7
rebuildPreOpenExp Data.Array.Accelerate.Trafo.Substitution 1.4 1.3
== Data.Array.Accelerate.Trafo.Sharing 1.4 1.8
codegenOpenExp.cvtE Data.Array.Accelerate.CUDA.CodeGen 1.2 1.2
withLifetime Data.Array.Accelerate.Lifetime 1.1 0.1
insert Data.HashTable.ST.Basic 1.1 0.3
readDelLoad Data.HashTable.ST.Basic 1.1 0.6
blockSize Data.Array.Accelerate.CUDA.Analysis.Launch 0.8 1.7
fmap Data.Array.Accelerate.Trafo.Substitution 0.8 1.2
readArray Data.HashTable.Internal.IntArray 0.7 1.0
```

And here's the output of -s:

```
6,372,076,320 bytes allocated in the heap
3,751,952,272 bytes copied during GC
164,069,960 bytes maximum residency (49 sample(s))
2,371,928 bytes maximum slop
328 MB total memory in use (0 MB lost due to fragmentation)

Tot time (elapsed) Avg pause Max pause
Gen 0 12011 colls, 0 par 3.477s 3.440s 0.0003s 0.0012s
Gen 1 49 colls, 0 par 5.114s 5.071s 0.1035s 0.2412s

TASKS: 6 (2 bound, 4 peak workers (4 total), using -N1)

SPARKS: 0 (0 converted, 0 overflowed, 0 dud, 0 GC'd, 0 fizzled)

INIT time 0.002s ( 0.002s elapsed)
MUT time 4.539s ( 6.391s elapsed)
GC time 6.669s ( 6.604s elapsed)
RP time 0.000s ( 0.000s elapsed)
PROF time 1.922s ( 1.907s elapsed)
EXIT time 0.003s ( 0.003s elapsed)
Total time 13.140s ( 12.999s elapsed)

Alloc rate 1,403,793,613 bytes per MUT second

Productivity 34.6% of total user, 35.0% of total elapsed

gc_alloc_block_sync: 0
whitehole_spin: 0
gen[0].sync: 0
gen[1].sync: 0
```

As you can see, a huge amount of time is spent in GC. That clearly should not be the case for a program that's just number crunching.

An obvious potential issue is that I'm building a very large kernel with each iteration and then running only at the end. The solution would be to express the iteration as an Acc -> Acc function and then use awhile to run it a fixed number of times. Unfortunately, it is very difficult for me to turn my real program into a pure Acc -> Acc function, since it does a large graph traversal that relies on IORefs. I've also tried inserting a compute (`id >-> id`) into each iteration, with no effect.

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.