AccelerateHS / AccelerateHS/accelerate

Locally mutable batched access to arrays

オープン
#216 コメント 2 件 リアクション 0 件 担当者 0 名 GitHub で見る
language new feature
主要言語
Haskell
スター
1k
フォーク
135
PR マージ指標
30日以内にマージされた PR はありません

説明

The batched LU decomposition of CUBLAS, namely `getrfBatched`, returns a pivot array, but the batched LU solver `trsmBatched` does not accept a pivot array as parameter. Thus I have to apply the permutations manually that are implied by the pivot array. Unfortunately, the pivot array cannot be translated into a permutation as required by `permute` or `backpermute`. Thus I have to fetch the pivot array from the GPU, convert the pivot array into a permutation vector on the CPU using the mutable vector type of the `vector` library and apply the resulting permutation vector with `backpermute` on the GPU. Here is the conversion function that works on `Vector.Storable.Mutable`:

```
permutationFromPivotsMutableBackward ::
V.Vector Word32 -> MV.MVector s Word32 -> ST s ()
permutationFromPivotsMutableBackward pivots perm = do
zipWithM_
(\k j -> do
MV.write perm k (fromIntegral k)
MV.swap perm k (fromIntegral j))
(iterate (subtract 1) $ V.length pivots - 1) (V.toList $ V.reverse pivots)
```

A batched loop which allows mutations (in this case: element swaps) of an array would be great to have in `accelerate` or at least in `accelerate-cuda`.

Unfortunately, I have no concrete idea, what API is both useful and implementable on CUDA. There must be a separation between the outer parallelizable loop for the batch operation and the inner sequential loop for mutable manipulations.

See according post to the Accelerate mailing list: https://groups.google.com/forum/#!topic/accelerate-haskell/OSb53e4yF4Q

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。