AccelerateHS / AccelerateHS/accelerate

Locally mutable batched access to arrays

未关闭
#216 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
language new feature
主要语言
Haskell
星标
1k
派生
135
PR 合并指标
30 天内没有已合并 PR

描述

The batched LU decomposition of CUBLAS, namely `getrfBatched`, returns a pivot array, but the batched LU solver `trsmBatched` does not accept a pivot array as parameter. Thus I have to apply the permutations manually that are implied by the pivot array. Unfortunately, the pivot array cannot be translated into a permutation as required by `permute` or `backpermute`. Thus I have to fetch the pivot array from the GPU, convert the pivot array into a permutation vector on the CPU using the mutable vector type of the `vector` library and apply the resulting permutation vector with `backpermute` on the GPU. Here is the conversion function that works on `Vector.Storable.Mutable`:

```
permutationFromPivotsMutableBackward ::
V.Vector Word32 -> MV.MVector s Word32 -> ST s ()
permutationFromPivotsMutableBackward pivots perm = do
zipWithM_
(\k j -> do
MV.write perm k (fromIntegral k)
MV.swap perm k (fromIntegral j))
(iterate (subtract 1) $ V.length pivots - 1) (V.toList $ V.reverse pivots)
```

A batched loop which allows mutations (in this case: element swaps) of an array would be great to have in `accelerate` or at least in `accelerate-cuda`.

Unfortunately, I have no concrete idea, what API is both useful and implementable on CUDA. There must be a separation between the outer parallelizable loop for the batch operation and the inner sequential loop for mutable manipulations.

See according post to the Accelerate mailing list: https://groups.google.com/forum/#!topic/accelerate-haskell/OSb53e4yF4Q

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。