AccelerateHS / AccelerateHS/accelerate

Locally mutable batched access to arrays

Aperta
#216 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
language new feature
Lingua principale
Haskell
Stelle
1k
Fork
135
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

The batched LU decomposition of CUBLAS, namely `getrfBatched`, returns a pivot array, but the batched LU solver `trsmBatched` does not accept a pivot array as parameter. Thus I have to apply the permutations manually that are implied by the pivot array. Unfortunately, the pivot array cannot be translated into a permutation as required by `permute` or `backpermute`. Thus I have to fetch the pivot array from the GPU, convert the pivot array into a permutation vector on the CPU using the mutable vector type of the `vector` library and apply the resulting permutation vector with `backpermute` on the GPU. Here is the conversion function that works on `Vector.Storable.Mutable`:

```
permutationFromPivotsMutableBackward ::
V.Vector Word32 -> MV.MVector s Word32 -> ST s ()
permutationFromPivotsMutableBackward pivots perm = do
zipWithM_
(\k j -> do
MV.write perm k (fromIntegral k)
MV.swap perm k (fromIntegral j))
(iterate (subtract 1) $ V.length pivots - 1) (V.toList $ V.reverse pivots)
```

A batched loop which allows mutations (in this case: element swaps) of an array would be great to have in `accelerate` or at least in `accelerate-cuda`.

Unfortunately, I have no concrete idea, what API is both useful and implementable on CUDA. There must be a separation between the outer parallelizable loop for the batch operation and the inner sequential loop for mutable manipulations.

See according post to the Accelerate mailing list: https://groups.google.com/forum/#!topic/accelerate-haskell/OSb53e4yF4Q

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.