JuliaGPU / JuliaGPU/CUDA.jl

Implement `mapslices` without scalar iteration

Open
#807 6 comments 1 reaction 0 assignees View on GitHub
hard
Dominant language
Julia
Stars
1.4k
Forks
281
Avg merge
1d 7h
Merged PRs (30d)
30

Description

When I use `mapslices(f,a,dims)` to manipulate CuArray, a warning appears. It reminds me that using scalar operations on the GPU is inefficient.

```julia
a=CUDA.rand(3,4,5)
b=CUDA.rand(2,3)
mapslices(a,dims=[1,2])do t
b*t
end
```
I had to use additional code to perform the operation.
```julia
c=map(eachslice(a,dims=3)) do t
b*t
end
cat(c...,dims=3)
```
In neural networks or machine learning, mini-batch is often used. When a sample is not a vector or matrix, the input of the model will have multiple dimensions each time, such as size(x)==(100,3,4,batch_size). In this case, `mapslices()` seems very convenient.

However, when the model and input are both CuArray, the GPU will be very inefficient due to too many scalar operations. Can the internal operations of `mapslices()` be optimized to make it more efficient?
**Describe the solution you'd like**

Is there a more elegant way to implement mapslices(f,a,dims) that enables it to use vectorized operations instead of scalar operations.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.