JuliaSIMD / JuliaSIMD/LoopVectorization.jl
Plain for loop faster than @turbo
- Dominant language
- Julia
- Stars
- 789
- Forks
- 73
- PR merge metrics
- No merged PRs in 30d
Description
Consider the following:
```julia
using LoopVectorization, StrideArraysCore, BenchmarkTools
randn_stridearray(size...) = StrideArray(randn(Float32, size...), static.(size))
y = randn_stridearray(64, 32, 256);
x = randn_stridearray(64, 32, 256);
b = randn_stridearray(32);
function foo!(y, x, b)
@turbo for i in axes(y, 1), c in axes(y, 2), n in axes(y, 3)
y[i, c, n] = x[i, c, n] + b[c]
end
return nothing
end
function foo2!(y, x, b)
for n in axes(y, 3)
@turbo for i in axes(y, 1), c in axes(y, 2)
y[i, c, n] = x[i, c, n] + b[c]
end
end
return nothing
end
```
I would expect `foo2` to perform slightly worse or the same as `foo`.
However, the result is
```
julia> @btime foo!($y, $x, $b)
153.958 μs (0 allocations: 0 bytes)
julia> @btime foo2!($y, $x, $b)
51.583 μs (0 allocations: 0 bytes)
```
What is happening here? Am I doing something wrong?
If it matters: I am running this on an Apple M1.
Edit:
A plain for loop is about as fast as `foo2!`:
```
julia> function foo3!(y, x, b)
for n in axes(y, 3)
for c in axes(y, 2)
for i in axes(y, 1)
y[i, c, n] = x[i, c, n] + b[c]
end
end
end
return nothing
end
foo3! (generic function with 1 method)
julia> @btime foo3!($y, $x, $b)
52.375 μs (0 allocations: 0 bytes)
```
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.