JuliaSIMD / JuliaSIMD/LoopVectorization.jl

Plain for loop faster than @turbo

Abierto
#449 5 comentarios 0 reacciones 0 asignados Ver en GitHub
Lenguaje dominante
Julia
Estrellas
789
Forks
73
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

Consider the following:

```julia
using LoopVectorization, StrideArraysCore, BenchmarkTools

randn_stridearray(size...) = StrideArray(randn(Float32, size...), static.(size))

y = randn_stridearray(64, 32, 256);
x = randn_stridearray(64, 32, 256);
b = randn_stridearray(32);

function foo!(y, x, b)
@turbo for i in axes(y, 1), c in axes(y, 2), n in axes(y, 3)
y[i, c, n] = x[i, c, n] + b[c]
end
return nothing
end

function foo2!(y, x, b)
for n in axes(y, 3)
@turbo for i in axes(y, 1), c in axes(y, 2)
y[i, c, n] = x[i, c, n] + b[c]
end
end
return nothing
end
```

I would expect `foo2` to perform slightly worse or the same as `foo`.
However, the result is

```
julia> @btime foo!($y, $x, $b)
153.958 μs (0 allocations: 0 bytes)

julia> @btime foo2!($y, $x, $b)
51.583 μs (0 allocations: 0 bytes)
```

What is happening here? Am I doing something wrong?
If it matters: I am running this on an Apple M1.

Edit:

A plain for loop is about as fast as `foo2!`:

```
julia> function foo3!(y, x, b)
for n in axes(y, 3)
for c in axes(y, 2)
for i in axes(y, 1)
y[i, c, n] = x[i, c, n] + b[c]
end
end
end
return nothing
end
foo3! (generic function with 1 method)

julia> @btime foo3!($y, $x, $b)
52.375 μs (0 allocations: 0 bytes)
```

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Evaluación

Este issue todavía no se ha evaluado.

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.