JuliaSIMD / JuliaSIMD/LoopVectorization.jl

Plot against theoretical peak performance

Ouverte
#355 23 commentaires 0 réactions 0 personnes assignées Voir sur GitHub
Langage dominant
Julia
Étoiles
789
Forks
73
Métriques de merge des PR
Aucune PR mergée en 30 j

Description

I started with these plots:

https://juliasimd.github.io/LoopVectorization.jl/stable/examples/matrix_vector_ops/

The issue is that I can't tell from the plot if the performance is already the best you can do, or if there is still room for improvement. To answer that question, I like to plot clock cycles (per loop body / array element) on the vertical axis, and also plot the theoretical performance peak into the same graph. Let's do that for the very first benchmark:
```console
julia> function jgemvavx!(𝐲, 𝐀, 𝐱)
@turbo for i ∈ eachindex(𝐲)
𝐲i = zero(eltype(𝐲))
for j ∈ eachindex(𝐱)
𝐲i += 𝐀[i,j] * 𝐱[j]
end
𝐲[i] = 𝐲i
end
end
jgemvavx! (generic function with 1 method)

julia> N = 64
64

julia> y = rand(N);

julia> A = rand(N, N);

julia> x = rand(N);

julia> @benchmark jgemvavx!(y, A, x)
BenchmarkTools.Trial: 10000 samples with 199 evaluations.
Range (min … max): 427.970 ns … 610.970 ns ┊ GC (min … max): 0.00% … 0.00%
Time (median): 428.603 ns ┊ GC (median): 0.00%
Time (mean ± σ): 430.210 ns ± 9.385 ns ┊ GC (mean ± σ): 0.00% ± 0.00%

▇█▅▁ ▃ ▂
████▅▄▄▄▃▄▅▄▅▅▅▅▆▄▅▄▃▁▃███▇▆▄▁▄▁▄▁▄▄▄▅▅▅▆▇▅▃▄▁▁▃▁▁▁▁▁▁▁▃▁▅▅▅▅ █
428 ns Histogram: log(frequency) by time 463 ns <

Memory estimate: 0 bytes, allocs estimate: 0.

julia> cpu_freq = 3.2e9 # Ghz
3.2e9

julia> 427.136e-9 * cpu_freq / N^2
0.3337
```

So for `N=64`, using double precision, it seems the body of the loop takes 0.333 clock cycles.

I use Apple M1. I installed the macOS Julia, I assume that is Intel based? Somehow it works for me.

Regarding the theoretical peak, the loop body is just:
```julia
𝐲i += 𝐀[i,j] * 𝐱[j]
```
Everything else I think gets amortized. So it is two memory reads from L1 cache, each takes 0.1665 clock cycles (`ldr q0, [x1]` takes 0.333). And one fma, which takes 0.125 (`fmla.2d v0, v0, v0` takes 0.25). The bottleneck in this case is the memory read, total of 0.333 clock cycles. The fma runs at the same time, so we do not include it.

So according to my analysis, the code runs at 100% of the theoretical peak. Which sounds too good to be true. I am worried I made some mistake in my analysis somewhere. But I am posting what I have so far.

Guide de contribution

Aucun guide de contribution indexé pour ce dépôt

Piste de recherche

Commencez par l’exemple matrix_vector_ops lié et son premier benchmark. Examinez comment les graphiques existants sont produits, puis déterminez comment afficher les cycles d’horloge par corps de boucle à côté du pic théorique pour le benchmark SIMD de Julia. La tâche est terminée lorsque le graphique de l’exemple rend la comparaison claire et que le calcul du pic est étayé par le contexte du benchmark.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Évaluation

Stack technique
julia
Domaine
data-visualization, performance
Type d'issue
Fonctionnalité
Difficulté
4/5
Temps estimé
3-5 jours
Activité
À l'abandon
Clarté
Plutôt claire
Accessibilité débutants
35/100

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.