OpenMathLib / OpenMathLib/OpenBLAS

OpenBLAS bottlenecks multithreading benefits in `symv.c` interface at 8 working threads due to memory allocator lock conflict

Offen
#5,589 8 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Vorherrschende Sprache
C
Sterne
7.6k
Forks
1.7k
Ø Merge
1 T. 3 Std.
Gemergte PRs (30 T.)
42

Beschreibung

I have code that heavily makes use of the LAPACK syevr functions via Julia. They're relatively small matrices (at most 15x15). I'm processing chunks of video frames across multiple threads, and each thread will perform millions of these operations. I've set BLAS threads to 1, which I understand to mean that OpenBLAS just uses the parent thread calling it. (Setting it to anything more than 1 tanks performance generally.)

However, what I've found is that no matter what size computer I run on, performance gains stop once I reach 8 working threads; even worsening with many more. Somehow it seems that OpenBLAS, without itself doing multithreaded computation, is interfering with higher-level multithreading?

If I switch to MKL with 1 thread, I see continued performance improvements through 48 CPUs.

I'm willing to poke around at this as much as I can myself, I'm just not sure where to begin. Where might the bottleneck be?

Image

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginnen Sie mit der symv.c-Schnittstelle und reproduzieren Sie die gemeldete Arbeitslast mit LAPACK syevr auf kleinen Matrizen, wobei BLAS auf einen Thread beschränkt ist. Vergleichen Sie die Skalierung über verschiedene Anzahlen von Arbeitsthreads mit OpenBLAS und MKL und ermitteln Sie anschließend, ob ein Allocator-Lock oder ein anderer OpenBLAS-Engpass das Plateau bei acht Threads erklärt.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
c
Bereich
performance
Issue-Typ
Bug
Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
35/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.