OpenMathLib / OpenMathLib/OpenBLAS

risc-v vector v1.0 support

Offen
#4,050 10 Kommentare 1 Reaktion 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Vorherrschende Sprache
C
Sterne
7.6k
Forks
1.7k
Ø Merge
1 T. 3 Std.
Gemergte PRs (30 T.)
42

Beschreibung

The discussions in #4049 inspire me to creat an issue for further discussions.
Differ from commercial ISAs, which have a clear development plan, the total amount of products supporting rvv may be large. Optimization for all individual products may lead to code bloat, and is contrary to the purpose of the vector isa, which is expected to be length-adaptive.
Until now the intrinsic spec of rvv 1.0 is stable enough to develop codes, and the support of rvv 1.0 has been fully submitted to openblas, based on sifive x280, an in-order cpu with vlen=512.
Would it be better to do more development, based on this x280 version? The final destination may be the compatibility in different vlen, instruction execution order, tail/mask policy. Of course the pursuing of compatibility may lead to suboptimum performance, a balance have to be considered.
There are some cpu specified features in kernels of x280 and may lead to incorrect results in other cpus. List as following

  1. Architecture specified cflags, such as -riscv-v-vector-bits-min=512 and -ffast-math.
  2. Changing vl in a loop, leading to tail cleared without tail undisturbed setted. Such as vl = VSETVL(k); in symv_L_rvv.c, line 96.
  3. Set vl by immediate value under the assumption of vlen=512. Such as size_t vl = 8; in gemm_tcopy_8_rvv.c, line 84.

In addition to above, the registers tiling in gemm of different vlen should be considered. Now we set GEMM_UNROLL_N_SHIFT 8, which may waste other vector registers. 12 or 14 may be better?

Beitragsleitfaden

Für dieses Repository ist kein Beitragsleitfaden indexiert

Erste Schritte

  1. Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
  3. Forke das Repository und arbeite in einem Branch.
  4. Öffne einen Pull Request, der die Issue-Nummer nennt.

Rechercherichtung

Beginne mit den im Issue erwähnten RVV-Kernels, insbesondere mit symv_L_rvv.c Zeile 96 und gemm_tcopy_8_rvv.c Zeile 84, und überprüfe die x280-spezifischen Compiler-Flags. Vergleiche die aktuelle Einstellung GEMM_UNROLL_N_SHIFT mit verschiedenen VLEN-Annahmen und berücksichtige die Tail-/Masken-Policy sowie die Kompatibilität der Befehlsreihenfolge. Als abgeschlossen gilt die Arbeit, wenn ein abgestimmter Implementierungsumfang für die RVV-1.0-Kompatibilität über verschiedene VLEN-Werte und CPUs hinweg festgelegt ist.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Bewertung

Tech-Stack
c
Bereich
performance
Issue-Typ
Feature
Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Aktivitätsstatus
Veraltet
Klarheit
Muss geklärt werden
Anfängerfreundlichkeit
25/100

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.