[Migrated] Conditionals and code generation performance
Dieses Issue hat noch niemand übernommen.
- Vorherrschende Sprache
- Rust
- Sterne
- 3.4k
- Forks
- 125
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
Issue automatically imported from old repo: https://github.com/EmbarkStudios/rust-gpu/issues/1110
Old labels: t: bug
Originally creatd by DGriffin91 on 2023-12-18T19:06:07Z
I noticed that a rust gpu shader was running much slower than the equivalent wgsl one.
The wgsl one takes 53ms, and the rust gpu version takes 67ms.
Looking at the SPIRV I tracked part of the issue down to this:
if uvt.x > 0.0 && uvt.y > 0.0 && uvt.z > 0.0 && uvt.x + uvt.y < 1.0 {
uvt
} else {
vec3(f32::MAX, f32::MAX, f32::MAX)
}
I used spirv-cross to look at the code rust-gpu was producing in glsl and noticed it was producing this:
if (_1039 > 0.0)
{
bool _1057;
bool _1058;
if (_1040 > 0.0)
{
bool _1047 = _1041 > 0.0;
bool _1053;
if (_1047)
{
_1053 = fma(_1028, _1023, _1040) < 1.0;
}
else
{
_1053 = _76;
}
_1057 = _1053;
_1058 = _1047 ? false : true;
}
else
{
_1057 = _76;
_1058 = true;
}
_1061 = _1057;
_1062 = _1058;
}
else
{
_1061 = _76;
_1062 = true;
}
Whereas if I take wgsl through the same path (wgsl -> spirv -> glsl) it looks like this:
if ((((_87.x > 0.0) && (_87.y > 0.0)) && (_87.z > 0.0)) && ((_87.x + _87.y) < 1.0)) {
return _87;
} else {
return vec3(F32MAX);
}
I tried forcing it to not branch but generate a bool, with u32(uvt.x > 0.0 && uvt.y > 0.0 && uvt.z > 0.0 && uvt.x + uvt.y < 1.0) == 1, and while it kept the conversion and the equality check, it still had this same nested branching structure.
I then tried this which got me a lot closer to the wgsl perf (now 58ms):
if (uvt.x > 0.0) as u32
& (uvt.y > 0.0) as u32
& (uvt.z > 0.0) as u32
& (uvt.x + uvt.y < 1.0) as u32
== 1
{
uvt
} else {
vec3(f32::MAX, f32::MAX, f32::MAX)
}
This actually results in it using mix here:
mix(vec3(F32MAX), vec3(_1024, _1025, _1026), bvec3((((uint(_1024 > 0.0) & uint(_1025 > 0.0)) & uint(_1026 > 0.0)) & uint(fma(_1013, _1008, _1025) < 1.0)) == 1u)).y;
Is it possible to improve the code generation in rust gpu to avoid the excessive branching in situations like this?
(I'm aware that this could also be written differently to avoid branching, I'm not concerned about this specific impl, but about the code generation in general)
Beitragsleitfaden
Erste Schritte
- Lies das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreib ins Issue, dass du es übernimmst — das erspart doppelte Arbeit.
- Forke das Repository und arbeite in einem Branch.
- Öffne einen Pull Request, der die Issue-Nummer nennt.
Rechercherichtung
Beginne damit, die gemeldete Rust-GPU-Shader-Bedingung mit dem generierten SPIR-V/GLSL und der entsprechenden WGSL-Konvertierung zu vergleichen. Untersuche, warum verkettete boolesche Bedingungen zu verschachtelten Verzweigungen werden; die Aufgabe ist abgeschlossen, wenn vergleichbare Bedingungen weniger übermäßige Verzweigungen erzeugen, ohne auf die manuelle Integer-Masken-Umschreibung des Issues angewiesen zu sein.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- rust
- Bereich
- compilers, performance
- Issue-Typ
- Bug
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Größtenteils klar
- Anfängerfreundlichkeit
- 35/100