[Migrated] Conditionals and code generation performance
まだ誰も着手していません。
- 主要言語
- Rust
- スター
- 3.4k
- フォーク
- 125
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
Issue automatically imported from old repo: https://github.com/EmbarkStudios/rust-gpu/issues/1110
Old labels: t: bug
Originally creatd by DGriffin91 on 2023-12-18T19:06:07Z
I noticed that a rust gpu shader was running much slower than the equivalent wgsl one.
The wgsl one takes 53ms, and the rust gpu version takes 67ms.
Looking at the SPIRV I tracked part of the issue down to this:
if uvt.x > 0.0 && uvt.y > 0.0 && uvt.z > 0.0 && uvt.x + uvt.y < 1.0 {
uvt
} else {
vec3(f32::MAX, f32::MAX, f32::MAX)
}
I used spirv-cross to look at the code rust-gpu was producing in glsl and noticed it was producing this:
if (_1039 > 0.0)
{
bool _1057;
bool _1058;
if (_1040 > 0.0)
{
bool _1047 = _1041 > 0.0;
bool _1053;
if (_1047)
{
_1053 = fma(_1028, _1023, _1040) < 1.0;
}
else
{
_1053 = _76;
}
_1057 = _1053;
_1058 = _1047 ? false : true;
}
else
{
_1057 = _76;
_1058 = true;
}
_1061 = _1057;
_1062 = _1058;
}
else
{
_1061 = _76;
_1062 = true;
}
Whereas if I take wgsl through the same path (wgsl -> spirv -> glsl) it looks like this:
if ((((_87.x > 0.0) && (_87.y > 0.0)) && (_87.z > 0.0)) && ((_87.x + _87.y) < 1.0)) {
return _87;
} else {
return vec3(F32MAX);
}
I tried forcing it to not branch but generate a bool, with u32(uvt.x > 0.0 && uvt.y > 0.0 && uvt.z > 0.0 && uvt.x + uvt.y < 1.0) == 1, and while it kept the conversion and the equality check, it still had this same nested branching structure.
I then tried this which got me a lot closer to the wgsl perf (now 58ms):
if (uvt.x > 0.0) as u32
& (uvt.y > 0.0) as u32
& (uvt.z > 0.0) as u32
& (uvt.x + uvt.y < 1.0) as u32
== 1
{
uvt
} else {
vec3(f32::MAX, f32::MAX, f32::MAX)
}
This actually results in it using mix here:
mix(vec3(F32MAX), vec3(_1024, _1025, _1026), bvec3((((uint(_1024 > 0.0) & uint(_1025 > 0.0)) & uint(_1026 > 0.0)) & uint(fma(_1013, _1008, _1025) < 1.0)) == 1u)).y;
Is it possible to improve the code generation in rust gpu to avoid the excessive branching in situations like this?
(I'm aware that this could also be written differently to avoid branching, I'm not concerned about this specific impl, but about the code generation in general)
コントリビューションガイド
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
調査の方向性
まず、報告された Rust GPU シェーダーの条件式を、生成された SPIR-V/GLSL および同等の WGSL 変換と比較します。連結されたブール条件がなぜネストした分岐になるのかを調査します。完了の条件は、Issue の手動による整数マスクへの書き換えに頼らず、同等の条件で過剰な分岐が少なくなることです。
索引モデルが issue の本文から書いたものです。
評価
- 技術スタック
- rust
- 領域
- compilers, performance
- issue の種類
- バグ
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 活発さ
- 停滞
- 明瞭さ
- おおむね明確
- 初心者へのやさしさ
- 35/100