[Migrated] Conditionals and code generation performance
Nadie ha tomado este issue todavía.
- Lenguaje dominante
- Rust
- Estrellas
- 3.4k
- Forks
- 126
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Descripción
Issue automatically imported from old repo: https://github.com/EmbarkStudios/rust-gpu/issues/1110
Old labels: t: bug
Originally creatd by DGriffin91 on 2023-12-18T19:06:07Z
I noticed that a rust gpu shader was running much slower than the equivalent wgsl one.
The wgsl one takes 53ms, and the rust gpu version takes 67ms.
Looking at the SPIRV I tracked part of the issue down to this:
if uvt.x > 0.0 && uvt.y > 0.0 && uvt.z > 0.0 && uvt.x + uvt.y < 1.0 {
uvt
} else {
vec3(f32::MAX, f32::MAX, f32::MAX)
}
I used spirv-cross to look at the code rust-gpu was producing in glsl and noticed it was producing this:
if (_1039 > 0.0)
{
bool _1057;
bool _1058;
if (_1040 > 0.0)
{
bool _1047 = _1041 > 0.0;
bool _1053;
if (_1047)
{
_1053 = fma(_1028, _1023, _1040) < 1.0;
}
else
{
_1053 = _76;
}
_1057 = _1053;
_1058 = _1047 ? false : true;
}
else
{
_1057 = _76;
_1058 = true;
}
_1061 = _1057;
_1062 = _1058;
}
else
{
_1061 = _76;
_1062 = true;
}
Whereas if I take wgsl through the same path (wgsl -> spirv -> glsl) it looks like this:
if ((((_87.x > 0.0) && (_87.y > 0.0)) && (_87.z > 0.0)) && ((_87.x + _87.y) < 1.0)) {
return _87;
} else {
return vec3(F32MAX);
}
I tried forcing it to not branch but generate a bool, with u32(uvt.x > 0.0 && uvt.y > 0.0 && uvt.z > 0.0 && uvt.x + uvt.y < 1.0) == 1, and while it kept the conversion and the equality check, it still had this same nested branching structure.
I then tried this which got me a lot closer to the wgsl perf (now 58ms):
if (uvt.x > 0.0) as u32
& (uvt.y > 0.0) as u32
& (uvt.z > 0.0) as u32
& (uvt.x + uvt.y < 1.0) as u32
== 1
{
uvt
} else {
vec3(f32::MAX, f32::MAX, f32::MAX)
}
This actually results in it using mix here:
mix(vec3(F32MAX), vec3(_1024, _1025, _1026), bvec3((((uint(_1024 > 0.0) & uint(_1025 > 0.0)) & uint(_1026 > 0.0)) & uint(fma(_1013, _1008, _1025) < 1.0)) == 1u)).y;
Is it possible to improve the code generation in rust gpu to avoid the excessive branching in situations like this?
(I'm aware that this could also be written differently to avoid branching, I'm not concerned about this specific impl, but about the code generation in general)
Guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Línea de trabajo
Comienza comparando la condición del shader de GPU en Rust reportada con su SPIR-V/GLSL generado y la conversión equivalente a WGSL. Investiga por qué las condiciones booleanas encadenadas se convierten en ramas anidadas; el trabajo estará terminado cuando las condiciones comparables generen menos ramificaciones excesivas sin depender de la reescritura manual con máscaras enteras del issue.
Escrito por el modelo de indexación a partir del texto del issue.
Evaluación
- Stack tecnológico
- rust
- Área
- compilers, performance
- Tipo de issue
- Error
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Estado de actividad
- Estancado
- Claridad
- Bastante claro
- Aptitud para principiantes
- 35/100