Rust-GPU / Rust-GPU/rust-gpu

[Migrated] Conditionals and code generation performance

Abierto
#71 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Lenguaje dominante
Rust
Estrellas
3.4k
Forks
126
Métricas de merge de PR
Sin PR fusionados en 30 d

Descripción

Issue automatically imported from old repo: https://github.com/EmbarkStudios/rust-gpu/issues/1110
Old labels: t: bug
Originally creatd by DGriffin91 on 2023-12-18T19:06:07Z


I noticed that a rust gpu shader was running much slower than the equivalent wgsl one.
The wgsl one takes 53ms, and the rust gpu version takes 67ms.
Looking at the SPIRV I tracked part of the issue down to this:

if uvt.x > 0.0 && uvt.y > 0.0 && uvt.z > 0.0 && uvt.x + uvt.y < 1.0 {
    uvt
} else {
    vec3(f32::MAX, f32::MAX, f32::MAX)
}

I used spirv-cross to look at the code rust-gpu was producing in glsl and noticed it was producing this:

if (_1039 > 0.0)
{
    bool _1057;
    bool _1058;
    if (_1040 > 0.0)
    {
        bool _1047 = _1041 > 0.0;
        bool _1053;
        if (_1047)
        {
            _1053 = fma(_1028, _1023, _1040) < 1.0;
        }
        else
        {
            _1053 = _76;
        }
        _1057 = _1053;
        _1058 = _1047 ? false : true;
    }
    else
    {
        _1057 = _76;
        _1058 = true;
    }
    _1061 = _1057;
    _1062 = _1058;
}
else
{
    _1061 = _76;
    _1062 = true;
}

Whereas if I take wgsl through the same path (wgsl -> spirv -> glsl) it looks like this:

if ((((_87.x > 0.0) && (_87.y > 0.0)) && (_87.z > 0.0)) && ((_87.x + _87.y) < 1.0)) {
    return _87;
} else {
    return vec3(F32MAX);
}

I tried forcing it to not branch but generate a bool, with u32(uvt.x > 0.0 && uvt.y > 0.0 && uvt.z > 0.0 && uvt.x + uvt.y < 1.0) == 1, and while it kept the conversion and the equality check, it still had this same nested branching structure.

I then tried this which got me a lot closer to the wgsl perf (now 58ms):

if (uvt.x > 0.0) as u32
    & (uvt.y > 0.0) as u32
    & (uvt.z > 0.0) as u32
    & (uvt.x + uvt.y < 1.0) as u32
    == 1
{
    uvt
} else {
    vec3(f32::MAX, f32::MAX, f32::MAX)
}

This actually results in it using mix here:

mix(vec3(F32MAX), vec3(_1024, _1025, _1026), bvec3((((uint(_1024 > 0.0) & uint(_1025 > 0.0)) & uint(_1026 > 0.0)) & uint(fma(_1013, _1008, _1025) < 1.0)) == 1u)).y;

Is it possible to improve the code generation in rust gpu to avoid the excessive branching in situations like this?

(I'm aware that this could also be written differently to avoid branching, I'm not concerned about this specific impl, but about the code generation in general)

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Línea de trabajo

Comienza comparando la condición del shader de GPU en Rust reportada con su SPIR-V/GLSL generado y la conversión equivalente a WGSL. Investiga por qué las condiciones booleanas encadenadas se convierten en ramas anidadas; el trabajo estará terminado cuando las condiciones comparables generen menos ramificaciones excesivas sin depender de la reescritura manual con máscaras enteras del issue.

Escrito por el modelo de indexación a partir del texto del issue.

Evaluación

Stack tecnológico
rust
Área
compilers, performance
Tipo de issue
Error
Dificultad
5/5
Tiempo estimado
Más de una semana
Estado de actividad
Estancado
Claridad
Bastante claro
Aptitud para principiantes
35/100

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.