Inquiry and Suggestions Regarding OpenBLAS Code Flow with OpenMP

Abierto
#4,418 8 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
20/100
Tipo de issue
Error
Claridad
Necesita aclaración
Estado de actividad
Estancado
Stack tecnológico
c
Área
hpc, performance

Línea de trabajo

Comienza siguiendo gemm_driver en driver/level3/level3_thread.c y exec_blas e inner_thread en driver/others/blas_server_omp.c. Reproduce el deadlock reportado con menos hilos de OpenMP que particiones de la cola y, después, compara las rutas de bloqueo y espera activa. La tarea se considerará completada cuando haya una explicación confirmada y un cambio de sincronización claramente delimitado o un resultado de documentación.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Hello,

I've been delving into the OpenBLAS codebase, specifically focusing on the gemm_driver function in level3_thread.c with the #USE_OPENMP=1 flag enabled. I've come across a section in the code where a lock is used in the Pthreads and Win32 backend before initializing and assigning the blasqueue. The lock is released only after the exec_blas call is completed.

My current understanding is :

  1. In the case of Pthreads The thread pool is initialized (POSIX), and when a BLAS call is made, the thread pool is utilized to execute it, after which the threads go back to sleep. The use of locks ensures that only one exec_blas call can be executed at a time.
  1. If I use a thread pool with No_of_threads < nthreads (number of queue partitions) defined in level3_thread.c, I encounter a deadlock within the inner_threads function.
  2. There seems to be significant busy-waiting (NOP) inside the inner_thread function call for thread synchronization.

I have a few questions and want to seek suggestions from the community:

  1. I noticed that OpenMP locking mechanisms like omp_set_lock are not used, and instead, busy-waiting is implemented inside the exec_blas function in blas_server_omp.c using max_parallel_number. Could you please provide insights into this choice?

  2. The inner_thread function appears to encounter deadlock when used with fewer threads. I tested this in OpenMP by creating a parallel region inside exec_blas in line 437 - blas_server_omp.c with a fixed number of threads less than nthreads. Could you shed some light on the reasons behind this behavior?

  3. Considering the busy-waiting in the code, have there been considerations for putting threads to sleep or employing other synchronization methods to enhance efficiency?

I would appreciate clarification on these points.

Lenguaje dominante
C
Estrellas
7.6k
Forks
1.7k
Merge medio
1 d 3 h
PR fusionados (30 d)
42

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de OpenMathLib/OpenBLAS

Todos los issues de OpenMathLib/OpenBLAS

Issues similares

Más issues de C

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.