riseproject-dev / riseproject-dev/python-wheels

libgomp dynamic and guided OpenMP schedules segfault on the riscv64 runners

オープン
#617 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

主要言語
Python
スター
0
フォーク
0
平均マージ
16時間 14分
マージ済み PR(30日)
952

説明

Summary

On the ubuntu-24.04-riscv runners, libgomp's dynamic and guided work-share schedules fault. A 40-line C program that allocates nothing in the loop body segfaults 3 times out of 3; schedule(static) on the same program is clean. This is independent of LightGBM, of the compiler version and of the container.

Any riscv64 wheel we publish whose extension uses #pragma omp parallel for schedule(dynamic) or schedule(guided) is exposed at runtime.

Reproducer
#include <omp.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>

static long calc_body(int i) { return (long)i * 3 % 7; }

int main(int argc, char **argv) {
  const char *kind = argc > 1 ? argv[1] : "guided";
  int rounds = argc > 2 ? atoi(argv[2]) : 100000;
  int n = argc > 3 ? atoi(argv[3]) : 300;
  int nthreads = omp_get_max_threads();
  long total = 0;
  for (int r = 0; r < rounds; r++) {
    long s = 0;
    if (!strcmp(kind, "guided")) {
#pragma omp parallel for num_threads(nthreads) schedule(guided) reduction(+ : s)
      for (int i = 0; i < n; i++) s += calc_body(i);
    } else if (!strcmp(kind, "dynamic")) {
#pragma omp parallel for num_threads(nthreads) schedule(dynamic) reduction(+ : s)
      for (int i = 0; i < n; i++) s += calc_body(i);
    } else {
#pragma omp parallel for num_threads(nthreads) schedule(static) reduction(+ : s)
      for (int i = 0; i < n; i++) s += calc_body(i);
    }
    total += s;
  }
  printf("OK %s %ld\n", kind, total);
  return 0;
}

gcc -O2 -fopenmp -o probe probe.c && ./probe guided 100000 300

Results

Runner riscv-runner-48, 4 harts, isa: rv64imafdcsu, mmu: sv39, 16 GB.

schedule body bare runner (GCC 13.3.0) manylinux_2_39_riscv64 (GCC 14.3.1)
guided arithmetic only SIGSEGV 3/3 SIGSEGV 3/3
dynamic malloc/free SIGSEGV 3/3 SIGSEGV 3/3
guided malloc/free SIGSEGV 3/3 SIGSEGV 3/3
static malloc/free 0/3 0/3
4 raw pthreads, 30M malloc/free each 0/3 0/3

Controls: the identical binaries are clean under QEMU riscv64 (6M guided regions) and on manylinux_2_39_aarch64 at 4, 16 and 64 threads.

So the fault is specific to libgomp's shared work-share iterator — the path schedule(static) does not use — on this hardware. glibc's allocator is not implicated.

How it surfaces

lightgbm (#616) is the package that found it. Its wheel builds and repairs cleanly, then the ranking tests fault inside gomp_iter_guided_next():

Thread 9 "python" received signal SIGSEGV
#0  gomp_iter_guided_next () from /lib64/lp64d/libgomp.so.1
#1  LightGBM::FeatureGroup::FinishLoad()._omp_fn.0 () at include/LightGBM/feature_group.h:366
#2  gomp_thread_start () from /lib64/lp64d/libgomp.so.1

Setting OMP_WAIT_POLICY=passive GOMP_SPINCOUNT=0, or OMP_NUM_THREADS=1, makes it disappear (0/10 each, against 9/10 at the default). Rebuilding LightGBM at -O1 does not (10/10), so it is not a codegen problem in the package.

Next steps

Worth narrowing to hardware vs kernel vs libgomp before reporting upstream: the same probe on a different riscv64 machine would say which.

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

調査の方向性

まず、提供された C 再現プログラムを ubuntu-24.04-riscv runner 上で gcc -O2 -fopenmp を使ってコンパイルします。guided および dynamic スケジュールを static および列挙された制御項目と比較します。include/LightGBM/feature_group.h:366 の LightGBM trace を使って、障害が riscv64 マシン、カーネル、または libgomp のバージョンによって異なるかどうかを絞り込み、確認済みの原因または upstream report を記録します。

索引モデルが issue の本文から書いたものです。

評価

技術スタック
c, linux, python, ubuntu
領域
compilers, infrastructure, operating-systems
issue の種類
バグ
難易度
4/5
見積もり時間
3〜5日
活発さ
活発
明瞭さ
おおむね明確
初心者へのやさしさ
48/100

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。