riseproject-dev / riseproject-dev/python-wheels

libgomp dynamic and guided OpenMP schedules segfault on the riscv64 runners

未关闭
#617 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

主要语言
Python
星标
0
派生
0
平均合并
16 小时 14 分钟
30 天内合并 PR
952

描述

Summary

On the ubuntu-24.04-riscv runners, libgomp's dynamic and guided work-share schedules fault. A 40-line C program that allocates nothing in the loop body segfaults 3 times out of 3; schedule(static) on the same program is clean. This is independent of LightGBM, of the compiler version and of the container.

Any riscv64 wheel we publish whose extension uses #pragma omp parallel for schedule(dynamic) or schedule(guided) is exposed at runtime.

Reproducer
#include <omp.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>

static long calc_body(int i) { return (long)i * 3 % 7; }

int main(int argc, char **argv) {
  const char *kind = argc > 1 ? argv[1] : "guided";
  int rounds = argc > 2 ? atoi(argv[2]) : 100000;
  int n = argc > 3 ? atoi(argv[3]) : 300;
  int nthreads = omp_get_max_threads();
  long total = 0;
  for (int r = 0; r < rounds; r++) {
    long s = 0;
    if (!strcmp(kind, "guided")) {
#pragma omp parallel for num_threads(nthreads) schedule(guided) reduction(+ : s)
      for (int i = 0; i < n; i++) s += calc_body(i);
    } else if (!strcmp(kind, "dynamic")) {
#pragma omp parallel for num_threads(nthreads) schedule(dynamic) reduction(+ : s)
      for (int i = 0; i < n; i++) s += calc_body(i);
    } else {
#pragma omp parallel for num_threads(nthreads) schedule(static) reduction(+ : s)
      for (int i = 0; i < n; i++) s += calc_body(i);
    }
    total += s;
  }
  printf("OK %s %ld\n", kind, total);
  return 0;
}

gcc -O2 -fopenmp -o probe probe.c && ./probe guided 100000 300

Results

Runner riscv-runner-48, 4 harts, isa: rv64imafdcsu, mmu: sv39, 16 GB.

schedule body bare runner (GCC 13.3.0) manylinux_2_39_riscv64 (GCC 14.3.1)
guided arithmetic only SIGSEGV 3/3 SIGSEGV 3/3
dynamic malloc/free SIGSEGV 3/3 SIGSEGV 3/3
guided malloc/free SIGSEGV 3/3 SIGSEGV 3/3
static malloc/free 0/3 0/3
4 raw pthreads, 30M malloc/free each 0/3 0/3

Controls: the identical binaries are clean under QEMU riscv64 (6M guided regions) and on manylinux_2_39_aarch64 at 4, 16 and 64 threads.

So the fault is specific to libgomp's shared work-share iterator — the path schedule(static) does not use — on this hardware. glibc's allocator is not implicated.

How it surfaces

lightgbm (#616) is the package that found it. Its wheel builds and repairs cleanly, then the ranking tests fault inside gomp_iter_guided_next():

Thread 9 "python" received signal SIGSEGV
#0  gomp_iter_guided_next () from /lib64/lp64d/libgomp.so.1
#1  LightGBM::FeatureGroup::FinishLoad()._omp_fn.0 () at include/LightGBM/feature_group.h:366
#2  gomp_thread_start () from /lib64/lp64d/libgomp.so.1

Setting OMP_WAIT_POLICY=passive GOMP_SPINCOUNT=0, or OMP_NUM_THREADS=1, makes it disappear (0/10 each, against 9/10 at the default). Rebuilding LightGBM at -O1 does not (10/10), so it is not a codegen problem in the package.

Next steps

Worth narrowing to hardware vs kernel vs libgomp before reporting upstream: the same probe on a different riscv64 machine would say which.

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

调研方向

首先,在 ubuntu-24.04-riscv runner 上使用 gcc -O2 -fopenmp 编译提供的 C 复现程序;将 guided 和 dynamic 调度与 static 以及列出的控制项进行比较。利用 include/LightGBM/feature_group.h:366 处的 LightGBM trace,确定该故障是否因不同的 riscv64 机器、内核或 libgomp 版本而有所不同;记录已确认的原因或 upstream 报告。

由索引模型根据 Issue 内容生成。

评估

技术栈
c, linux, python, ubuntu
领域
compilers, infrastructure, operating-systems
Issue 类型
缺陷
难度
4/5
预计耗时
3-5 天
活跃度
活跃
描述清晰度
基本清楚
新手友好度
48/100

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。