agentscope-ai / agentscope-ai/Trinity-RFT

[BUG] NCCL hang during rollout_weight_sync (ProcessGroup watchdog timeout)

未关闭
#283 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看
bug
主要语言
Python
星标
701
派生
79
平均合并
8 小时 7 分钟
30 天内合并 PR
1

描述

### Bug Description

Please provide a detailed description of the issue you encountered.

### Environment Information

- Python Version: 3.12.4
- GPU: NVIDIA L20-40G * 8
- CUDA Version: 12.4
- Installation Method: git clone
- Trinity-RFT Version: 0.3.0.dev0

### Steps to Reproduce

Please provide a minimal, self-contained, and reproducible example.

1. trinity run --config examples/XXX/XXX.yaml

### Expected Behavior

No interruptions.

### Actual Behavior

During multi-GPU training, the process occasionally crashes with a NCCL watchdog hang. The error happens at rollout_weight_sync and terminates the job with SIGABRT.

### Log Information

```vbnet
ProcessGroupNCCL.cpp:1554 [PG ID 6 PG GUID rollout_weight_sync Rank 3]
ProcessGroup watchdog hang due to timeout
*** SIGABRT received ...
Fatal Python error: Aborted
```

Image

### Question
Is this a known NCCL communication hang issue?
Any recommended configuration or workaround to prevent the crash?

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。