agentscope-ai / agentscope-ai/Trinity-RFT
[BUG] NCCL hang during rollout_weight_sync (ProcessGroup watchdog timeout)
- Ngôn ngữ chính
- Python
- Star
- 701
- Fork
- 79
- Merge trung bình
- 8 giờ 7 phút
- Pull request đã merge (30 ngày)
- 1
Mô tả
### Bug Description
Please provide a detailed description of the issue you encountered.
### Environment Information
- Python Version: 3.12.4
- GPU: NVIDIA L20-40G * 8
- CUDA Version: 12.4
- Installation Method: git clone
- Trinity-RFT Version: 0.3.0.dev0
### Steps to Reproduce
Please provide a minimal, self-contained, and reproducible example.
1. trinity run --config examples/XXX/XXX.yaml
### Expected Behavior
No interruptions.
### Actual Behavior
During multi-GPU training, the process occasionally crashes with a NCCL watchdog hang. The error happens at rollout_weight_sync and terminates the job with SIGABRT.
### Log Information
```vbnet
ProcessGroupNCCL.cpp:1554 [PG ID 6 PG GUID rollout_weight_sync Rank 3]
ProcessGroup watchdog hang due to timeout
*** SIGABRT received ...
Fatal Python error: Aborted
```
### Question
Is this a known NCCL communication hang issue?
Any recommended configuration or workaround to prevent the crash?
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.