Megvii-BaseDetection / Megvii-BaseDetection/YOLOX

In Windows, multiple Gpus train my VOC datasets to report NCCL problems

Open
#1,394 3 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
10.6k
Forks
2.5k
PR merge metrics
No merged PRs in 30d

Description

Hello, the model can be trained normally when I use 1 GPU. But the following problem occurred when I tried to convert the model to 4 Gpus for training

image-20220622161317189

(yolox_train) D:\mengxianchi\YOLOX\YOLOX>python tools/train.py -f exps/example/yolox_voc/yolox_voc_s.py -d 4 -b 32 --fp1
6 -o -c weights/yolox_s.pth
2022-06-22 14:58:49.750 | INFO | yolox.core.launch:_distributed_worker:116 - Rank 1 initialization finished.
2022-06-22 14:58:49.844 | INFO | yolox.core.launch:_distributed_worker:116 - Rank 3 initialization finished.
2022-06-22 14:58:50.089 | INFO | yolox.core.launch:_distributed_worker:116 - Rank 0 initialization finished.
2022-06-22 14:58:50.170 | INFO | yolox.core.launch:_distributed_worker:116 - Rank 2 initialization finished.
2022-06-22 14:58:50.183 | ERROR | yolox.core.launch:distributed_worker:126 - Process group URL: tcp://127.0.0.1:4931
5
Traceback (most recent call last):
File "tools/train.py", line 133, in
launch(
File "d:\mengxianchi\yolox\yolox\yolox\core\launch.py", line 82, in launch
mp.start_processes(
File "C:\ProgramData\Anaconda3\envs\yolox_train\lib\site-packages\torch\multiprocessing\spawn.py", line 188, in start

processes
while not context.join():
File "C:\ProgramData\Anaconda3\envs\yolox_train\lib\site-packages\torch\multiprocessing\spawn.py", line 150, in join
raise ProcessRaisedException(msg, error_index, failed_process.pid)
torch.multiprocessing.spawn.ProcessRaisedException:

-- Process 2 terminated with the following error:
Traceback (most recent call last):
File "C:\ProgramData\Anaconda3\envs\yolox_train\lib\site-packages\torch\multiprocessing\spawn.py", line 59, in _wrap
fn(i, *args)
File "d:\mengxianchi\yolox\yolox\yolox\core\launch.py", line 118, in _distributed_worker
dist.init_process_group(
File "C:\ProgramData\Anaconda3\envs\yolox_train\lib\site-packages\torch\distributed\distributed_c10d.py", line 503, in
init_process_group
_update_default_pg(_new_process_group_helper(
File "C:\ProgramData\Anaconda3\envs\yolox_train\lib\site-packages\torch\distributed\distributed_c10d.py", line 597, in
_new_process_group_helper
raise RuntimeError("Distributed package doesn't have NCCL "
RuntimeError: Distributed package doesn't have NCCL built in

There are many ways to try to solve it online, and want to ask you.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the reported command from tools/train.py on Windows, then inspect yolox/core/launch.py around _distributed_worker and dist.init_process_group. Determine whether the failure is supported by the current Windows setup and distributed backend configuration. Done means the issue has a confirmed cause and a documented or implemented resolution, but the report does not specify which is expected.

Written by the indexing model from the issue text.

Assessment

Tech stack
python, pytorch
Domain
distributed-systems, machine-learning, operating-systems
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.