huggingface / huggingface/accelerate

Need to do distributed training using 2 separate machines

Open
#924 2 comments 1 reaction 1 assignee Claimed by @muellerzr View on GitHub
enhancement feature request
Dominant language
Python
Stars
9.9k
Forks
1.5k
Avg merge
5d 2h
Merged PRs (30d)
27

Description

Hello, I am trying to do distributed training using 2 separate machines. Can anyone please guide me towards any tutorial / demo on this? The configs created using accelerate are:

_Machine 1_:

compute_environment: LOCAL_MACHINE
deepspeed_config: {}
distributed_type: MULTI_CPU
fsdp_config: {}
machine_rank: 0
main_process_ip: 20.160.27.77
main_process_port: 8080
main_training_function: main
mixed_precision: fp16
num_machines: 2
num_processes: 2
use_cpu: true

_Machine 2:_

compute_environment: LOCAL_MACHINE
deepspeed_config: {}
distributed_type: MULTI_CPU
fsdp_config: {}
machine_rank: 1
main_process_ip: 20.160.27.77
main_process_port: 8080
main_training_function: main
mixed_precision: fp16
num_machines: 2
num_processes: 2
use_cpu: true

Then I ran both the launch commands , they started training separately. Then I ran only the main machine , which again started training on it's own. I am not able to get any concrete direction on this.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.