FedML-AI / FedML-AI/FedCV

Problems of distributed computing in federated learning

Open
#36 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
No language data
Stars
66
Forks
23
PR merge metrics
No merged PRs in 30d

Description

When using distributed operation, I have four Gpus, each of which has a client. During the training process, each GPU has a huge difference. Two gpus even ran out of memory. By the way, I also found that gpu training with overflow was extremely slow and seemed to have gpu utilization close to zero.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.