google-deepmind / google-deepmind/asyncdiloco

Why is Async Local SGD better even without momentum

Open
#1 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
51
Forks
3
PR merge metrics
No merged PRs in 30d

Description

I wondered why, in the experiment where momentum was turned off, the async local sgd still outperformed vanilla local-sgd. Shouldn't the gradient staleness significantly impact wall clock training time, which is why async-local sgd suggests a delayed momentum update?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.