allenai / allenai/OLMo

Scaling of multinode training

Aperta
#845 2 commenti 0 reazioni 0 assegnatari Vedi su GitHub
type/question
Lingua principale
Python
Stelle
6.7k
Fork
797
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

### ❓ The question

Hi

I am trying to optimise the scaling behaviour of multi-node training with OLMo, and I would like to know what is expected. Running an a slurm cluster, with A100 40G, I get ~12000 tokens/sec/device with 1 node (4 GPUs) but only ~9000 toks/sec/device when running with 2 nodes.

I have tested various settings, and discussed with the cluster admins, but I haven't managed to improve the scaling. This is with FSDP.

So my question is, is this the expected scaling performance of OLMo?

best
Barry

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.