Scaling of multinode training
Aperta
type/question
- Lingua principale
- Python
- Stelle
- 6.7k
- Fork
- 797
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
### ❓ The question
Hi
I am trying to optimise the scaling behaviour of multi-node training with OLMo, and I would like to know what is expected. Running an a slurm cluster, with A100 40G, I get ~12000 tokens/sec/device with 1 node (4 GPUs) but only ~9000 toks/sec/device when running with 2 nodes.
I have tested various settings, and discussed with the cluster admins, but I haven't managed to improve the scaling. This is with FSDP.
So my question is, is this the expected scaling performance of OLMo?
best
Barry
Guida per i contributori
Apri la guida per i contributori
Valutazione
Questa issue non è ancora stata valutata.