AI-Hypercomputer / AI-Hypercomputer/gpu-recipes

Maxtext llama2-7b on 32 nodes has issue running on A3Mega

Aperta
#2 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
140
Fork
87
Merge medio
7g 21h
PR unite (30g)
1

Descrizione

Following the [instructions](https://github.com/AI-Hypercomputer/gpu-recipes/tree/main/training/a3mega/llama-2-7b/maxtext-pretraining-gke) in this repo for running 32-node maxtext llama2-7b workload.

I have the following error

```
Stopping coordination service as cluster registration failed. This may be due to 1) some tasks crashed earlier before connecting, 2) some tasks were never scheduled, or 3) scheduling delays. Consider setting a longer initialization timeout if such delays are expected, the timeout is currently set to: 1h.

Original error: DEADLINE_EXCEEDED: Barrier timed out. Id: [Init]Wait_for_all_tasks_to_register::0. This usually happens because a task triggered the barrier too early or too slowly. Please look at the task logs (both timed out and first task) to debug further.
```

Guida per i contributori

Apri la guida per i contributori

Valutazione

Questa issue non è ancora stata valutata.

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.