AI-Hypercomputer / AI-Hypercomputer/gpu-recipes

Maxtext llama2-7b on 32 nodes has issue running on A3Mega

Aberta
#2 0 comentários 0 reações 0 responsáveis Ver no GitHub
Linguagem predominante
Python
Estrelas
140
Forks
87
Merge médio
7d 21h
PRs com merge (30d)
1

Descrição

Following the [instructions](https://github.com/AI-Hypercomputer/gpu-recipes/tree/main/training/a3mega/llama-2-7b/maxtext-pretraining-gke) in this repo for running 32-node maxtext llama2-7b workload.

I have the following error

```
Stopping coordination service as cluster registration failed. This may be due to 1) some tasks crashed earlier before connecting, 2) some tasks were never scheduled, or 3) scheduling delays. Consider setting a longer initialization timeout if such delays are expected, the timeout is currently set to: 1h.

Original error: DEADLINE_EXCEEDED: Barrier timed out. Id: [Init]Wait_for_all_tasks_to_register::0. This usually happens because a task triggered the barrier too early or too slowly. Please look at the task logs (both timed out and first task) to debug further.
```

Guia de contribuição

Abrir o guia de contribuição

Avaliação

Esta issue ainda não foi avaliada.

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.