AI-Hypercomputer / AI-Hypercomputer/gpu-recipes

Maxtext llama2-7b on 32 nodes has issue running on A3Mega

未關閉
#2 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視
主要語言
Python
星號
140
分支
87
平均合併
7 天 21 小時
30 天內合併 PR
1

描述

Following the [instructions](https://github.com/AI-Hypercomputer/gpu-recipes/tree/main/training/a3mega/llama-2-7b/maxtext-pretraining-gke) in this repo for running 32-node maxtext llama2-7b workload.

I have the following error

```
Stopping coordination service as cluster registration failed. This may be due to 1) some tasks crashed earlier before connecting, 2) some tasks were never scheduled, or 3) scheduling delays. Consider setting a longer initialization timeout if such delays are expected, the timeout is currently set to: 1h.

Original error: DEADLINE_EXCEEDED: Barrier timed out. Id: [Init]Wait_for_all_tasks_to_register::0. This usually happens because a task triggered the barrier too early or too slowly. Please look at the task logs (both timed out and first task) to debug further.
```

貢獻指南

開啟貢獻指南

評估

這個 Issue 還沒有評估資料。

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。