AI-Hypercomputer / AI-Hypercomputer/gpu-recipes

Maxtext llama2-7b on 32 nodes has issue running on A3Mega

Đang mở
#2 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
Ngôn ngữ chính
Python
Star
140
Fork
87
Merge trung bình
7 ngày 21 giờ
Pull request đã merge (30 ngày)
1

Mô tả

Following the [instructions](https://github.com/AI-Hypercomputer/gpu-recipes/tree/main/training/a3mega/llama-2-7b/maxtext-pretraining-gke) in this repo for running 32-node maxtext llama2-7b workload.

I have the following error

```
Stopping coordination service as cluster registration failed. This may be due to 1) some tasks crashed earlier before connecting, 2) some tasks were never scheduled, or 3) scheduling delays. Consider setting a longer initialization timeout if such delays are expected, the timeout is currently set to: 1h.

Original error: DEADLINE_EXCEEDED: Barrier timed out. Id: [Init]Wait_for_all_tasks_to_register::0. This usually happens because a task triggered the barrier too early or too slowly. Please look at the task logs (both timed out and first task) to debug further.
```

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Đánh giá

Issue này chưa được đánh giá.

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.