AI-Hypercomputer / AI-Hypercomputer/maxtext
Conversion scripts fail on standard TPU VMs for large models
- Ngôn ngữ chính
- Python
- Star
- 2.4k
- Fork
- 607
- Merge trung bình
- 2 ngày 19 giờ
- Pull request đã merge (30 ngày)
- 158
Mô tả
### Bug report
The [official instruction]([https://maxtext.readthedocs.io/en/latest/guides/checkpointing_solutions/convert_checkpoint.html#hugging-face-to-maxtext) for model conversion failed on a standard TPU-v5p VM for large models such as QWen3-235B, with CPU OOM errors when sharding MoE layers. Since on GCP the CPU memory is fixed (400GB), I wonder if we can improve the script to bypass this issue, or is the doc outdated?
Also for the model conversion the sharding process with default simulated_cpu_devices_count=16 is very very slow (even for 30B model).
### Logs/Output
_No response_
### Environment Information
_No response_
### Additional Context
_No response_
Hướng dẫn đóng góp
Đánh giá
Issue này chưa được đánh giá.