AI-Hypercomputer / AI-Hypercomputer/maxtext
Regression: HF JSONL training halts early with
- 主要語言
- Python
- 星號
- 2.4k
- 分支
- 607
- 平均合併
- 2 天 19 小時
- 30 天內合併 PR
- 158
描述
### Bug report
@aireenmei
I keep getting an error related to running out of shards. Training is halted early, and gives the error:
“Run out of shards on host 0, shard 256 is not available”
followed by StopIteration → “You may have run out of training data.”
Setup:
- MaxText @ main (Aug 2025)
- TPU v5-32 (8 hosts)
- dataset_type=hf, streaming JSONL from GCS
- 256 shards matched by: hf_train_files='gs://mybucket/train*.jsonl'
(files named train_aa.jsonl … train_jo.jsonl)
- Command: `python -m MaxText.train ... hf_path=json hf_data_dir= hf_train_files='gs://…/train*.jsonl' ...`
I think this is a regression, because earlier this did run until none of the workers had any data.
### Logs/Output
_No response_
### Environment Information
_No response_
### Additional Context
_No response_
貢獻指南
評估
這個 Issue 還沒有評估資料。