AI-Hypercomputer / AI-Hypercomputer/maxtext

Regression: HF JSONL training halts early with

未关闭
#2,147 4 条评论 0 个 reaction 已指派 1 人 已被 @aireenmei 认领 在 GitHub 查看
bug
主要语言
Python
星标
2.4k
派生
607
平均合并
2 天 19 小时
30 天内合并 PR
158

描述

### Bug report

@aireenmei

I keep getting an error related to running out of shards. Training is halted early, and gives the error:
“Run out of shards on host 0, shard 256 is not available”
followed by StopIteration → “You may have run out of training data.”

Setup:
- MaxText @ main (Aug 2025)
- TPU v5-32 (8 hosts)
- dataset_type=hf, streaming JSONL from GCS
- 256 shards matched by: hf_train_files='gs://mybucket/train*.jsonl'
(files named train_aa.jsonl … train_jo.jsonl)
- Command: `python -m MaxText.train ... hf_path=json hf_data_dir= hf_train_files='gs://…/train*.jsonl' ...`

I think this is a regression, because earlier this did run until none of the workers had any data.

### Logs/Output

_No response_

### Environment Information

_No response_

### Additional Context

_No response_

贡献指南

打开贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。