[QUESTION] Why does GPTDataset not directly cache all samples document_index and sample_index, and then construct different shuffle_index for different parameters?
Open
community-request
enhancement
module: data pipeline
- Dominant language
- Python
- Stars
- 17.9k
- Forks
- 4.5k
- Avg merge
- 4d 6h
- Merged PRs (30d)
- 271
Description
**Your question**
Why does GPTDataset not directly cache all samples document_index and sample_index, and then construct different shuffle_index for different parameters?
Contributor guide
Assessment
This issue has not been assessed yet.