NVIDIA / NVIDIA/Megatron-LM

[QUESTION] Why does GPTDataset not directly cache all samples document_index and sample_index, and then construct different shuffle_index for different parameters?

Open
#1,139 3 comments 0 reactions 0 assignees View on GitHub
community-request enhancement module: data pipeline
Dominant language
Python
Stars
17.9k
Forks
4.5k
Avg merge
4d 6h
Merged PRs (30d)
271

Description

**Your question**
Why does GPTDataset not directly cache all samples document_index and sample_index, and then construct different shuffle_index for different parameters?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.