huggingface / huggingface/alignment-handbook
Can we please add the option to work with a tokenized dataset, escpailly for the CPT task.
Open
- Dominant language
- Python
- Stars
- 5.7k
- Forks
- 490
- Avg merge
- 2m
- Merged PRs (30d)
- 1
Description
Since we have the CPT task now, it would be nice to have the ability to feel a tokenized and packed dataset directly.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.