bigscience-workshop / bigscience-workshop/tokenization
Is there anyway to merge multiple tokenize vocab?
Open
- Dominant language
- Python
- Stars
- 11
- Forks
- 2
- PR merge metrics
- No merged PRs in 30d
Description
Hi guys, My PC RAM Is overused because the large files to tokenize, So I have to train the tokenizers for part to part, but is there any way to merge them? Thanks a lot!
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.