bigscience-workshop / bigscience-workshop/tokenization
Is there anyway to merge multiple tokenize vocab?
Open
- Dominant language
- Python
- Stars
- 11
- Forks
- 2
- PR merge metrics
- No merged PRs in 30d
Description
Hi guys, My PC RAM Is overused because the large files to tokenize, So I have to train the tokenizers for part to part, but is there any way to merge them? Thanks a lot!
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names no files, tests, tokenizer implementation, or vocabulary format. First inspect the repository to identify the tokenizer entry points and supported vocabulary artifacts, then clarify which independently trained vocabularies should be merged and what successful merging must preserve.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100