bigscience-workshop / bigscience-workshop/tokenization
Is there anyway to merge multiple tokenize vocab?
Aperta
- Lingua principale
- Python
- Stelle
- 11
- Fork
- 2
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Descrizione
Hi guys, My PC RAM Is overused because the large files to tokenize, So I have to train the tokenizers for part to part, but is there any way to merge them? Thanks a lot!
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Direzione di ricerca
The issue names no files, tests, tokenizer implementation, or vocabulary format. First inspect the repository to identify the tokenizer entry points and supported vocabulary artifacts, then clarify which independently trained vocabularies should be merged and what successful merging must preserve.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Valutazione
- Stack tecnologico
- python
- Ambito
- machine-learning
- Tipo di issue
- Funzionalità
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Stato di attività
- Ferma
- Chiarezza
- Da chiarire
- Idoneità per principianti
- 25/100