bigscience-workshop / bigscience-workshop/tokenization

Is there anyway to merge multiple tokenize vocab?

Aperta
#3 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub
Lingua principale
Python
Stelle
11
Fork
2
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Descrizione

Hi guys, My PC RAM Is overused because the large files to tokenize, So I have to train the tokenizers for part to part, but is there any way to merge them? Thanks a lot!

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Direzione di ricerca

The issue names no files, tests, tokenizer implementation, or vocabulary format. First inspect the repository to identify the tokenizer entry points and supported vocabulary artifacts, then clarify which independently trained vocabularies should be merged and what successful merging must preserve.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Valutazione

Stack tecnologico
python
Ambito
machine-learning
Tipo di issue
Funzionalità
Difficoltà
5/5
Tempo stimato
Più di una settimana
Stato di attività
Ferma
Chiarezza
Da chiarire
Idoneità per principianti
25/100

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.