bigscience-workshop / bigscience-workshop/tokenization
Is there anyway to merge multiple tokenize vocab?
Offen
- Vorherrschende Sprache
- Python
- Sterne
- 11
- Forks
- 2
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Beschreibung
Hi guys, My PC RAM Is overused because the large files to tokenize, So I have to train the tokenizers for part to part, but is there any way to merge them? Thanks a lot!
Beitragsleitfaden
Für dieses Repository ist kein Beitragsleitfaden indexiert
Rechercherichtung
The issue names no files, tests, tokenizer implementation, or vocabulary format. First inspect the repository to identify the tokenizer entry points and supported vocabulary artifacts, then clarify which independently trained vocabularies should be merged and what successful merging must preserve.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Bewertung
- Tech-Stack
- python
- Bereich
- machine-learning
- Issue-Typ
- Feature
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Aktivitätsstatus
- Veraltet
- Klarheit
- Muss geklärt werden
- Anfängerfreundlichkeit
- 25/100