huggingface / huggingface/transformers

How to tokenize big dataset

Open
#13,844 6 comments 0 reactions 0 assignees View on GitHub
Feature request
Dominant language
Python
Stars
166k
Forks
34.6k
Avg merge
3d 9h
Merged PRs (30d)
281

Description

Based on examples, I am trying to train a tokenizer and a model for T5.
I use Google Colab pro,
when I tried to run the following code:

```
import datasets

from t5_tokenizer_model import SentencePieceUnigramTokenizer

vocab_size = 32_000
input_sentence_size = None # change to 100_000 works

# Initialize a dataset
dataset = datasets.load_dataset("oscar", name="unshuffled_deduplicated_fa", split="train")

tokenizer = SentencePieceUnigramTokenizer(unk_token="", eos_token="", pad_token="")

print("len dataset:", len(dataset))

# Build an iterator over this dataset
def batch_iterator(input_sentence_size=None):
if input_sentence_size is None:
input_sentence_size = len(dataset)
batch_length = 100
for i in range(0, input_sentence_size, batch_length):
yield dataset[i: i + batch_length]["text"]

# Train tokenizer
tokenizer.train_from_iterator(
iterator=batch_iterator(input_sentence_size=input_sentence_size),
vocab_size=vocab_size,
show_progress=True,
)

# Save files to disk
tokenizer.save("/content/drive/MyDrive/Pouramini/tokenizer.json")
```
It get stuck in `train_from_iterator` because the size of dataset is large (`input_sentence_size` is around 8M sentences)
How can I divide the dataset and run the code on each block and then merge them to a tokenizer output?

Contributor guide

Open the contributing guide

Research direction

Start with the `SentencePieceUnigramTokenizer.train_from_iterator` call and the `batch_iterator` shown in the issue, then inspect how tokenizer training consumes iterators. Determine whether training can be split across the 8M-sentence dataset and whether resulting tokenizer state can be merged; done means a documented, supported workflow or a clearly identified limitation.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.