huggingface / huggingface/datasets

Excessive RAM Usage After Dataset Concatenation concatenate_datasets

Open
#7,373 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Describe the bug

When loading a dataset from disk, concatenating it, and starting the training process, the RAM usage progressively increases until the kernel terminates the process due to excessive memory consumption.

https://github.com/huggingface/datasets/issues/2276

### Steps to reproduce the bug

```python
from datasets import DatasetDict, concatenate_datasets

dataset = DatasetDict.load_from_disk("data")

...
...

combined_dataset = concatenate_datasets(
[dataset[split] for split in dataset]
)

#start SentenceTransformer training
```

### Expected behavior

I would not expect RAM utilization to increase after concatenation. Removing the concatenation step resolves the issue

### Environment info

sentence-transformers==3.1.1
datasets==3.2.0

python3.10

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.