huggingface / huggingface/datasets

Request for text deduplication feature

Open
#5,877 4 comments 1 reaction 0 assignees View on GitHub
enhancement
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

### Feature request

It would be great if there would be support for high performance, highly scalable text deduplication algorithms as part of the datasets library.

### Motivation

Motivated by this blog post https://huggingface.co/blog/dedup and this library https://github.com/google-research/deduplicate-text-datasets, but slightly frustrated by how its not very easy to work with these tools I am proposing this feature.

### Your contribution

I would be happy to contribute to the development effort of this feature. would love to collaborate with others in the development effort.

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.