huggingface / huggingface/datatrove

Pipeline for data contamination

Open
#373 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.3k
Forks
302
Avg merge
2h 18m
Merged PRs (30d)
2

Description

Hi guys,

How can I implement a pipeline to check for contamination between two different datasets (e.g., pre-training vs. fine-tuning datasets) and eventually delete marked documents from the pre-training dataset?

Thanks.

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are named. Start by reviewing datatrove's customizable pipeline processing blocks and dataset handling, then clarify how contamination is detected, how documents are marked, and whether deletion is required. Done should include an agreed pipeline design and verified behavior for identifying and removing contaminated pre-training documents.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering, machine-learning
Issue type
Feature
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.