huggingface / huggingface/datatrove
Pipeline for data contamination
- Dominant language
- Python
- Stars
- 3.3k
- Forks
- 302
- Avg merge
- 2h 18m
- Merged PRs (30d)
- 2
Description
Hi guys,
How can I implement a pipeline to check for contamination between two different datasets (e.g., pre-training vs. fine-tuning datasets) and eventually delete marked documents from the pre-training dataset?
Thanks.
Contributor guide
No contributing guide indexed for this repository
Research direction
No files, tests, or entry points are named. Start by reviewing datatrove's customizable pipeline processing blocks and dataset handling, then clarify how contamination is detected, how documents are marked, and whether deletion is required. Done should include an agreed pipeline design and verified behavior for identifying and removing contaminated pre-training documents.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data-engineering, machine-learning
- Issue type
- Feature
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100