CarperAI / CarperAI/Code-Pile

Data Processing

Open
#37 6 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
110
Forks
32
PR merge metrics
No merged PRs in 30d

Description

We should follow a similar process to the BigScience workshop's dataset processing. They include many of the tools ready for us to use such as data deduplication, both exact match and near dedup, filtering of low information content examples, removal of potentially hateful documents, and removal of PII.

They have all their tools available and discussions of them here: https://github.com/bigscience-workshop/data_tooling

Here is an initial set of tasks to perform:
- [ ] Filtering of low quality documents
- [ ] Filtering of documents with specific removal words
- [ ] Filtering of exact duplicate content
- [ ] Filtering of near duplicate content
- [ ] Removal of PII

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.