bigscience-workshop / bigscience-workshop/data_tooling
Create license-compliant version of the Pile
- Dominant language
- HTML
- Stars
- 91
- Forks
- 47
- PR merge metrics
- No merged PRs in 30d
Description
As discussed with @StellaAthena, this would be useful to constitute the English-language component of the dataset, possibly augmented by the Spotify transcripts dataset.
The creation of this dataset might be decomposed into smaller subsets, as reported by @StellaAthena (see https://github.com/bigscience-workshop/data_tooling/issues/65#issuecomment-971138275):
- [ ] Pile CC: Unclear, see below
- [x] #74: this was downloaded in a license-compliant fashion
- [ ] Books3: excluded
- [ ] OpenWebText2: Unclear, see below
- [ ] arXiv: needs to be redownloaded and filtered by license
- [ ] GitHub: to be replaced by a license-compliant code dataset compiled by Google
- [ ] #75
- Good as-is, I have acquired permission to use this from the org that owns the data
- [ ] #376
- [x] Good as-is. Dataset link: https://huggingface.co/datasets/the_pile_stack_exchange
- [ ] #297
- Good as-is
- [x] PubMed
- Good as-is
- [ ] #301
- [ ] Project Gutenberg
- Good as-is
- [ ] OpenSubtitles: Excluded. Although their website claims to be license complaint this is an obvious lie. They even posted the script of Wonder Woman before the movie debuted. there’s no way in hell they had Disney’s permission to do that.
- [ ] #311
- Good as-is
- [ ] DM Mathematics
- Good as-is
- [x] Ubuntu IRC
- Good as-is
- [ ] #301
- [ ] BookCorpus2: Excluded
- [x] #378
- Good as-is
- [x] HackerNews
- Good as is
- [ ] #301
- [ ] YouTube Subtitles: excluded
- [ ] PhilPapers: I need to double check but this is either good as-is or needs to be redownloaded and filtered by license
- [x] NIH ExPorter
- Good as-is
- [ ] #301
- [x] #310
- Good as-is
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.