bigscience-workshop / bigscience-workshop/data_tooling
Create license-compliant version of the Pile
- 主要言語
- HTML
- スター
- 91
- フォーク
- 47
- PR マージ指標
- 30日以内にマージされた PR はありません
説明
As discussed with @StellaAthena, this would be useful to constitute the English-language component of the dataset, possibly augmented by the Spotify transcripts dataset.
The creation of this dataset might be decomposed into smaller subsets, as reported by @StellaAthena (see https://github.com/bigscience-workshop/data_tooling/issues/65#issuecomment-971138275):
- [ ] Pile CC: Unclear, see below
- [x] #74: this was downloaded in a license-compliant fashion
- [ ] Books3: excluded
- [ ] OpenWebText2: Unclear, see below
- [ ] arXiv: needs to be redownloaded and filtered by license
- [ ] GitHub: to be replaced by a license-compliant code dataset compiled by Google
- [ ] #75
- Good as-is, I have acquired permission to use this from the org that owns the data
- [ ] #376
- [x] Good as-is. Dataset link: https://huggingface.co/datasets/the_pile_stack_exchange
- [ ] #297
- Good as-is
- [x] PubMed
- Good as-is
- [ ] #301
- [ ] Project Gutenberg
- Good as-is
- [ ] OpenSubtitles: Excluded. Although their website claims to be license complaint this is an obvious lie. They even posted the script of Wonder Woman before the movie debuted. there’s no way in hell they had Disney’s permission to do that.
- [ ] #311
- Good as-is
- [ ] DM Mathematics
- Good as-is
- [x] Ubuntu IRC
- Good as-is
- [ ] #301
- [ ] BookCorpus2: Excluded
- [x] #378
- Good as-is
- [x] HackerNews
- Good as is
- [ ] #301
- [ ] YouTube Subtitles: excluded
- [ ] PhilPapers: I need to double check but this is either good as-is or needs to be redownloaded and filtered by license
- [x] NIH ExPorter
- Good as-is
- [ ] #301
- [x] #310
- Good as-is
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
評価
この issue はまだ評価されていません。