bigscience-workshop / bigscience-workshop/data_tooling
Create license-compliant version of the Pile
- 主要语言
- HTML
- 星标
- 91
- 派生
- 47
- PR 合并指标
- 30 天内没有已合并 PR
描述
As discussed with @StellaAthena, this would be useful to constitute the English-language component of the dataset, possibly augmented by the Spotify transcripts dataset.
The creation of this dataset might be decomposed into smaller subsets, as reported by @StellaAthena (see https://github.com/bigscience-workshop/data_tooling/issues/65#issuecomment-971138275):
- [ ] Pile CC: Unclear, see below
- [x] #74: this was downloaded in a license-compliant fashion
- [ ] Books3: excluded
- [ ] OpenWebText2: Unclear, see below
- [ ] arXiv: needs to be redownloaded and filtered by license
- [ ] GitHub: to be replaced by a license-compliant code dataset compiled by Google
- [ ] #75
- Good as-is, I have acquired permission to use this from the org that owns the data
- [ ] #376
- [x] Good as-is. Dataset link: https://huggingface.co/datasets/the_pile_stack_exchange
- [ ] #297
- Good as-is
- [x] PubMed
- Good as-is
- [ ] #301
- [ ] Project Gutenberg
- Good as-is
- [ ] OpenSubtitles: Excluded. Although their website claims to be license complaint this is an obvious lie. They even posted the script of Wonder Woman before the movie debuted. there’s no way in hell they had Disney’s permission to do that.
- [ ] #311
- Good as-is
- [ ] DM Mathematics
- Good as-is
- [x] Ubuntu IRC
- Good as-is
- [ ] #301
- [ ] BookCorpus2: Excluded
- [x] #378
- Good as-is
- [x] HackerNews
- Good as is
- [ ] #301
- [ ] YouTube Subtitles: excluded
- [ ] PhilPapers: I need to double check but this is either good as-is or needs to be redownloaded and filtered by license
- [x] NIH ExPorter
- Good as-is
- [ ] #301
- [x] #310
- Good as-is
贡献指南
这个仓库没有索引到贡献指南
评估
这个 Issue 还没有评估数据。