bigscience-workshop / bigscience-workshop/data_tooling

Create license-compliant version of the Pile

オープン
#65 コメント 2 件 リアクション 0 件 担当者 1 名 @albertvillanova が担当を希望しています GitHub で見る
data catalog language modeling script
主要言語
HTML
スター
91
フォーク
47
PR マージ指標
30日以内にマージされた PR はありません

説明

As discussed with @StellaAthena, this would be useful to constitute the English-language component of the dataset, possibly augmented by the Spotify transcripts dataset.

The creation of this dataset might be decomposed into smaller subsets, as reported by @StellaAthena (see https://github.com/bigscience-workshop/data_tooling/issues/65#issuecomment-971138275):
- [ ] Pile CC: Unclear, see below
- [x] #74: this was downloaded in a license-compliant fashion
- [ ] Books3: excluded
- [ ] OpenWebText2: Unclear, see below
- [ ] arXiv: needs to be redownloaded and filtered by license
- [ ] GitHub: to be replaced by a license-compliant code dataset compiled by Google
- [ ] #75
- Good as-is, I have acquired permission to use this from the org that owns the data
- [ ] #376
- [x] Good as-is. Dataset link: https://huggingface.co/datasets/the_pile_stack_exchange
- [ ] #297
- Good as-is
- [x] PubMed
- Good as-is
- [ ] #301
- [ ] Project Gutenberg
- Good as-is
- [ ] OpenSubtitles: Excluded. Although their website claims to be license complaint this is an obvious lie. They even posted the script of Wonder Woman before the movie debuted. there’s no way in hell they had Disney’s permission to do that.
- [ ] #311
- Good as-is
- [ ] DM Mathematics
- Good as-is
- [x] Ubuntu IRC
- Good as-is
- [ ] #301
- [ ] BookCorpus2: Excluded
- [x] #378
- Good as-is
- [x] HackerNews
- Good as is
- [ ] #301
- [ ] YouTube Subtitles: excluded
- [ ] PhilPapers: I need to double check but this is either good as-is or needs to be redownloaded and filtered by license
- [x] NIH ExPorter
- Good as-is
- [ ] #301
- [x] #310
- Good as-is

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。