tensorflow / tensorflow/datasets
[data request] WikiText-103
Open
@cuent is already working on this.
Since Mar 1, 2019.
dataset request
- Dominant language
- Python
- Stars
- 4.6k
- Forks
- 1.6k
- Avg merge
- 3h 54m
- Merged PRs (30d)
- 1
Description
- Name of dataset: WikiText-103
- URL of dataset: https://blog.einstein.ai/the-wikitext-long-term-dependency-language-modeling-dataset/
- License of dataset: CC BY-SA 3.0 Unported
- Short description of dataset and use case(s): The WikiText language modeling dataset is a collection of over 100 million tokens extracted from the set of verified Good and Featured articles on Wikipedia.
Folks who would also like to see this dataset in tensorflow/datasets, please +1/thumbs-up so the developers can know which requests to prioritize.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.