tensorflow / tensorflow/datasets
[data request] Add DiscoFuse dataset
Open
@sklan is already working on this.
Since Apr 8, 2019.
dataset request
- Dominant language
- Python
- Stars
- 4.6k
- Forks
- 1.6k
- Avg merge
- 3h 54m
- Merged PRs (30d)
- 1
Description
- Name of dataset: DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion
- URL of dataset: https://github.com/google-research-datasets/discofuse
- License of dataset: Creative Commons Attribution-ShareAlike 3.0
- Short description of dataset and use case(s): Sentence fusion is the task of joining several independent sentences into a single coherent text. Current datasets for sentence fusion are small and insufficient for training modern neural models. It's arxiv link is 1902.10526.
Folks who would also like to see this dataset in tensorflow/datasets, please thumbs-up so the developers can know which requests to prioritize.
And if you'd like to contribute the dataset (thank you!), see our guide to adding a dataset.
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Assessment
This issue has not been assessed yet.