huggingface / huggingface/datasets

Add web_split dataset for Paraphase and Rephrase benchmark

Open
#2,648 1 comment 1 reaction 1 assignee Claimed by @bhadreshpsavani View on GitHub
enhancement
Dominant language
Python
Stars
22k
Forks
3.4k
Avg merge
5d 7h
Merged PRs (30d)
17

Description

## Describe:
For getting simple sentences from complex sentence there are dataset and task like wiki_split that is available in hugging face datasets. This web_split is a very similar dataset. There some research paper which states that by combining these two datasets we if we train the model it will yield better results on both tests data.

This dataset is made from web NLG data.

All the dataset related details are provided in the below repository

Github link: https://github.com/shashiongithub/Split-and-Rephrase

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.