epfl-dlab / epfl-dlab/SynthIE

Rebel clean questions

Open
#7 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
63
Forks
5
PR merge metrics
No merged PRs in 30d

Description

Hello! Thank you for you work. Really good job working on the higher quality datasets.

However, I have a question on the Rebel-clean dataset. In the original paper you mentioned that you manually created the dataset consisting of 360 samples with input texts that fit two specified criteria. But in the rebel dataset on the [hugging face](https://huggingface.co/datasets/martinjosifoski/SynthIE) there are much more data points, several of which by my manual exploration still have the problems that you tried to tackle using filtering by above-mentioned criteria. Can you please clarify, is the rebel version on hugging face supposed to be rebel_clean mentioned in the paper? Otherwise, can you please provide rebel_clean 360 quality data points for the reproducibility of your work?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.