bigscience-workshop / bigscience-workshop/catalogue_data
Repeated lines across examples
Open
- Dominant language
- Jupyter Notebook
- Stars
- 8
- Forks
- 1
- PR merge metrics
- No merged PRs in 30d
Description
Several datasets have repeated text across examples:
- crawled newspapers tend to have the links to other articles at the bottom, which are nearly always the same
- datasets like the wiki datasets tend to have templates at the start, also always the same.
The difficulty is that some datasets have legitimate repetitions, such as parliamentary proceedings (`lm_en_the_pile_europarl` f.e.)
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.