bigscience-workshop / bigscience-workshop/catalogue_data

Repeated lines across examples

Open
#6 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
8
Forks
1
PR merge metrics
No merged PRs in 30d

Description

Several datasets have repeated text across examples:
- crawled newspapers tend to have the links to other articles at the bottom, which are nearly always the same
- datasets like the wiki datasets tend to have templates at the start, also always the same.

The difficulty is that some datasets have legitimate repetitions, such as parliamentary proceedings (`lm_en_the_pile_europarl` f.e.)

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.