bigscience-workshop / bigscience-workshop/catalogue_data

Removing dataset lm_en_a_million_news_headlines_abc_australia

Open
#8 2 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
8
Forks
1
PR merge metrics
No merged PRs in 30d

Description

I think it would be better to remove this dataset from the list

Here are some random examples of documents, and they are all like that

Doc 0:
adrian bayley minimum prison term extended 10 years over rapes

Doc 1:
egg farm break in

Doc 2:
stoner claims grand prix in portugal

Doc 3:
palau typhoon bopha watch

Doc 4:
dna breakthrough on unsolved rape

Doc 5:
labor says mortgage stress at record high

Doc 6:
concerns raised over carbon capture

Doc 7:
nigeria to set up regional anti boko haram force

Doc 8:
habib says torturers used information from

Doc 9:
mixed bag for wine production

Doc 10:
sue butler said it

Doc 11:
dog hitches ride from queensland to sa

Doc 12:
tests show beach algae harmless

Doc 13:
more support sought for chamber of commerce

Doc 14:
australia india engaged together to stop people

Doc 15:
tim costello on financial crisis

Doc 16:
push to save womens army camp ruins from roe highway extension

Doc 17:
serial rapist convicted over knifepoint attacks

Doc 18:
grandmother lorn cheng jailed for smuggling heroin from cambodia

Doc 19:
fact check bradfield scheme barnaby joyce drought

Doc 20:
gold coast man attacked with tomahawk

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.