google-research / google-research/language

CC-news reproduction

Open
#70 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.8k
Forks
362
PR merge metrics
No merged PRs in 30d

Description

Dear authors,

I want to use the CC-news dataset to train my model.

Now, I use https://github.com/fhamborg/news-please to construct CC-news corpus from CC.

But I don't know if it's the right way to obtain CC-news by directing running

`python3 -m newsplease.examples.commoncrawl`

Questions:

- Did you use the same tool? Did you add some extra filtering rules?
- Since you can't release the CC-news corpus, could you tell me how to collect CC-news corpus by myself?

Contributor guide

Open the contributing guide

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.