huggingface / huggingface/olm-datasets

CC data Language Splits

Open
#6 3 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
179
Forks
21
PR merge metrics
No merged PRs in 30d

Description

Thanks a lot for putting this repo together and providing the fresh CC dumps at HF. I was looking for a way to find dataset splits for other languages but couldn't find a way to do it. Are datasets `olm/olm-CC-MAIN-*` monolingual by chance?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.