huggingface / huggingface/olm-datasets
CC data Language Splits
Open
- Dominant language
- Python
- Stars
- 179
- Forks
- 21
- PR merge metrics
- No merged PRs in 30d
Description
Thanks a lot for putting this repo together and providing the fresh CC dumps at HF. I was looking for a way to find dataset splits for other languages but couldn't find a way to do it. Are datasets `olm/olm-CC-MAIN-*` monolingual by chance?
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.