bigscience-workshop / bigscience-workshop/data-preparation
Question about ROOTS corpus: availability & earlier web data
- Dominant language
- Jupyter Notebook
- Stars
- 318
- Forks
- 42
- PR merge metrics
- No merged PRs in 30d
Description
Hi ROOTS / BigScience,
First, many thanks for ROOTS — it's an awesome multilingual dataset that’s super helpful.
I have a few questions:
Is there a way to access the full ROOTS corpus (beyond the “large initial subset”)? Or is the full version publicly downloadable?
Does anyone know whether ROOTS or related BigScience projects have plans or workflows for collecting web text from before 2008? Any archives, tools, or datasets people have used for that time period?
If I wanted to combine ROOTS with other historical web datasets (or reconstruct earlier web snapshots), would the preprocessing / filtering tools from the data-preparation GitHub repo be helpful for that?
Thanks a lot for any pointers or suggestions!
Best,
Patrick
Contributor guide
No contributing guide indexed for this repository
Research direction
The issue names the ROOTS corpus and the data-preparation GitHub repository, but no file, test, or entry point. Start by reviewing the repository's documented corpus-access and preprocessing resources; done means answering the questions about full-corpus availability, pre-2008 data sources, and whether the tools support combining historical datasets.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, documentation
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 20/100