bigscience-workshop / bigscience-workshop/data-preparation

Question about ROOTS corpus: availability & earlier web data

Open
#45 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
318
Forks
42
PR merge metrics
No merged PRs in 30d

Description

Hi ROOTS / BigScience,

First, many thanks for ROOTS — it's an awesome multilingual dataset that’s super helpful.

I have a few questions:

Is there a way to access the full ROOTS corpus (beyond the “large initial subset”)? Or is the full version publicly downloadable?

Does anyone know whether ROOTS or related BigScience projects have plans or workflows for collecting web text from before 2008? Any archives, tools, or datasets people have used for that time period?

If I wanted to combine ROOTS with other historical web datasets (or reconstruct earlier web snapshots), would the preprocessing / filtering tools from the data-preparation GitHub repo be helpful for that?

Thanks a lot for any pointers or suggestions!

Best,
Patrick

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names the ROOTS corpus and the data-preparation GitHub repository, but no file, test, or entry point. Start by reviewing the repository's documented corpus-access and preprocessing resources; done means answering the questions about full-corpus availability, pre-2008 data sources, and whether the tools support combining historical datasets.

Written by the indexing model from the issue text.

Assessment

Domain
data, documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
20/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.