bigscience-workshop / bigscience-workshop/data_tooling

Crawling curated list of sites: BigScience catalog app URLs

オープン
#298 コメント 2 件 リアクション 0 件 担当者 1 名 @sebastian-nagel が担当を希望しています GitHub で見る
data catalog
主要言語
HTML
スター
91
フォーク
47
PR マージ指標
30日以内にマージされた PR はありません

説明

We want to be able to obtain all web and media content associated with a specific list pre-identified domain names.

This issue tracks domain names identified in the [**BigScience Data Cataloging Event**](https://bigscience.huggingface.co/data-catalogue)

The steps to follow are:
1. filter the CommonCrawl (or another archive) for all WARC records with one of the given domain names
- filtering all dumps form the last two years
2. obtain overall metrics and metrics per domain name
- page counts, content languages, content types, etc.
3. upload all of the relevant WARC records for each domain name to a HF dataset in the [BigScience Catalogue Data Organization](https://huggingface.co/bigscience-catalogue-data)
- minimal filtering of WARC records to include human-readable pages AND pages that reference links to objects we want to download (e.g. PDFs)
- Extract the HTML tags corresponding to all URLs in the WARC entries
- optional: post-process the above list to identify outgoing links, extract their domain name, and content type
- optional: run text extraction

In particular, the list of domain names mentioned in outgoing link may be used to obtain a "depth 1 pseudo-crawl" by running the same process again

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

評価

この issue はまだ評価されていません。

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。