bigscience-workshop / bigscience-workshop/data_tooling

Crawling curated list of sites: BigScience catalog app URLs

未关闭
#298 2 条评论 0 个 reaction 已指派 1 人 已被 @sebastian-nagel 认领 在 GitHub 查看
data catalog
主要语言
HTML
星标
91
派生
47
PR 合并指标
30 天内没有已合并 PR

描述

We want to be able to obtain all web and media content associated with a specific list pre-identified domain names.

This issue tracks domain names identified in the [**BigScience Data Cataloging Event**](https://bigscience.huggingface.co/data-catalogue)

The steps to follow are:
1. filter the CommonCrawl (or another archive) for all WARC records with one of the given domain names
- filtering all dumps form the last two years
2. obtain overall metrics and metrics per domain name
- page counts, content languages, content types, etc.
3. upload all of the relevant WARC records for each domain name to a HF dataset in the [BigScience Catalogue Data Organization](https://huggingface.co/bigscience-catalogue-data)
- minimal filtering of WARC records to include human-readable pages AND pages that reference links to objects we want to download (e.g. PDFs)
- Extract the HTML tags corresponding to all URLs in the WARC entries
- optional: post-process the above list to identify outgoing links, extract their domain name, and content type
- optional: run text extraction

In particular, the list of domain names mentioned in outgoing link may be used to obtain a "depth 1 pseudo-crawl" by running the same process again

贡献指南

这个仓库没有索引到贡献指南

评估

这个 Issue 还没有评估数据。

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。