bigscience-workshop / bigscience-workshop/data_tooling

Crawling curated list of sites: Data Sourcing Candidate seeds spreadsheet

Đang mở
#299 0 bình luận 0 reaction 0 người được giao Xem trên GitHub
data catalog
Ngôn ngữ chính
HTML
Star
91
Fork
47
Chỉ số merge pull request
Không có pull request nào được merge trong 30 ngày

Mô tả

We want to be able to obtain all web and media content associated with a specific list pre-identified domain names.

This issue tracks potential crawling seeds identified [**BigScience Data Sourcing Participants**](https://docs.google.com/spreadsheets/d/1DNLAGz--qvLh-0qQ7pMPGiNeUMgp-fRgn-8mbLagC7U/edit#gid=513216703), primarily in Spanish and SEA English (and three Chinese).

The steps to follow are:
1. filter the CommonCrawl (or another archive) for all WARC records with one of the given domain names
- filtering all dumps form the last two years
2. obtain overall metrics and metrics per domain name
- page counts, content languages, content types, etc.
3. upload all of the relevant WARC records for each domain name to a HF dataset in the [BigScience Catalogue Data Organization](https://huggingface.co/bigscience-catalogue-data)
- minimal filtering of WARC records to include human-readable pages AND pages that reference links to objects we want to download (e.g. PDFs)
- Extract the HTML tags corresponding to all URLs in the WARC entries
- optional: post-process the above list to identify outgoing links, extract their domain name, and content type
- optional: run text extraction

In particular, the list of domain names mentioned in outgoing link may be used to obtain a "depth 1 pseudo-crawl" by running the same process again

cc @sebastian-nagel

Hướng dẫn đóng góp

Chưa lập chỉ mục được hướng dẫn đóng góp cho kho mã nguồn này

Hướng nghiên cứu

Bắt đầu với bảng tính «BigScience Data Sourcing Participants» được liên kết và các yêu cầu lọc CommonCrawl/WARC. Xác định các chỉ số cần thu thập cho từng domain, lượng nội dung WARC tối thiểu cần giữ lại và đích của dataset trên Hugging Face. Được xem là hoàn tất khi các domain đã chọn đã được xử lý, các chỉ số đã được báo cáo và các bản ghi liên quan đã được tải lên cùng với thông tin URL được yêu cầu.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Đánh giá

Công nghệ
html
Lĩnh vực
data-engineering, web-dev
Loại issue
Tính năng
Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức độ hoạt động
Đình trệ
Độ rõ ràng
Khá rõ ràng
Mức phù hợp với người mới
25/100

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.