bigscience-workshop / bigscience-workshop/metadata

Create Dataset with metadata

Open
#124 0 comments 3 reactions 0 assignees View on GitHub
#dataset Epic
Dominant language
Python
Stars
29
Forks
11
PR merge metrics
No merged PRs in 30d

Description

Steps:
- [x] pseudo crawl ~10% of C4 web page from Common Crawl @tianjianjiang
- [x] import pseudo crawled dataset on JZ @SaulLu
- [x] run 1st step of extraction:
1. Extract text, HTML head sections, HTML footer sections, HTML Titles section and HTML metadata @SaulLu
2. Change format of URL @SaulLu
3. Extract Timestamp @cccntu @SaulLu
4. Extract Generation Length Sentence @chkla @SaulLu
5. Extract Generation Length Text @chkla @SaulLu
6. Extract Data source @chkla @SaulLu
- [x] run 2nd step of extraction:
1. Extract Website descriptions @shanyas10 @SaulLu
- [x] run 3rd step of extraction:
1. Extract Entities @manandey @SaulLu
2. (option) Extract Entities descriptions @manandey @SaulLu
- [x] run 4th step of extraction:
1. Extract Paragraph @tianjianjiang @SaulLu
* #114
* #125
* annotator (preprocessor) of the metadata
2. Modify entities metadata with paragraph information @manandey @SaulLu
3. Modify generation length with paragraph information @chkla @SaulLu
- [ ] (optional) clean final dataset:
1. Remove empty lines @SaulLu
2. Remove "errors" columns @SaulLu
3. (optional) Gather all metadata into same column @cccntu @timoschick @SaulLu
- [ ] push dataset to Hub @SaulLu

Contributor guide

No contributing guide indexed for this repository

Research direction

No files or tests are named. Start by reviewing the completed extraction steps and the current dataset output, then work through the remaining checklist items: clean empty lines and error columns as appropriate, gather metadata if required, and push the finished dataset to the Hub.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering
Issue type
Feature
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.