bigscience-workshop / bigscience-workshop/metadata
Create Dataset with metadata
- Dominant language
- Python
- Stars
- 29
- Forks
- 11
- PR merge metrics
- No merged PRs in 30d
Description
Steps:
- [x] pseudo crawl ~10% of C4 web page from Common Crawl @tianjianjiang
- [x] import pseudo crawled dataset on JZ @SaulLu
- [x] run 1st step of extraction:
1. Extract text, HTML head sections, HTML footer sections, HTML Titles section and HTML metadata @SaulLu
2. Change format of URL @SaulLu
3. Extract Timestamp @cccntu @SaulLu
4. Extract Generation Length Sentence @chkla @SaulLu
5. Extract Generation Length Text @chkla @SaulLu
6. Extract Data source @chkla @SaulLu
- [x] run 2nd step of extraction:
1. Extract Website descriptions @shanyas10 @SaulLu
- [x] run 3rd step of extraction:
1. Extract Entities @manandey @SaulLu
2. (option) Extract Entities descriptions @manandey @SaulLu
- [x] run 4th step of extraction:
1. Extract Paragraph @tianjianjiang @SaulLu
* #114
* #125
* annotator (preprocessor) of the metadata
2. Modify entities metadata with paragraph information @manandey @SaulLu
3. Modify generation length with paragraph information @chkla @SaulLu
- [ ] (optional) clean final dataset:
1. Remove empty lines @SaulLu
2. Remove "errors" columns @SaulLu
3. (optional) Gather all metadata into same column @cccntu @timoschick @SaulLu
- [ ] push dataset to Hub @SaulLu
Contributor guide
No contributing guide indexed for this repository
Research direction
No files or tests are named. Start by reviewing the completed extraction steps and the current dataset output, then work through the remaining checklist items: clean empty lines and error columns as appropriate, gather metadata if required, and push the finished dataset to the Hub.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- data-engineering
- Issue type
- Feature
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100