EleutherAI / EleutherAI/pilev2

Documents scraped from Open Directories

Open
#11 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
13
Forks
10
PR merge metrics
No merged PRs in 30d

Description

https://odcrawler.xyz/
You have the ability to search by document type, you should get all the html, htm, pdf, doc, docx, json, xls, xlsx, java, xml, js, css, py, ppt, pptx, txt, csv, md, odf, vtt, srt, and tex files in the Collection and scrape each and every one of them, including the directory listings themselves, and make sure to integrate their names. Make sure to exclude certain directories that are datasets themselves, which will be added separately.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.