deepseek-ai / deepseek-ai/DeepSeek-Math

About raw common crawl data

Open
#12 0 comments 2 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.4k
Forks
592
PR merge metrics
No merged PRs in 30d

Description

Hi,

I'm trying to reproduce your paper. However, I find that many math-related contents are filtered out in many popular text extraction pipeline. I'm wondering which version of the common crawl data you used to mined high-quality math contents? Did you use the custom pipeline for web data processing or something more specific? I cannot find any details regarding this in your paper.

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are named. Start by locating the paper materials and any data-processing documentation, then document the Common Crawl version and whether a custom extraction pipeline was used; done means those reproduction details are available.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
data-engineering, documentation
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.