deepseek-ai / deepseek-ai/DeepSeek-Math

Question about the way to extract text from CC HTML

Open
#18 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.4k
Forks
592
PR merge metrics
No merged PRs in 30d

Description

Hi guys @DeepSeekPH , thanks so much for sharing such an excellent work. I note that Openwebmath uses a specialized pipeline to extract content from HTML instead of directing using the WET file from Common Crawl. I just wonder how you guys deal with this problem? Do you also follow openwebmath to process the html with a private diagram? sincerely wait for your feedback.

Contributor guide

No contributing guide indexed for this repository

Research direction

No files, tests, or entry points are mentioned. Clarify whether the project processes Common Crawl HTML or WET data and document the extraction pipeline, including whether a private diagram is involved. Completion requires a maintainer answer or linked documentation describing the approach.

Written by the indexing model from the issue text.

Assessment

Tech stack
html
Domain
data-engineering
Issue type
Documentation
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.