deepseek-ai / deepseek-ai/DeepSeek-Math
Question about the way to extract text from CC HTML
- Dominant language
- Python
- Stars
- 3.4k
- Forks
- 592
- PR merge metrics
- No merged PRs in 30d
Description
Hi guys @DeepSeekPH , thanks so much for sharing such an excellent work. I note that Openwebmath uses a specialized pipeline to extract content from HTML instead of directing using the WET file from Common Crawl. I just wonder how you guys deal with this problem? Do you also follow openwebmath to process the html with a private diagram? sincerely wait for your feedback.
Contributor guide
No contributing guide indexed for this repository
Research direction
No files, tests, or entry points are mentioned. Clarify whether the project processes Common Crawl HTML or WET data and document the extraction pipeline, including whether a private diagram is involved. Completion requires a maintainer answer or linked documentation describing the approach.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- html
- Domain
- data-engineering
- Issue type
- Documentation
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100