deepseek-ai / deepseek-ai/DeepSeek-Math

SFT的数据分布

Open
#11 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
3.4k
Forks
592
PR merge metrics
No merged PRs in 30d

Description

恭喜你们的效果取得了非常好的效果! 我有一个问题想要请教一下各位大佬:

我想了解一下SFT的数据分布。看到training examples 是 776K,但是可能是我对于数据集的估算可能出现了一些问题。English mathematical datasets:GSM8K和MATH部分我看是根据ToRA进行标注的,所以根据ToRA那篇文章的估算应该是69K,MathInstruct 260K 的子集不是特别好估算我就按照200K来估算,Lila-OOD是32.2K。总计300K左右,而且MathInstruct里面的MATH和GSM8K应该会与前面的69K的数据重复。那么Chinese mathematical datasets的数据应该是476K,这个数据集是你们收集,后续会开源的嘛?

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the SFT training-example count and dataset names listed in this issue, including ToRA, MathInstruct, Lila-OOD, GSM8K, and MATH. Verify the arithmetic and identify the source and release status of the reported 476K Chinese mathematical examples. Done means documenting a definitive dataset breakdown and clarifying whether that Chinese data will be open-sourced.

Written by the indexing model from the issue text.

Assessment

Domain
data, machine-learning
Issue type
Documentation
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.