deepseek-ai / deepseek-ai/DeepSeek-Math
SFT的数据分布
- Dominant language
- Python
- Stars
- 3.4k
- Forks
- 592
- PR merge metrics
- No merged PRs in 30d
Description
恭喜你们的效果取得了非常好的效果! 我有一个问题想要请教一下各位大佬:
我想了解一下SFT的数据分布。看到training examples 是 776K,但是可能是我对于数据集的估算可能出现了一些问题。English mathematical datasets:GSM8K和MATH部分我看是根据ToRA进行标注的,所以根据ToRA那篇文章的估算应该是69K,MathInstruct 260K 的子集不是特别好估算我就按照200K来估算,Lila-OOD是32.2K。总计300K左右,而且MathInstruct里面的MATH和GSM8K应该会与前面的69K的数据重复。那么Chinese mathematical datasets的数据应该是476K,这个数据集是你们收集,后续会开源的嘛?
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the SFT training-example count and dataset names listed in this issue, including ToRA, MathInstruct, Lila-OOD, GSM8K, and MATH. Verify the arithmetic and identify the source and release status of the reported 476K Chinese mathematical examples. Done means documenting a definitive dataset breakdown and clarifying whether that Chinese data will be open-sourced.
Written by the indexing model from the issue text.
Assessment
- Domain
- data, machine-learning
- Issue type
- Documentation
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100