Quality Evals that stress longer decode / 压力测试长解码的质量评估
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 48/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python
- Domain
- machine-learning, testing-qa
Research direction
Locate the existing Eval benchmarks and the dataset or configuration entry points used for GSM8K or MMLU. Add AIME 2024/2025 as a longer-decode quality evaluation without sandbox requirements, then run the benchmark flow and verify that its quantization results are reported alongside the existing evaluations.
Written by the indexing model from the issue text.
Description
I love the new Eval benchmarks idea and think you pick one that is less saturated to expose quantization gaps.
AIME 2024/2025 is less saturated and requires more thinking (but still no sandbox) and exposes more quantization gaps in my experiments than gsm8k or MMLU. It will take slightly longer to run but worth it!
中文说明
建议选择 AIME 2024/2025 作为质量评估基准测试。与 GSM8K 或 MMLU 相比,AIME 不那么饱和,需要更深入的推理(但无需沙盒),能更好地暴露量化精度差异。运行时间稍长但值得。
- Dominant language
- Python
- Stars
- 1.7k
- Forks
- 303
- Avg merge
- 1d 13h
- Merged PRs (30d)
- 284
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from SemiAnalysisAI/InferenceX
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
SemiAnalysisAI/InferenceX#2125 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
SemiAnalysisAI/InferenceX#1587 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 78/100
SemiAnalysisAI/InferenceX#1369 · 3 comments ·
-
Difficulty 1/5 1-3 hours Newbie friendliness 76/100
SemiAnalysisAI/InferenceX#1359 · 1 comment ·
-
Difficulty 5/5 Over a week Newbie friendliness 30/100
SemiAnalysisAI/InferenceX#3122 · 3 comments ·
All issues in SemiAnalysisAI/InferenceX
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
bancolombia/sentinel#23 ·
-
test md OpenCI
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
integration:quickjs org:external priority:backlog topic:code-interpreter topic:middleware type:feature
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
langchain-ai/deepagents#6450 ·
-
bug client
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100