Quality Evals that stress longer decode / 压力测试长解码的质量评估

Open
#1,920 2 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
48/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Quiet
Tech stack
python

Research direction

Locate the existing Eval benchmarks and the dataset or configuration entry points used for GSM8K or MMLU. Add AIME 2024/2025 as a longer-decode quality evaluation without sandbox requirements, then run the benchmark flow and verify that its quantization results are reported alongside the existing evaluations.

Written by the indexing model from the issue text.

Description

I love the new Eval benchmarks idea and think you pick one that is less saturated to expose quantization gaps.
AIME 2024/2025 is less saturated and requires more thinking (but still no sandbox) and exposes more quantization gaps in my experiments than gsm8k or MMLU. It will take slightly longer to run but worth it!

中文说明

建议选择 AIME 2024/2025 作为质量评估基准测试。与 GSM8K 或 MMLU 相比,AIME 不那么饱和,需要更深入的推理(但无需沙盒),能更好地暴露量化精度差异。运行时间稍长但值得。

Dominant language
Python
Stars
1.7k
Forks
303
Avg merge
1d 13h
Merged PRs (30d)
284

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from SemiAnalysisAI/InferenceX

All issues in SemiAnalysisAI/InferenceX

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.