open-compass / open-compass/opencompass

[Bug] 关于评测集社区-daily benchmark模块AI解读的问题

Open
#2,339 0 comments 0 reactions 1 assignee View on GitHub

@bittersweet1999 is already working on this.

Since Dec 1, 2025.

Dominant language
Python
Stars
7.5k
Forks
869
Avg merge
17h 52m
Merged PRs (30d)
13

Description

Prerequisite
Type

I'm evaluating with the officially supported tasks/models/datasets.

Environment

暂无

Reproduces the problem - code/configuration sample

暂无

Reproduces the problem - command or script

暂无

Reproduces the problem - error message

暂无

Other information

最近,关注到了司南官网添加了daily benchmark这个模块,帮助非常大。但是在看当中的AI解读的时候,发现一些幻觉现象。
比如,10月30日的Automating Benchmark Design的一文中,AI解读道 “成本黑洞:构建新基准需要数百专家月的人力投入(如MMLU耗资超200万美元)“,发现原文并没有提及MMLU的构建成本,不知道出处何来,还是说这个是大模型的幻觉导致

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.