open-compass / open-compass/VLMEvalKit
Ovis1.5-Llama3-8B在Hallusion Bench上的指标和榜单上的指标差距过大
Open
Nobody has claimed this yet.
- Dominant language
- Python
- Stars
- 4.4k
- Forks
- 768
- Avg merge
- 1d 10h
- Merged PRs (30d)
- 17
Description
1、OpenCompass排行榜的指标是45,但是我们本地测试只有41.30
2、这个差距不是由评判模型造成的。因为需要评判模型处理的'unknown'预测只有14个问题,而这14个问题本身就不是Yes/No问题,我参考了官方给出的预测结果,这14个问题同样回答错误。
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
Research direction
Start by reproducing the Ovis1.5-Llama3-8B evaluation on Hallusion Bench and comparing the local 41.30 result with the OpenCompass leaderboard value of 45. Check the handling of the 14 “unknown” predictions against the official predictions; done means identifying and fixing or clearly explaining the metric discrepancy.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, testing-qa
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 25/100