deepseek-ai / deepseek-ai/Janus

Cannot Reproduce MMBench-EN Results for Janus-Pro-7B

Open
#205 0 comments 11 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
17.8k
Forks
2.2k
PR merge metrics
No merged PRs in 30d

Description

I am unable to reproduce the MMBench-EN results reported for Janus-Pro-7B. Specifically, when evaluating the model using `lmms-eval`, I obtained the following scores:

- `mmbench_en_dev`: **65.81**
- `mmbench_en_test`: **65.98**

These are significantly lower than the score of **79.2** reported in the paper:
> "Janus-Pro-7B achieved a score of 79.2 on the multimodal understanding benchmark MMBench."

Additionally, the score for Janus-Pro-7B on the [MMBench leaderboard](https://mmbench.opencompass.org.cn/leaderboard) and [opencompass](https://rank.opencompass.org.cn/leaderboard-multimodal) is also relatively low, consistent with my observations.

Could you please clarify:
- What evaluation settings or prompts were used to achieve the reported score of 79.2?
- Is there any preprocessing or special configuration required?
- Are the leaderboard and paper scores based on different evaluation protocols or updated model versions?

Thank you in advance for your help.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.