deepseek-ai / deepseek-ai/Janus
Cannot Reproduce MMBench-EN Results for Janus-Pro-7B
- Dominant language
- Python
- Stars
- 17.8k
- Forks
- 2.2k
- PR merge metrics
- No merged PRs in 30d
Description
I am unable to reproduce the MMBench-EN results reported for Janus-Pro-7B. Specifically, when evaluating the model using `lmms-eval`, I obtained the following scores:
- `mmbench_en_dev`: **65.81**
- `mmbench_en_test`: **65.98**
These are significantly lower than the score of **79.2** reported in the paper:
> "Janus-Pro-7B achieved a score of 79.2 on the multimodal understanding benchmark MMBench."
Additionally, the score for Janus-Pro-7B on the [MMBench leaderboard](https://mmbench.opencompass.org.cn/leaderboard) and [opencompass](https://rank.opencompass.org.cn/leaderboard-multimodal) is also relatively low, consistent with my observations.
Could you please clarify:
- What evaluation settings or prompts were used to achieve the reported score of 79.2?
- Is there any preprocessing or special configuration required?
- Are the leaderboard and paper scores based on different evaluation protocols or updated model versions?
Thank you in advance for your help.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.