open-compass / open-compass/VLMEvalKit

Can't Reproduce the results of Qwen2_5_vl -3B on MMMU_DEV_VAL and AI2D_TEST

Open
#1,092 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
4.4k
Forks
768
Avg merge
1d 10h
Merged PRs (30d)
17

Description

My result on MMMU is 50.67, while the official score is 53.1. On the AI2D_TEST set, the official result is 81.6, and mine is 80.53. I have confirmed that I used the OpenAI API key with Transformers version 4.49.0.

-----------------------------------  -------------------  -------------------
split                                dev                  validation
Overall                              0.4866666666666667   0.5066666666666667
Accounting                           0.6                  0.4
Agriculture                          0.4                  0.5666666666666667
Architecture_and_Engineering         0.0                  0.23333333333333334
Art                                  0.4                  0.5666666666666667
Art_Theory                           0.8                  0.7333333333333333
Basic_Medical_Science                0.8                  0.6666666666666666
Biology                              0.6                  0.4666666666666667
Chemistry                            0.4                  0.36666666666666664
Clinical_Medicine                    0.4                  0.5333333333333333
Computer_Science                     0.6                  0.5666666666666667
Design                               0.6                  0.7666666666666667
Diagnostics_and_Laboratory_Medicine  0.6                  0.4666666666666667
Economics                            0.6                  0.36666666666666664
Electronics                          0.4                  0.5666666666666667
Energy_and_Power                     0.4                  0.3
Finance                              0.8                  0.4
Geography                            0.0                  0.5
History                              1.0                  0.7333333333333333
Literature                           0.8                  0.8333333333333334
Manage                               0.4                  0.5
Marketing                            0.0                  0.5666666666666667
Materials                            0.4                  0.36666666666666664
Math                                 0.6                  0.36666666666666664
Mechanical_Engineering               0.4                  0.3
Music                             

split                      none
Overall                    0.8053756476683938
atomStructure              0.75
eclipses                   0.9354838709677419
faultsEarthquakes          0.7142857142857143
foodChainsWebs             0.8950086058519794
lifeCycles                 0.7793764988009593
moonPhaseEquinox           0.6642599277978339
partsOfA                   0.7967145790554415
partsOfTheEarth            0.7884615384615384
photosynthesisRespiration  0.7215189873417721
rockCycle                  0.6567164179104478
rockStrata                 0.7073170731707317
solarSystem                0.8888888888888888
typesOf                    0.7346938775510204
volcano                    0.8125
waterCNPCycle              0.6136363636363636

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Start by reproducing the reported Qwen2.5-VL-3B evaluations on MMMU_DEV_VAL and AI2D_TEST with Transformers 4.49.0, using the reported OpenAI API setup. Compare the resulting scores with the official MMMU and AI2D_TEST values and identify whether the discrepancy is reproducible; the issue names no repository file or test to inspect.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, testing-qa
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.