open-compass / open-compass/VLMEvalKit

Weird Evaluation Output when using Qwen2.5-VL-7B-Instruct

Open
#1,148 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Dominant language
Python
Stars
4.4k
Forks
768
Avg merge
1d 10h
Merged PRs (30d)
17

Description

Hi,

Thanks for your contribution to the community for bringing this utility!

I am running into a problem when inference with Qwen2.5-VL models. The outputs for MMStar are as follows,

Infer Qwen2.5-VL-7B-Instruct/MMStar, Rank 2/4:   0%|                                                                                                                                                       | 0/375 [00:00<?, ?it/s]
 A A 144441441111144141454445144414542444224441414241424242424441414244142414241414441414442414141414144441414144414144414141414141414141411414124412444412444124412144412141214412141412141141211412114121141211421211211214412111
41211121112144121111211114412111114114121111121111211111441211112111121111211111411112111121111211112111121111211112111121111211112111114411111111211111111111111111111111111111111111111111111121111111111212112111111112212212212
21221221221221221221221221111221111221111221111221111212212212212122122122112212112121212212121212122121212121212121221211212112121212121212121212121212121211441211111111112122122121121121212121212114212212121212122121212121212
21212112121212122221221211212121221212121121212121221212112121121121212122121212112121121212112121121121212212212112121212122221221211212121221221211211212121221221211211211212212121212122121221211212121221212212112112121212212
21211211212112121212122121211211212112121212121221212112212112112121212212112112112112122121221211211212122122122122122121221121221121212121221221212121221212212112121212212212121212212212122122121221212211221211212121212212212
1122112122121122112121212122212212112211212212122121221212212122122121221221211221121221121221121212121221221211221121212121221221211221121221121221121212121221121211212112121212122222122122121121211211212212112211212112122212112212112112212112122121221221211221121212122122122121221122121121211212211212112112121221221221211211211212122121121121211212212211212211212112112122112112112112122222222222222221222121121122121221221211221121221121221212211212
21121221121221121212212122112211212121221221221211222122112211212211212211212211212212212211222121221122112122112122112122112122112122112122112122112112221212211221121221122212122122121122112212112211212211212211212211212211212
21121221121221121221121221121221121221121221121221151221221221221221221221221121221121221121221121122212212212211221121221122212122121121221121221121121221121221121121212221221221221112222121221212211221211212222222222222222212
2222222                                                                                                                                                                                                                            
Infer Qwen2.5-VL-7B-Instruct/MMStar, Rank 1/4:   0%|▍                                                                                                                                            | 1/375 [00:44<4:37:02, 44.44s/it]
 A A A A  A  A  A  A  A  A  A  A  44  11  11  11  11  11  11  11  11  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1
  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  
1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1 
 1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1
  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  
1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1 
 1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1 1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1 
 1  1 1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1 
 1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1  1
  1  1  1  1  1                        

I have tried other models such as InternVL3-8B and works well. Any suggestions on solution to this? Thanks.

Contributor guide

No contributing guide indexed for this repository

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

Research direction

Reproduce the reported Qwen2.5-VL-7B-Instruct evaluation on MMStar and compare its output with InternVL3-8B, which reportedly works. Trace the inference path used by the evaluation toolkit and identify why Qwen2.5-VL produces repeated token-like output; done means explaining the discrepancy and verifying a corrected evaluation result.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
computer-vision, machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.