DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA2

Can not reproduce results on MVbench.

Open
#171 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.3k
Forks
90
PR merge metrics
No merged PRs in 30d

Description

![Image](https://github.com/user-attachments/assets/a46aa6a6-6e45-4700-bf03-ef1473c8e803)

VideoLLaMA2.1-7B-16F reported is 57.3, but I got 55.3, is there anything wrong?

BTW: I modified here in videollama2_arch.py, otherwise it will be 6 dims, is this problem?

![Image](https://github.com/user-attachments/assets/d46feea7-6a32-4c19-ac71-2f7192a73a61)

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the reported MVbench score for VideoLLaMA2.1-7B-16F and the modification mentioned in videollama2_arch.py. Compare the evaluation setup and model output dimensions with the reported result, then determine whether the change explains the 57.3 versus 55.3 discrepancy. Done means identifying the cause or documenting the exact reproducible evaluation conditions.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
3/5
Estimated time
1-2 days
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
30/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.