DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA2

Cannot reproduce results on vllava datasets

Open
#81 27 comments 0 reactions 0 assignees View on GitHub
good first issue
Dominant language
Python
Stars
1.3k
Forks
90
PR merge metrics
No merged PRs in 30d

Description

Dear authors of VideoLLaMA2,
Thanks for the great work. We tried to reproduce your results on vllava datasets using the latest version of the code. However, we observe a large discrepancy in the three test datasets.

Model | MVBench | Egoschema | ActivityNet | Avg
-- | -- | -- | -- | --
reported | 45.5 | 42.2 | 47.6 | 45.1
reproduced | 44.475 | 38.5 | 43.55 | 42.175

We directly use your code, and follow your instructions to download the vllava datasets as well as three test sets, i.e. MVBench, Egoschema, and ActivityNet.

Can you hint at how you achieved the average 45.1 results?

Best
Yijiang

Contributor guide

No contributing guide indexed for this repository

Research direction

Start with the repository's instructions for downloading the vllava, MVBench, Egoschema, and ActivityNet datasets, then follow the documented evaluation path used to reproduce the reported scores. Compare the reproduced results with the reported averages and identify which evaluation step or dataset preparation accounts for the discrepancy; done means the source of the mismatch is documented or the reported results are reproduced.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning
Issue type
Bug
Difficulty
4/5
Estimated time
3-5 days
Activity status
Stale
Clarity
Mostly clear
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.