DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA2
Cannot reproduce results on vllava datasets
- Dominant language
- Python
- Stars
- 1.3k
- Forks
- 90
- PR merge metrics
- No merged PRs in 30d
Description
Dear authors of VideoLLaMA2,
Thanks for the great work. We tried to reproduce your results on vllava datasets using the latest version of the code. However, we observe a large discrepancy in the three test datasets.
Model | MVBench | Egoschema | ActivityNet | Avg
-- | -- | -- | -- | --
reported | 45.5 | 42.2 | 47.6 | 45.1
reproduced | 44.475 | 38.5 | 43.55 | 42.175
We directly use your code, and follow your instructions to download the vllava datasets as well as three test sets, i.e. MVBench, Egoschema, and ActivityNet.
Can you hint at how you achieved the average 45.1 results?
Best
Yijiang
Contributor guide
No contributing guide indexed for this repository
Research direction
Start with the repository's instructions for downloading the vllava, MVBench, Egoschema, and ActivityNet datasets, then follow the documented evaluation path used to reproduce the reported scores. Compare the reproduced results with the reported averages and identify which evaluation step or dataset preparation accounts for the discrepancy; done means the source of the mismatch is documented or the reported results are reproduced.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning
- Issue type
- Bug
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Activity status
- Stale
- Clarity
- Mostly clear
- Newbie friendliness
- 25/100