DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA2
lmms-eval evaluation
- Dominant language
- Python
- Stars
- 1.3k
- Forks
- 90
- PR merge metrics
- No merged PRs in 30d
Description
Hello, I am currently trying to run various video benchmarks, including the VideoMME and Egoschema, through the lmms-eval evaluation framework (https://github.com/EvolvingLMMs-Lab/lmms-eval).
Since lmm-eval does not support videollama2, I personally implemented videollama2 on the lmm-eval, but failed to reproduce evaluation benchmark results similar to videomme.
It seems that videollama2 should work on the lmms-eval, as indicated in https://github.com/LLaVA-VL/LLaVA-NeXT/blob/main/docs/LLaVA-NeXT-Video_0716.md
I wonder if you have tried to evaluate the models on the lmms-eval and successfully reproduced the performance.
Thanks for the help!
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing the lmms-eval framework and the LLaVA-NeXT-Video evaluation documentation linked in the issue, then compare the attempted VideoLLaMA2 integration with the documented VideoMME and Egoschema evaluation process. Done means determining whether the benchmarks can be reproduced successfully and identifying the cause of any discrepancy.
Written by the indexing model from the issue text.
Assessment
- Tech stack
- python
- Domain
- machine-learning, testing
- Issue type
- Bug
- Difficulty
- 5/5
- Estimated time
- Over a week
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 15/100