DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA2

lmms-eval evaluation

Open
#144 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Python
Stars
1.3k
Forks
90
PR merge metrics
No merged PRs in 30d

Description

Hello, I am currently trying to run various video benchmarks, including the VideoMME and Egoschema, through the lmms-eval evaluation framework (https://github.com/EvolvingLMMs-Lab/lmms-eval).

Since lmm-eval does not support videollama2, I personally implemented videollama2 on the lmm-eval, but failed to reproduce evaluation benchmark results similar to videomme.

It seems that videollama2 should work on the lmms-eval, as indicated in https://github.com/LLaVA-VL/LLaVA-NeXT/blob/main/docs/LLaVA-NeXT-Video_0716.md
I wonder if you have tried to evaluate the models on the lmms-eval and successfully reproduced the performance.

Thanks for the help!

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the lmms-eval framework and the LLaVA-NeXT-Video evaluation documentation linked in the issue, then compare the attempted VideoLLaMA2 integration with the documented VideoMME and Egoschema evaluation process. Done means determining whether the benchmarks can be reproduced successfully and identifying the cause of any discrepancy.

Written by the indexing model from the issue text.

Assessment

Tech stack
python
Domain
machine-learning, testing
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
15/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.