DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

What is the max_frames for inference on long videos in the paper?

Open
#4 2 comments 1 reaction 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

Thank you for the excellent work.

I saw that in the example code, max_frames is set to 128, but when I use this parameter, I encounter an out-of-memory error. I am using an 80GB A800 GPU.

To replicate the results in your paper's tables, how many frames should I use? I am planning to reproduce the performance of VideoLlama3 on long video datasets.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.