DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3
What is the max_frames for inference on long videos in the paper?
Open
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 89
- PR merge metrics
- No merged PRs in 30d
Description
Thank you for the excellent work.
I saw that in the example code, max_frames is set to 128, but when I use this parameter, I encounter an out-of-memory error. I am using an 80GB A800 GPU.
To replicate the results in your paper's tables, how many frames should I use? I am planning to reproduce the performance of VideoLlama3 on long video datasets.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.