DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

How to use video reasoning for video frames?

Open
#51 1 comment 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

Hello author! Thank you very much for your open source and I am looking forward to trying your model for inference. However, I am currently facing a problem where I have extracted my video as frames. If I only have frames, can I still use video for inference? Or can we only use multi graph reasoning?
I have tried to put the frames in a list and input them into the model, but still got the error Number of images does not match the number of image tokens in the text
Looking forward to your reply! Wishing you all the best

Contributor guide

No contributing guide indexed for this repository

Research direction

The issue names no file, test, or entry point. Start by checking the documented inference interface for frame-list inputs and reproduce the reported image-token mismatch. Done means establishing whether extracted frames are supported and recording the required input format or a clear limitation.

Written by the indexing model from the issue text.

Assessment

Tech stack
machine-learning
Domain
machine-learning
Issue type
Bug
Difficulty
5/5
Estimated time
Over a week
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
25/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.