DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3
Steaming video understanding
Open
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 89
- PR merge metrics
- No merged PRs in 30d
Description
Hi, thank you for sharing your code! It's a great value to the community.
I am interested in using your model for online video understanding, using a stream of images as opposed to batch processing a video file. Is there some way to achieve this without retraining your model? For example by decomposing a sequence and forwarding latent tokens for context? Essentially I want to be able to comment video streams with TTS.
Cheers,
Theo
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.