DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

Steaming video understanding

Open
#8 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

Hi, thank you for sharing your code! It's a great value to the community.
I am interested in using your model for online video understanding, using a stream of images as opposed to batch processing a video file. Is there some way to achieve this without retraining your model? For example by decomposing a sequence and forwarding latent tokens for context? Essentially I want to be able to comment video streams with TTS.

Cheers,
Theo

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.