DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

could you tell me which part of code is implemented for "Any-resolutionVisionTokenization(AVT)"

Open
#53 0 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

Does this setting really activate "Any-resolutionVisionTokenization(AVT)" fashion?

https://github.com/DAMO-NLP-SG/VideoLLaMA3/blob/0898cf8092b852697657b9154bcd32ed30cd79b7/videollama3/train.py#L118C5-L118C23

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.