DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3
could you tell me which part of code is implemented for "Any-resolutionVisionTokenization(AVT)"
Open
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 89
- PR merge metrics
- No merged PRs in 30d
Description
Does this setting really activate "Any-resolutionVisionTokenization(AVT)" fashion?
https://github.com/DAMO-NLP-SG/VideoLLaMA3/blob/0898cf8092b852697657b9154bcd32ed30cd79b7/videollama3/train.py#L118C5-L118C23
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.