DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3
Trainable parameters in each stage
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 89
- PR merge metrics
- No merged PRs in 30d
Description
Thanks for great work!
I have a question about trainable parameters in each stage.
Here's what I think, but is it right?
- Stage 1: Vision Encoder + Projector
- Stage 2: Vision Encoder + Projector + LLM
- Stage 3: Vision Encoder + Video Compressor + Projector + LLM (**I'm curious about this part.**)
- Stage 4: Vision Encoder + Video Compressor + Projector + LLM
Thank you.
Contributor guide
No contributing guide indexed for this repository
Research direction
Start by reviewing the repository's training-stage documentation and Jupyter notebooks to identify which components are trainable in stages 1–4. Done means documenting or confirming the parameter status for every listed component, with an explicit answer for the stage 3 video compressor.
Written by the indexing model from the issue text.
Assessment
- Domain
- documentation, machine-learning
- Issue type
- Documentation
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Activity status
- Stale
- Clarity
- Needs clarification
- Newbie friendliness
- 28/100