DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

Trainable parameters in each stage

Open
#15 2 comments 1 reaction 0 assignees View on GitHub
good first issue
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

Thanks for great work!

I have a question about trainable parameters in each stage.

Here's what I think, but is it right?
- Stage 1: Vision Encoder + Projector
- Stage 2: Vision Encoder + Projector + LLM
- Stage 3: Vision Encoder + Video Compressor + Projector + LLM (**I'm curious about this part.**)
- Stage 4: Vision Encoder + Video Compressor + Projector + LLM

Thank you.

Contributor guide

No contributing guide indexed for this repository

Research direction

Start by reviewing the repository's training-stage documentation and Jupyter notebooks to identify which components are trainable in stages 1–4. Done means documenting or confirming the parameter status for every listed component, with an explicit answer for the stage 3 video compressor.

Written by the indexing model from the issue text.

Assessment

Domain
documentation, machine-learning
Issue type
Documentation
Difficulty
1/5
Estimated time
Under an hour
Activity status
Stale
Clarity
Needs clarification
Newbie friendliness
28/100

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.