DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3
does learning rate affect much for further finetune?
Open
- Dominant language
- Jupyter Notebook
- Stars
- 1.2k
- Forks
- 89
- PR merge metrics
- No merged PRs in 30d
Description
I'm finetuning beyond stage4 and using the learning rates listed in paper for stage4. I wonder if you have any intuition how I should tune these learning rates (because I have limited computation resource), and whether this would change the performance a lot on video understanding tasks.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.