DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA3

multi images sft

Open
#41 4 comments 0 reactions 0 assignees View on GitHub
Dominant language
Jupyter Notebook
Stars
1.2k
Forks
89
PR merge metrics
No merged PRs in 30d

Description

Thanks for this great work. I want to finetune VideoLLama3 in a scenario where video has been decomposed to multiple images. I construct the dataset jsonl file as followed.
{"image": ["valid_frame/mp4_0.jpg", "valid_frame/mp4_1.jpg", "valid_frame/mp4_2.jpg", "valid_frame/mp4_3.jpg", "valid_frame/mp4_4.jpg"], "conversations": [{"from": "human", "value": "\n"}, {"from": "gpt", "value": "good"}]}
When I finetune this model using stage4 script, I always get this error **IndexError: The shape of the mask [2964] at index 0 does not match the shape of the indexed tensor [14820, 1536] at index 0,** what can I do to solve this issue?

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.