DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA2

How to process multiple consecutive frames for fine-tuning VideoLLaMA2?

Open
#154 1 comment 1 reaction 0 assignees View on GitHub
Dominant language
Python
Stars
1.3k
Forks
90
PR merge metrics
No merged PRs in 30d

Description

Hi, I am working on fine-tuning VideoLLaMA2 on a dataset where video sequences have already been decomposed into sequences of consecutive frames. I would like the model to process a sequence of tens of frames at once, rather than generating a response for each individual frame.

My goal is:
The model should look at a sequence of 30-50 frames and predict.

Currently, I don't know how should the fine-tuning data format looks like, maybe like this? :

```
[
{
"id": 1,
"video_frames": ["frames/001.jpg", "frames/002.jpg", "frames/003.jpg", ..., "frames/030.jpg"],
"conversations": [
{
"from": "human",
"value": "\nWhat type of accident is happening in this sequence?"
},
{
"from": "gpt",
"value": "This sequence depicts a rear-end collision caused by a sudden stop."
}
]
}
]
```

However, I am unsure if this is the best way to structure the input for VideoLLaMA2.

Contributor guide

No contributing guide indexed for this repository

Assessment

This issue has not been assessed yet.

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.