DAMO-NLP-SG / DAMO-NLP-SG/VideoLLaMA2
How to process multiple consecutive frames for fine-tuning VideoLLaMA2?
- Dominant language
- Python
- Stars
- 1.3k
- Forks
- 90
- PR merge metrics
- No merged PRs in 30d
Description
Hi, I am working on fine-tuning VideoLLaMA2 on a dataset where video sequences have already been decomposed into sequences of consecutive frames. I would like the model to process a sequence of tens of frames at once, rather than generating a response for each individual frame.
My goal is:
The model should look at a sequence of 30-50 frames and predict.
Currently, I don't know how should the fine-tuning data format looks like, maybe like this? :
```
[
{
"id": 1,
"video_frames": ["frames/001.jpg", "frames/002.jpg", "frames/003.jpg", ..., "frames/030.jpg"],
"conversations": [
{
"from": "human",
"value": "\nWhat type of accident is happening in this sequence?"
},
{
"from": "gpt",
"value": "This sequence depicts a rear-end collision caused by a sudden stop."
}
]
}
]
```
However, I am unsure if this is the best way to structure the input for VideoLLaMA2.
Contributor guide
No contributing guide indexed for this repository
Assessment
This issue has not been assessed yet.